This episode of Make Me Data Literate features Associate Professor Nicole White discussing the nuances and assumptions that can impact data collection and analysis.
“When we collect data, you know, it’s never purely objective. It’s it’s sort of like this very complex series of assumptions and decisions that we make along the way”
“The thing that I think was really missing from my education was what happens when you’ve got messy data? What happens when it’s not, you know, a nice clean example where you know you don’t have this black and white answer. “
“Subjectivity rules the world, I think, even data.“
Linda McIver (00:00)
Welcome back to another episode of Make Me Data Literate. I am back on my strange habit of nabbing people at conferences and things and and saying, Please be on my podcast. and I think this is gonna be a really fun chat. we have I think a lot of a lot of things in common. So welcome, Associate Professor Nicole White.
Nicole White (00:26)
Hi Linda. Thanks for having me. Looking forward to the chat.
Linda McIver (00:29)
Yeah, I think it’s gonna be fun. so can you tell us who are you and what do you do?
Nicole White (00:34)
Sure. So my name’s Nicole White. I’m a researcher and statistician based at the Queensland University of Technology here in Brisbane. as a statistician, like many, I wear quite a few hats in my day-to-day work. so a lot of my time is spent working with hospital decision makers and healthcare workers, clinicians, and so on, to help design clinical trials. so in my particular space, I’m interested in new initiatives or programs that help to reduce hospital infections for patients. So it’s a very topical area of work and really important too, because the last thing you want when you go to hospital is to get an infection unexpectedly and prolong your stay and increase your risk of all sorts of nasty things. So that’s predominantly where I work day to day.
The other hat that I like to wear is as a statistician, I’m really passionate about how people report their analyses and conduct their analyses in general as researchers and try to find ways of how we can do that better. So we’re more transparent about the data we’re reporting, the types of assumptions that we’ve made when we’ve tried to analyse and interpret that data. And to also also make sure that it’s useful and reproducible by other people as well. I call that my side hustle, but that’s sort of my second hat that I wear.
Linda McIver (01:58)
I love that. It’s so I have a lot of feelings about clinical trials that don’t measure up and and the way we handle data. And
Nicole White (02:06)
Me too.
Linda McIver (02:07)
I think having it open and transparent and, you know, reproducible is something we haven’t always done very well in the past. So it’s great to see such focused and passionate efforts directed to-
Nicole White (02:24)
Yeah.
Linda McIver (02:24)
-solving that problem.
Nicole White (02:26)
Yeah, absolutely. I mean it’s it’s always been, I guess, a very heavily invested area of research and with good reason too. I mean, when we talk about, you know, how strong does evidence need to be about whether it’s the effectiveness of a new treatment or, you know, a change in the healthcare system to help patients move through more smoothly. having sort of that randomization or that clinical trial is really considered to be, you know, the gold standard, the strongest evidence that we can actually collect.
But there’s so much that goes into that too, like even from you know, even just deciding on the question and how to measure the impact of new interventions or treatments. it’s a a lot of nuances in that. but yeah, it can go wrong in a lot of places too. yeah.
Linda McIver (03:14)
We’ll I I wanna I wanna chase that up, but let’s come back to that so that I don’t lose my thread in the questions. but that’s
Nicole White (03:19)
Yeah, let’s circle back.
Linda McIver (03:23)
yeah, I really want to hear more about how it goes wrong and and how we fix that. what did you have to learn to do your work? was there anything missing from your formal education that you had to kind of teach yourself or, you know, in your-
Nicole White (03:35)
Yeah.
Linda McIver (03:36)
-PhD or in becoming an associate professor?
Nicole White (03:40)
Sure. I mean it’s been a long road. I’ve sort of been in this gig for about fifteen years now. but in terms of I guess my formal training, I started off with a mathematics degree, as an undergraduate. So I sort of got sort of I was very interested in maths and data in general and sort of drilled down into statistics sort of towards my later years of undergraduate. and really just followed my nose after that in terms of, you know, going through those sort of the research qualifications of honours and and PhD eventually. I think just having sort of that grounding in mathematics, my training was predominantly sort of methods and theory. so you know, whenever sort of learning about new theory, like particularly statistics, like building methods from the ground up, it was always demonstrated on like really clean examples of data. Because, you know, as a student, you you wanna you wanna get a handle on the skills and sort of how everything works. So when you’re out in the big wide world, you can you can just roll with it and you know what you’re doing. I mean, that’s the idea, right?
Linda McIver (04:39)
Yeah.
Yep.
Nicole White (04:41)
I mean, reflecting on sort of my job and even sort of people that I work with now, the thing that I think was really missing from my education was what happens when you’ve got messy data? What happens when it’s not, you know, a nice clean example where you know you don’t have this black and white answer? And I think particularly when I’ve sort of first started out as a researcher, I still had that mindset of there’s a very there’s just one approach, it’s very clean, you follow the steps.
And you get the answer. but yeah, it’s much more nuanced than that, I think, because particularly with health data, people be people, you know, they say don’t-
Linda McIver (05:19)
Ha ha.
Nicole White (05:19)
-work with animals, children, or humans, don’t you know, we’re very-
Linda McIver (05:22)
Yeah.
Nicole White (05:23)
-unpredictable creatures, right? So it doesn’t really fit that,
Linda McIver (05:25)
Chaotic.
Nicole White (05:27)
yeah, that sort of procedural approach that people think statistics is and particularly in terms of managing data to get it to the point that you can analyze it as well.
Yeah, so that unpredictable element I think is something that I wish I’d had exposure to earlier, I guess.
Linda McIver (05:45)
You are so singing my song there. I get very-
Nicole White (05:47)
Yeah.
Linda McIver (05:48)
-intense about pointlessly clean toy data sets ’cause I f they’re teaching all of the wrong kind of attitudes, if not skills. I actually just put together some examples for ACARA, the curriculum authority in Australia,
Nicole White (06:03)
Right.
Linda McIver (06:05)
for their senior maths curriculum and one of the examples I found for normal distributions was the Southern Oscillation Index, which is wonderful-
Nicole White (06:16)
Okay.
Linda McIver (06:17)
-because the Bureau of Meteorology has you can get the last 10 years, the last 20 years, the last 50 years, or all the way back to I think it was 1867. and every one of those has a normal distribution, but it’s different. Like it’s every one of them was wonky-
Nicole White (06:36)
nice.
Linda McIver (06:36)
-every one is slightly different. So it’s just the most beautiful illustration of real data doing its real thing. one of my favorite sayings is there’s no such thing as a perfect d data set. Like that real data is-
Nicole White (06:50)
I love that.
Linda McIver (06:50)
-always broken, so you gotta ask what’s wrong with it, or how is it, you know, messy or complicated and
I we can be doing that from the start. There’s no reason we we should be giving clean data sets when we’re teaching this stuff. Even even to five year olds.
Nicole White (07:06)
No, absolutely.
I mean, I think my if I I don’t teach these days I’m full time research, but if I had my time as sort of in a teaching role again, I think my ideal assignment for even for first year a first year statistics course is here’s a data set. It’s super messy. How how would you approach it? I mean, there’s lots of different ways of approaching, you know, the same data set, the same problem. and you know, sort of something related to the messy data is when we collect data, you know, it’s never purely objective. It’s it’s sort of like this very complex series of assumptions and decisions that we make along the way around even, you know, how do you collect you know, discharge from hospital? Is
Linda McIver (07:50)
Mm.
Nicole White (07:51)
it a simple, yes, they were discharged, no they weren’t. Or is it when were they discharged? Is it the time of day? Is it the exact date? Is it down to the second? You know, there’s all these little mini trade offs that go on when we collect data.
Linda McIver (08:05)
Yeah.
Nicole White (08:05)
that you know, make it more important or detailed and informative, of course, but it makes it messy too. So sort of yeah, sort of my my soapbox that I stand on is data doesn’t speak for itself. You know, the context and how it’s collected before you even receive it is just absolutely everything. And understanding that is actually a big part of my job as well.
Linda McIver (08:27)
I I love that. That is Yep. It’s it’s it’s really, really important. and you know, just down to the definitions. Wha what what actually does discharge mean? You know, are we are we counting people who leave on their own? Are we counting people who, got fed up and left before the papers were done? Are we counting people who were officially told they could go home?
Nicole White (08:50)
Exactly.
Linda McIver (08:50)
You know, does that change the the definition?
We we need to be thinking about that and thinking about how we collect it, what we’re collecting, why we’re collecting it, all of those things that you don’t get from a toy data set.
Nicole White (09:05)
Yeah, exactly. And then at the other end being transparent about that too. Like it’s one thing to read, you know, in a report or a publication, we followed people up to discharge or discharge was our outcome. It’s like, well, what what’s encompassed in that? what assumptions are you br are you combining very different people together, for instance?
Linda McIver (09:25)
Yeah.
Yeah. And and what do you assume about the state of someone when they’re discharged? You know, are you assuming that they’re-
Nicole White (09:32)
Exactly.
Linda McIver (09:33)
-they’re that means they’re healthy and, you know, we’re all done? Or does it like that yeah. So much-
Nicole White (09:39)
Yeah. Precisely
Linda McIver (09:41)
-richness and complexity to to human data, to any real data, really. it’s
Nicole White (09:47)
Yeah, definitely not limited to humans at all.
Linda McIver (09:51)
No, no, very much so. Is there one thing that you wish everybody knew about data? That, you know, if if people just got this one thing you think the world would be better or your job would be easier or something?
Nicole White (10:09)
That is a good question. I thought I yeah, in prep preparing for this, I thought I thought long and hard about it. I think, yeah, as I s I guess as I said before, you know, and this goes back to even my days as an early career researcher, this idea that I’ll just let the data speak for itself. You know, like it’s this universal truth and it’s just a matter of, you know, digging into it enough or in the right way to find those nuggets of information.
Linda McIver (10:36)
Yeah.
Yeah.
Nicole White (10:38)
you know, I, I think as I’ve sort of gone on in my career, I I kind of regret having that mindset in a way. because, you know, as I mentioned, you know, the context in which the data collected are everything.
Linda McIver (10:51)
Mm.
Nicole White (10:51)
but also more than that, it’s what the data doesn’t tell you, which I think is something that people need to think about more too. So, you know, who’s left out of that data set? Now, with statistics, we’re all about describing, you know, the majority.
of people,
Linda McIver (11:07)
Mm.
Nicole White (11:08)
you know, whether it’s through a mean or, you know, an average effect of a particular treatment. You know, people who sort of fall, you know, in the extremes of that distribution are quite often sort of not really considered. You know, they’re sort of the the outliers, if you like. you know, when you’re running analysis, you know, it’s very easy to assume, outliers, they’re just a nuisance. They’re skewing my results, you know. It’s just easier to
Linda McIver (11:32)
Yes.
Nicole White (11:33)
justify justify them away and leave them out.
Linda McIver (11:35)
Yeah.
Mm-hmm.
Nicole White (11:37)
But why are they in those extremes? I think that’s I think that’s where an area of data literacy I guess that needs to be challenged in a way. you know, what what is the
Linda McIver (11:48)
A hundred percent.
Nicole White (11:49)
data not saying about those people? I think is really
Linda McIver (11:52)
Mm-hmm.
Nicole White (11:52)
important.
Linda McIver (11:53)
I love it when you work with students and they’re like, W how do I know when I can call something an outlier and discard it? And I’m like, Well, there’s actually no easy answer to that. There’s no definitive answer. There are techniques, but they all come with assumptions and then some of them sometimes they’re wrong for a particular
Nicole White (12:12)
Yeah.
Linda McIver (12:12)
data set or a particular outlier. So you have to really think about it. They hate that.
Nicole White (12:19)
everybody does. There’s sort of this adage
Linda McIver (12:21)
Yeah.
Nicole White (12:21)
of statisticians is that we always sit on the fence, right? We always say,
Linda McIver (12:24)
Yeah.
Nicole White (12:25)
well, you know, it depends. We net like I get that that can be frustrating for someone who’s come to us for advice because we sort of are experts in our field and we say, and we don’t have that yes no answer for them. But it really does depend.
Linda McIver (12:37)
Mm. Mm. Yep. I and I love that and there should be more of that. Like I think, you know, it depends. We we should be quicker to it depends and and slower to this is the black and white answer to everything.
Nicole White (12:51)
Yeah.
Yeah. And it’s more of a why should we keep them in our data?
Linda McIver (12:58)
Yeah.
Nicole White (12:59)
Or you know, what if what’s the trade off if we do leave these people or other data points out of the status? What everything’s a trade off, right? Or an assumption. So sort of those decisions I was talking about earlier.
Linda McIver (13:11)
Yep. And they’re very human decisions and we might kid ourselves that we can be objective and rigorous and all the rest of it, but we are still human and humans actually don’t really do objectivity. We pretend. Hm.
Nicole White (13:25)
No, no. We all exercise our own judgment based on, you know, our experiences leading up like through our career or leading up to any particular decision. it’s like when you consult you know, ten statisticians. how do I answer this problem? You’ll probably get ten different answers. and that’s a combination of our different training and our expertise and no the mistakes that we’ve made as well along the way.
Linda McIver (13:53)
Yes, a hundred percent. So I wanna come back to that thing we talked about very early on that you touched on, the idea of flawed or mistakes in clinical trials, the stuff that goes wrong.
Nicole White (14:07)
Mm.
Linda McIver (14:09)
tell me about that.
Nicole White (14:12)
Yeah, sure. So I think sort of when I talk about, you know, all the things that can go wrong in clinical trials, there’s sort of very somewhat minor, most of the time inadvertent decisions or misjudgments, I guess, that happen in the very early stages of the trial, you know, even before the first patient is recruited or the first data point is collected. so as a statistician, I, I’m sort of involved in terms of the trial process ideally from the very beginning and even the conceptualization stage. because a key part of a clinical trial, I guess, is, you know, you want to be able to recruit enough patients or participants in general to be confident that if something is effective, say it’s a new treatment or a change in hospital policy, is a genuine change and not something that’s purely due to chance or, you know, we put invest all this time and money into collecting these data points only to say well, we actually don’t know because we haven’t collected enough or there’s too much uncertainty around this effect or some yeah, some little decision that’s happened during that planning phase. So in my role where I typically see things go wrong. and it can sort of be, you know, when when someone like myself is involved in that very beginning, that sort of a conversation that we have and sort of that meaningful collaboration where we can sort of nut out these issues to make sure that we have a smooth run further down the line. It just in terms of where it does get a bit messy is so when we’re brought in at the very end. So all the data’s collected, all those decisions have already been made, and it’s now, let’s find out if this worked or not. so those sorts of mistakes or errors in judgment at the very beginning, they can’t be fixed.
Linda McIver (16:01)
Yeah. Yeah, you can’t take broken data and get a good result from it.
Nicole White (16:09)
No, you can’t yeah, you can’t travel back in time and change the the endpoint that you’re interested in. Yeah, so it goes back to your example of discharge before. What do we mean by discharge in this trial? If we get that wrong,
Linda McIver (16:22)
Mm. Mm.
Nicole White (16:24)
then we’ve lost that opportunity to revise it to make it more suitable for the question that we’re trying to answer.
Linda McIver (16:31)
Yeah, you can’t go back and re-collect that that data.
Nicole White (16:35)
Yeah. Yeah.
Linda McIver (16:37)
It doesn’t work that way. so what are the worst data mistakes that you’ve seen? And whether it’s clinical trials or in the media or any examples that speak to you?
Nicole White (16:50)
Any examples of a data mistake?
I guess in terms of the actual collection of data, I’ve been fortunate enough to sort of not see those sorts of mistakes. I think where the the mistakes that I’m most familiar with is how then that how is that data analysed and reported to people? So, you know, something even just taking sort of media articles, for example, where you know, there might be you will see it all the time, reports of AI or predictive analytics can predict our risk of dementia twenty years down the track.
You know, it’s got 90% accuracy.
Linda McIver (17:25)
Yeah.
Nicole White (17:28)
but 90% accuracy compared to what?
Linda McIver (17:31)
Yep.
Nicole White (17:32)
So, you know, once you sort of start digging into the peer review publications that the media are citing, you find, well, if we just base that prediction on somebody’s age and their general medical history today, we could probably predict it with 88% accuracy. So
Linda McIver (17:49)
Yeah, so we haven’t actually gained much.
Nicole White (17:51)
Yeah, so this big shot it’s big and shiny and it’s new and it sounds exciting, but what does it what value does it actually add up upon what we’re already doing?
Linda McIver (18:01)
Yeah.
Nicole White (18:02)
and sort of the I guess the parallel of that is now what’s the denominator? I think people quite often go, you know, ninety percent of people have this particular outcome. Isn’t that terrible? I’m like,
Linda McIver (18:14)
Mm.
Nicole White (18:14)
but out of how many people? How many
Linda McIver (18:16)
Mm.
Nicole White (18:17)
people are we talking? Ten. hundred, a thousand, what’s the absolute impact of this? so it’s very, very easy to just sort of report like those headline stats that, you know, are easily understandable by everybody.
Linda McIver (18:30)
Yes.
Nicole White (18:32)
but yeah, what’s the denominator?
Linda McIver (18:35)
Yeah. Yeah. And also, is it in mice or people? I saw one recently. I can’t
Nicole White (18:39)
Yeah, that’s true. I hadn’t thought of that actually. Context, right?
Linda McIver (18:46)
remember what it was, but it was something that that something that that strikes close to home to me or someone I love who I can’t remember w even what the condition was, but it was like,
Nicole White (18:55)
Mm-hmm.
Linda McIver (18:56)
you know, in in some crazy high percentage of cases this solves the problem in mice. So
Nicole White (19:04)
no.
Linda McIver (19:05)
I was like
Nicole White (19:08)
yeah.
Linda McIver (19:10)
We’re really not anywhere near a you know functional treatment.
Nicole White (19:12)
No. I mean yeah, promising, but yeah, a long road ahead. Yeah. And
Linda McIver (19:17)
Mm. Mm. Yeah.
Nicole White (19:19)
that yeah, that comes back to transparency, right? Reporting those details
Linda McIver (19:22)
Yes.
Nicole White (19:23)
or yeah, taking the position
Linda McIver (19:25)
Yeah, and
Nicole White (19:26)
of the reader or the interpreter and what they know and what they don’t know about all the research that you’ve done.
Linda McIver (19:32)
Yes, and and and how the media constructs the headline as well. Like is it is it is it for the purposes of clear explanation or is it for the purposes of clickbait? And usually it’s for the purposes of clickbait and those two,
Nicole White (19:45)
Yes.
Linda McIver (19:45)
you know, come into conflict a lot.
Nicole White (19:47)
Yes,
absolutely agree.
Linda McIver (19:49)
Very frustrating. have you ever seen data deliberately misused and how do we spot it? What do we look for?
Nicole White (19:59)
Yes, so I think I might give a shout out to my PhD student here if that’s okay. so an area
Linda McIver (20:05)
Please
Nicole White (20:06)
so an area that my student Alex has been exploring is around sort of the idea of open access data and fabricated data and how that’s used or justified in people’s analyses. so he conducted a study last year that looked at data sets on the competition website called Kaggle.
Linda McIver (20:28)
yes.
Nicole White (20:29)
so they I think in the very beginnings they hosted sort of data challenges where you would download the data set and people would sort of try and answer the same problem in different teams. but
Linda McIver (20:38)
Mm-hmm.
Nicole White (20:38)
it’s now become, I guess, a more broader data repository where people can deposit data sets with the with good intentions of making data available for teaching purposes and so on. but what Alex found, he found two particular data sets sort of in the health and medical field that showed signs that they were fabricated.
So made up, basically.
Linda McIver (21:00)
Oof.
Nicole White (21:01)
and he traced that through to quite a number of publications who had since used that data to build predictive
Linda McIver (21:08)
no.
Nicole White (21:09)
models. so one was on the risk of stroke in hospitalized patients. and these papers were reporting this data set as legitimate. So
Linda McIver (21:18)
Mm-hmm.
Nicole White (21:19)
he was finding lots of different sort of statements around where the data came from, some papers didn’t even state where the data came from at all, but we were able to trace it back to this competition source. And some of these people had taken these predictive models and actually made them available as online tools. So a person could come along and enter their own information and it would spit out a risk of quite a serious outcome.
Linda McIver (21:45)
no.
Nicole White (21:47)
So it’s not so much data being misused in the traditional sense, but we’re now in this age where.
You know, we want to promote sharing and open data and open science. But there’s now this dark side to this in that when data are open, is it real data? And how are people disclosing, you know, the sovereignty of that data? it’s yeah, and sort of the downstream consequences of that in terms of risk to patients
Linda McIver (22:12)
That’s horrific.
Nicole White (22:14)
and general patient safety.
Linda McIver (22:16)
Yeah, that’s really disturbing. this wouldn’t be Alex Gibson by any chance. Excellent, because he’s coming up on the podcast very soon, thanks to our friend Karen who introduced us. That’s gonna be a great conversation. You have we’re we’re ready.
Nicole White (22:21)
It is Alex Gibson, the very same. Excellent. I’ve I’ve given him an excellent plug. Not stolen his thunder, thankfully. Excellent.
Linda McIver (22:38)
That’s fantastic. so what do you look
Nicole White (22:40)
Yeah.
Linda McIver (22:40)
for? How do you how do you spot things like that? I mean, most of us don’t have a PhD student handy to do all the the what I imagine was laborious and quite tedious legwork. What do what do you look for? What are the what are the tells?
Nicole White (22:54)
What are the tells? something that my first year stats lecturer taught me, which stays to me today, so before you do anything else, look at the data. Like actually graph it. Like that’s the first step that we should do. and
Linda McIver (23:10)
Yeah.
Nicole White (23:11)
I think we’re in sort of a like to just talking about sort of the research environment in general. We’re sort of in this environment now where you know we’ve always heard of publish or perish, but we’re in like this, let’s produce research and outcomes, as much as possible because data
Linda McIver (23:24)
Mm. Mm.
Nicole White (23:26)
is becoming cheaper, analysis is becoming quicker, particularly, you know, with sort of generative AI tools and so on.
Linda McIver (23:33)
Mm.
Nicole White (23:34)
We’re so quick to just generate outputs and then move on to the next thing. We don’t, as far as I see, we don’t do that fundamental step anymore. Let’s look at the data first.
Linda McIver (23:45)
Yeah.
Nicole White (23:46)
Let’s just plot it. Like going back to the outlier conversation that we’re
Linda McIver (23:49)
Yeah.
Nicole White (23:50)
having earlier.
If we simply plot our data, we go, look, there’s something, there’s a few outliers off to the side here. Well, what does that mean? Or even just looking at the distribution of the data. So with Alex’s study, one of the signs that we could tell that it was fabricated is I think he was looking at body mass index and how that was distributed in this data set. You sort of see this nice bell curve and then this massive spike at a particular value, which just seemed completely implausible.
But without plotting that data, you’d be none the wiser having this assumption, this here’s a data set, it’s ready to go. Let’s jump in and go for it and produce a result. so you know, in terms of how do we know you know, sort of the the quality or the sovereignty or the legitimacy of a particular data set and what to look for, it’s
Linda McIver (24:43)
Mm.
Nicole White (24:44)
for me it’s very visual. Look at the data. Just do something really simple. Like I know it’s boring.
Linda McIver (24:50)
Yep.
Nicole White (24:50)
Just doing a histogram or a box plot,
a basic five number summary, but those simple tools are there for a reason, because
Linda McIver (24:58)
Yeah.
Nicole White (24:59)
w you can’t analyse your data until you understand understand it at a at that fundamental level.
Linda McIver (25:06)
Yeah, and I feel like that’s also something we don’t teach enough when we teach statistics and data science. We don’t teach understanding the data. We teach cookie cutter here is a process, here is a formula, plug the data in, pull out the result, move on. That’s like
Nicole White (25:20)
Yeah. Yeah.
Linda McIver (25:23)
that’s not it’s not how it should work.
Nicole White (25:26)
Yeah,
it’s very methods focused, very testing focused.
Linda McIver (25:30)
Mm.
Nicole White (25:31)
you know, sort of one of the things that sort of gets me on my soapbox again is sort of this idea of statistics and data by flow chart. Is your data measured continuously? If yes,
Linda McIver (25:40)
Yeah. Yeah.
Nicole White (25:44)
are you interested in comparing two groups or three groups? Two, and
Linda McIver (25:47)
Yeah, yeah.
Nicole White (25:48)
so you follow this path down.
Linda McIver (25:50)
Yeah.
Nicole White (25:51)
it completely ignores the what’s the question?
What are sort of the nuances in the data that you can only see by actually opening it up and looking at it and visualizing it?
Linda McIver (26:01)
Yeah. Yeah.
Nicole White (26:04)
But yeah, it’s sort of the predominant w way, particularly when I was sort of teaching as an early career researcher, that statistics is taught to people who don’t go on to pursue sort of maths, data science, stats careers.
Linda McIver (26:17)
Hm. I I see it even more if anything in in data science courses where the first thing they teach is the tech. And and I always start with data literacy. I’m like y w we don’t need a computer for this. We we actually need to teach the thinking before
Nicole White (26:34)
Yeah.
Linda McIver (26:34)
we teach the tech. You know, you can’t go into using these techniques and using these systems or languages or platforms or whatever
Nicole White (26:43)
Yeah.
Linda McIver (26:43)
it is that you’re gonna use until you understand the fundamentals of data literacy. And if you’re not starting with data literacy, then you’ve you’ve you’ve messed up from the get go.
Nicole White (26:53)
Yeah, absolutely. And I think sort of something going back to sort of my training and how I got here, something that I really valued, which I’ve since realized, you know, isn’t universal in everyone’s education, is, you know, you need to understand what’s going on under the hood. ‘Cause it’s so easy to just click a button and produce a result these days.
Linda McIver (27:12)
Mm. Mm. Yeah.
Nicole White (27:16)
but what’s going on behind that button? Because how can you understand the result and what it means and you know what’s
What’s assumed and not assumed in generating that, how can you trust it?
Linda McIver (27:28)
Yep. Yep. When you have systems that do things like produce auto-scaled axes or you know, automatically
Nicole White (27:39)
Yeah.
Linda McIver (27:40)
throw out outliers outside a particular number of standard deviations or, you know, whatever and and can fundamentally change your result or your understanding of the result and if you don’t know to look for it, then you won’t know what’s just happened.
Nicole White (27:55)
Yeah. Something that I see quite often in my work is, you know, this automatic throwing out of people who have missing information, where the data on a a really key variable, so let’s go heart rate, for instance, is
Linda McIver (28:08)
Mm, mm.
Nicole White (28:09)
missing. you know, particular I did a lot of work during COVID around ICU outcomes for severe COVID patients. And a lot of the lot of the research that you see,
excludes patients who don’t have information collected on particular outcomes. But you know, if you think back to that environment, like even the healthcare environment and what was going on at that time,
Linda McIver (28:32)
Mm.
Nicole White (28:34)
what is it about why is that information missing from those patients?
Linda McIver (28:38)
Hmm.
Nicole White (28:38)
Because, you know, was it a resource situation? So, you know, countries with high income countries with a lot of resources could contribute a lot of data. But if you’re looking at data from a lower middle income country that was overwhelmed at the time, they wouldn’t necessarily have the resources to complete all that information, say in a data collection form. Or was it just so much going on with the patient that was so severely ill that
Linda McIver (29:04)
Yeah.
Nicole White (29:04)
there wasn’t time to collect that blood measurement or order that test? But to leave
Linda McIver (29:08)
Yeah.
Nicole White (29:09)
them out of the analysis entirely, and the results that come out of that, the skew and the bias that comes from that could be enormous.
Linda McIver (29:20)
Yeah. I mean especially if you wind up leaving out the the sickest patients all the time, then you get a completely different result. That’s it’s obvious when you
Nicole White (29:32)
Yeah.
Linda McIver (29:33)
can sit back and look at it, but when you’re in the thick of it, maybe it’s not not so straightforward.
Nicole White (29:41)
Yeah, there was a particular publication in the Lancet very early on in the COVID pandemic where they were looking at factors for mortality in the ICU. So at the time where there was very little information.
Linda McIver (29:54)
Mm.
Nicole White (29:55)
this these types of papers are really important.
Linda McIver (29:58)
Hm.
Nicole White (29:58)
there was one particular particular paper, I think it’s now been cited thousands of times. they left completely excluded patients who were still in hospital at the time of the analysis.
So that either unfortunately died or had recovered and been discharged from hospital. But if you’re still in hospital, you’re
Linda McIver (30:14)
Yeah. Yeah.
Nicole White (30:15)
still in hospital for a reason. And so
Linda McIver (30:18)
Yeah.
Nicole White (30:18)
to exclude those patients entirely, you’re actually your interpretation is very different.
Linda McIver (30:24)
Yeah, you’re really skewing the results.
Nicole White (30:26)
Mm.
Linda McIver (30:27)
There was there were a lot of issues too around the definition of someone having COVID initially.
Nicole White (30:36)
Mm.
Linda McIver (30:37)
It’s like, okay, so you’ve had a positive test. When was your positive test? When did you start showing symptoms and at what point do we arbitrarily declare you recovered? and and everyone had different
Definitions, which meant that everyone’s data was different. Like, how many people currently have COVID? Well, define have COVID, you
Nicole White (31:00)
Yeah.
Linda McIver (31:01)
know.
Nicole White (31:02)
Yeah. And then even at a higher level, you know, comparing all the evidence across all these studies, it becomes futile because everyone has their own definition that they’re working with. So you’re not
Linda McIver (31:13)
Hm.
Nicole White (31:14)
comparing like with like anymore.
Linda McIver (31:16)
Mm. Yeah. Yeah, and that comes back to that idea of definitions and the you know,
Nicole White (31:22)
Mm.
Linda McIver (31:22)
the the human thumbprint on the data, if you like. The
Nicole White (31:24)
Yes.
Linda McIver (31:25)
you know that it’s it’s it’s not you know, you think how many people have COVID? Well that’s easy. That’s you know, that’s straightforward, that’s a black and white number. Well, no, it’s not.
Nicole White (31:36)
Yeah, no.
Linda McIver (31:37)
It’s it’s complicated.
Nicole White (31:39)
Yeah. Hundred percent.
Linda McIver (31:42)
I guess that’s the that’s my kind of what everyone what I wish everyone knew about data was just that it’s complicated. It’s
Nicole White (31:50)
Yeah.
Linda McIver (31:50)
always complicated and it’s more complicated than you think. I think if
Nicole White (31:53)
Yeah. Yeah.
Linda McIver (31:56)
you start from there you you got a better better chance of a good outcome.
Nicole White (32:00)
Yep. Subjectivity rules the world,
Linda McIver (32:03)
Yeah.
Nicole White (32:03)
I think, even data.
Linda McIver (32:05)
Yeah, a hundred percent. I I love that that message because we have this tendency. I think Luke Stark, who’s a Canadian researcher, calls it the the charisma of numbers. that you know, as soon as you put a number on something, as soon as you’ve got data, it looks you know real and straightforward and you can’t argue with it and it you know, you believe it and it’s similar
Nicole White (32:35)
Yeah.
Linda McIver (32:36)
to what we talk about a lot now with with AI, with the automation bias, where even when you know the system is flawed and can give you the wrong answer, you you trust it because
Nicole White (32:46)
Yeah.
Linda McIver (32:46)
it came out of a computer and and it’s the same with data. Like you trust it because it’s a data set. It’s like, there could be so many things wrong with that data set.
Nicole White (32:55)
Yeah.
And I think that problem compounds too with sort of this big data phenomena as well. I call it phenomena, but it’s been around for you know more than a decade now. But this idea of if only we had more data, we could get a better answer. But the trade-off is the more data you collect, the more nuances and assumptions and, you know, potential errors creep in.
Linda McIver (33:19)
Yeah.
Nicole White (33:20)
how much more value is having this huge data set compared to something that
Like a small clinical trial, for instance, where an individual person is collecting every single data point. It may may not be a huge data set, but it’s designed
Linda McIver (33:32)
Mm. Mm.
Nicole White (33:34)
well. it’s well thought through in terms of, you know, the outcomes and how it will be analyzed and how the results will be interpreted. but yeah, we t we
Linda McIver (33:45)
And there.
Nicole White (33:46)
tend to chase this big data idea. and don’t get
Linda McIver (33:49)
Yeah.
Nicole White (33:49)
me wrong, it definitely has value, but it’s not the universal solution.
To every single question.
Linda McIver (33:57)
Yeah. And doesn’t necessarily have the value we think it has.
Nicole White (34:00)
Yeah.
Linda McIver (34:01)
And and, you know, it that that small clinical trial you describe still has the definitions and decisions, but they’re all being made by the same person, then at least you’re likely to
Nicole White (34:12)
Correct.
Linda McIver (34:12)
get some consistency. Whereas, you know, w my field of education, just looking at inter-rater of reliability, you know, do two people mark the same exam the same way? Almost never.
Nicole White (34:22)
Yeah.
Linda McIver (34:23)
And that and and you know, exams are pretty straightforward compared to
people, you know, clinical trials and
Nicole White (34:32)
Yeah.
Linda McIver (34:32)
that those kind of measurements.
Nicole White (34:34)
Yeah. I mean exams have black and white answers. Well in my training they did. With the formulas and so on and so forth.
Linda McIver (34:38)
Yeah.
Yeah. Which is one of the issues with exams, of course, but that’s a Yeah, that’s a that’s a separate podcast.
Nicole White (34:44)
Yes, yeah, don’t get me started on that. Yeah.
Linda McIver (34:51)
what’s the first question you ask when you look at graphs in the media?
Nicole White (34:59)
Graphs in the media. Mm.
That’s a good question. I think it goes back to what’s the denominator in the graph.
Linda McIver (35:09)
Mm-hmm. Mm-hmm.
Nicole White (35:10)
so often we might see a bar chart, for instance, that just has percentages on the y-axis. You know, 80%
Linda McIver (35:16)
Yeah. Yep.
Nicole White (35:18)
of people bought this item, 20% bought this item. You
Linda McIver (35:20)
Mm-hmm. Mm-hmm.
Nicole White (35:22)
know, a difference of, you know, 10% between sort of different bars. What does that mean in terms of total people or total people impacted?
That’s probably something yeah, I think about when I see graphs in the media.
Linda McIver (35:39)
Did we look at
ten people or ten thousand?
Nicole White (35:42)
Yeah. Or, you know, even graphs that sort of might indicate, thanks to this new technology, you know, ninety percent of people with c this condition will now live longer.
Linda McIver (35:56)
Mm.
Nicole White (35:56)
again, it’s like, how rare or how common is this disease that you’re talking about?
Linda McIver (36:03)
Yeah.
Nicole White (36:03)
you know, is sort of this big headline percentage, whether it be ninety percent or ten percent or one percent, what does that actually mean in terms of
Individuals themselves.
Linda McIver (36:15)
Yep. Yep. and that brings me back to relative risk as well, which I know is one of Karen’s bugbears and I see the reaction
Nicole White (36:25)
Yes.
Linda McIver (36:27)
on your face. It’s like this will triple your chance of your your risk of getting the thing. It’s like, yes, but it’s taking it from, you know, one in a billion
Nicole White (36:39)
Yeah.
Linda McIver (36:40)
to three in a billion. I’m not really worried.
Nicole White (36:43)
Yeah.
Linda McIver (36:44)
But triple the risk.
Nicole White (36:44)
Or we want yeah. my goodness. Yes,
it’s like, yes, we can we can power this trial to have a twenty percent reduction in infection, say. It’s like, well,
Linda McIver (36:56)
Mm-hmm.
Nicole White (36:56)
are you starting from a rate of five percent or twenty percent? Because that that
Linda McIver (37:00)
Mm.
Nicole White (37:02)
twenty percent fall could be, you know, one in a hundred people. Or
Linda McIver (37:06)
Mm.
Nicole White (37:06)
it could be, you know, ten, twenty in a hundred people. Yeah, in terms of the return on investment.
I know which one I would go for.
Linda McIver (37:15)
Yep. Yep. But it’s all in the presentation. It’s you know, data is a can
Nicole White (37:23)
Yeah.
Linda McIver (37:23)
data results can be a form of marketing if we’re
Nicole White (37:28)
Yeah, for sure.
Linda McIver (37:29)
if we’re not careful.
This has been a great conversation, and it’s been very gratifying to have so many of my personal bugbears brought out and shaken in the light. Yes, yes
Nicole White (37:45)
It’s always nice to chat with like minded people.
Linda McIver (37:49)
it is. and it brings me to the last question, which is always my favourite. What excites you about data?
Nicole White (37:58)
What excites me about data so much. we talked about messy data a bit at the beginning and sort of how I wasn’t introduced to that early on in my career. I think now
Linda McIver (38:09)
Yeah.
Nicole White (38:09)
that’s what excites me the most is working
Linda McIver (38:12)
Yeah.
Nicole White (38:13)
out the mess. See, you know, I always think about, you know, real life data. It’s kind of like this, you know, a knotted up ball of wool or string.
Linda McIver (38:20)
Yeah. Yeah.
Nicole White (38:22)
actually, you know, it’s tedious and it takes time and it takes a lot of critical thinking, but just gently untangling that mess to get something that’s interpretable. that’s what excites me about data, I think, is sort of that process of making sense of it.
Linda McIver (38:40)
That’s that’s beautiful. The puzzle of it, the challenge of it is the yeah. And you don’t get that. You don’t get that from the toy data sets. Yeah.
Nicole White (38:41)
Hmm. Yeah. Yeah. It’s not everyone’s bag. I get that, but it’s satisfying. Yeah.
Linda McIver (38:51)
Yeah. It really is. Thank you so much. This has been a wonderful conversation. I really appreciate you sharing your time and expertise.
Nicole White (38:59)
Yeah, absolutely. It was lovely to chat. Thanks so much, Linda.
Linda McIver (39:02)
Thanks, Nicole.
Outro (39:07)
Thanks for listening to Make Me Data Literate. You can find more episodes at ADSEI.org/podcast and you can support our work at givenow.com.au/ADSEI. Have a great day.
