Associate Professor Karen Lamb on Rigorous Research Design

This episode of Make Me Data Literate is a super interesting chat with Associate Professor Karen Lamb about biostatistics, rigorous experiment design, the human side of statistics and more.

“Because to me the whole thing is statistics. If it doesn’t work from the start, if the design isn’t right, then there’s no point in me having a nice, beautiful sample size calculation. So it’s really about the rigor robustness of the science from the outset and then continuing with that”

“I really, really hate pie charts. Don’t get me started on how much I hate pie charts.”

Linda (00:00)
Welcome back to another episode of Make Me Data Literate. This is a very exciting one for me and a very special one because my longtime friend Associate Professor Karen Lamb, who is also director on the board of the Australian Data Science Education Institute, is here with us today. Welcome Karen.

Karen Lamb (00:19)
Thanks for having me, Linda. Sorry it’s taken so long.

Linda (00:24)
I finally pinned you down, we’re good.

Karen Lamb (she/her) (00:25)
Yeah.

Linda (00:27)
So can you tell me who are you and what do you do?

Karen Lamb (00:31)
So my name is Karen Lamb, as you said. I am a biostatistician, and I co-head the biostats node for the Methods and Implementation Support for Clinical and Health Research at the University of Melbourne. so I’ll call it MISCH for short, but MISCH essentially is an idea of a kind of one-stop shop for your methodological support for high quality health research. So we’ve got people in co design, we have people in health informatics, data informatics, myself obviously biostatistics, we help with design and conduct of of health studies, trials, observational studies. We also have people from the end of looking at the health economics perspective. So seeing if if you found something to be efficacious, is it actually going to be possible to move forward with it because of the the costs involved. So yeah, I’ve been here for about seven years now and we provide support to numerous health researchers and clinicians across the University of Melbourne and the affiliated hospital partners.

Linda (01:42)
So that’s what MISCH does. What is a biostatistician?

Karen Lamb (01:47)
Yeah. Biostatistician is someone that is from a statistical background, that’s a really great start, you know, using the name of the profession to help. But essentially a role of a statistician is someone that will help, I think, ensure robust study design and data capture and analysis. And particularly as a biostatistician, it’s my role to help particularly with health research, clinical research. so lots of work with human populations.

I originally came from a background of studying more what we call a cohort study, so looking at observational research erm to see different things that could cause different health outcomes, diseases, and so on. So you take data from a sample of the population, follow them over time to see what could potentially cause different health outcomes. But more recently, I guess the primary piece of work that I do as a biostatistician here is working in randomized control trials where we will help health researchers, clinicians, when they’ve got a particular intervention in mind that they want to study, and we help them to design something to try and estimate whether that intervention is worthwhile using and actually is as I say has a has a desired health effect.

Linda (03:15)
So an intervention might be something like a drug or a lifestyle change, things like that.

Karen Lamb (03:20)
Yeah, so typically for us it’s less often drug interventions. That’s typically more, I guess, the pharmaceutical sector than than what I do. a lot of the work I do is more behavioral interventions or exercise interventions. I guess I w I work a lot with physiotherapy group in particular at the University of Melbourne that do a lot of work in osteoarthritis. And so a lot of their interventions they’re interested in are exercise interventions. And so they or or even educational interventions to sort of people often think a good response to some of these these challenges is to go and have surgery. And it’s to try and actually say there are alternatives there and to try and promote things like exercise that we know could have multiple benefits. So so it’s more around that than the than drugs as such, although sometimes that is a component of the kind of work we do.

Linda (04:20)
I love that. So basically your work is to ensure that the studies that happen are as rigorous as possible, that the evidence is as strong as possible, that the outcomes are as clear as they can be and not muddied by problems with the design of the study and and things like that.

Karen Lamb (04:41)
Absolutely, because I think I used to say to people that would think of statistics as the science of disappointment because early in my career, and actually still now, but you often would get researchers coming to you at the end. So people think of statisticians as being the people that you go to when you’ve got your data, and we work our magic with it and analyze it and get you the result that you were looking for. but there’s so much more to the job, which is why I love the job, because that’s such a tiny piece of actually what the role of a statistician is.

My job largely involves talking to people, and particularly where I am now sort of heading a group. I do still occasionally do some of the analysis, but largely I talk to researchers to try and help them frame their questions in a manner that can be answered. And then taking that well-framed, refined research question, thinking about what’s the appropriate design to enable them to answer that question? Then taking that design into a well-developed protocol, where again people would have the perception that statisticians may only do your sample size calculation, how many people do you need for this study, and then write your couple of sentences the stats method. But actually, the role is ensuring consistency and rigor throughout the whole proposal. So I would never get a protocol and just look at, just do the stats sections.

Because to me the whole thing is statistics. If it doesn’t work from the start, if the design isn’t right, then there’s no point in me having a nice, beautiful sample size calculation. So it’s really about the rigor robustness of the science from the outset and then continuing with that. I then don’t leave them alone. I want to look at your questionnaire. How are you capturing the data you need? Is the outcome, is the tool you’re using to capture your measures, whether they’re outcomes, exposures, patient characteristics, are they validated, reliable measures? Have we thought out potential problems with the responses you have there? Are we capturing all the variables that are of key interest to you? We’re not just focusing on outcomes here, but there are other things in the process of your trial that might be important or your observational study that you need to make sure you’ve got. Then looking at if you’ve designed the survey, does the database align with your survey? So some people might just design the survey straight away in an electronic database that means it’s all in one step, but sometimes people like to work in Word or whatever, create their survey, and then I check to see is the database robust to actually get that information? And can we make have some efficiencies there? Now I never design a database, that’s not my role. That is more your data managers health informatics people.

My role is just to test it to make sure that the data that we get out is actually going to be useful. I can’t tell you the amount of times I’ve helped with a survey but then looked at the database and what I’d hoped were these nice clear drop-down menu responses or whatever. I’m not saying I’m I’m all for qualitative data as well, by the way. I’m gung-ho for mixed methods and having qualitative information. But largely what we’re needing in these studies, if you’ve got me, is quantitative data.

And so if you have all these people writing potential free text fields that I then would have to try and code up, then or well I wouldn’t, I would go back to your data manager, but essentially that’s why we need the front end all sorted before we even get to the point of analysis reporting all the rest of it.

Linda (08:28)
So you must have the same occupational hazard that I do, which is seeing studies reported in the news and just tearing your hair out at the both the reporting and

Karen Lamb (08:39)
Absolutely.

Linda (08:42)
often the study design and just going, no, it doesn’t show that. What are you doing?

Karen Lamb (08:48)
Yeah, I feel like, you know, you do this qualification or, you work in this field and you’re just an eternal skeptic. So it is a lot of as you probably see many things, like show me the data, but beyond that, show me where the data came from. Tell me how you got to this. Tell me, you know, how you cleaned your data and or or what subset of the data you took to try and get this message. there’s so many different ways to to get the answer out. And so what I work very hard to do with anyone working with is have a clear plan for what you want to do. And even if you come to me when your data is collected, I still want to work with you to have a clearly defined statistical analysis plan that outlines what we’re trying to do, what the limitations are of the data that we’ve got, what what all the steps are we’re going to take in place for the cleaning, just so that everything is transparent because, you know, there’s no correct model per se. Everything is based on assumptions, but let’s be transparent about the assumptions we’re making and justify them as much as we can and be mindful that our assumptions someone might disagree with, but at least they’ll know what assumptions we’ve made and why we made them.

Linda (10:03)
Think I can stop teaching workshops and just play an hour of Karen saying that.

Karen Lamb (10:08)
Me on my soapbox.

Linda (10:11)
It’s perfect. I love it. what did you have to learn to do your work? Was there anything missing from your formal education?

Karen Lamb (10:22)
Chat.

Really, I’ve talked to you about this before, Linda, I’m sure. But it’s really interesting recently reflecting on it because I think had you asked me this, and you probably did years ago when we first met, the thing that I really wished I’d done more of or felt was a real weakness in my background was the computing elements. so you know, I’m I’m statistics trained. Like n rarely people wake up and go, I want to be a statistician when I’m older, but and I certainly didn’t do that.

Linda (10:51)
Ha ha ha.

Karen Lamb (10:52)
I’m fortunate that my dad’s a mathematician, happens to be a pure mathematician, so gung ho for the beauty of maths, less excited, although actually no, he’s excited about everything, but more me, I’m on the-

Linda (11:06)
Ha ha ha.

Karen Lamb (11:06)
-other end of the spectrum where I’m like, the utility of it. How can it help people? So it was all

Linda (11:10)
Yeah. What is this good for?

Karen Lamb (11:13)
Well, it was always the maths element that brought me into this, not the computing. And I think it was a real shock to me. So I started doing mathematics at university and took statistics because it was like maths.

That was it. That was basically it. Like this sounds like math, let’s do it. but I really took to statistics because when I was studying at the University of Glasgow, they’d separate statistics and maths departments. And the statistics department, not only were they much more on the first name term basis, John, this, Tom, Jim, whatever, bearing in mind, yes, largely male names here. but I just really liked that vibe compared to going upstairs to the maths department and it was professor so-and-so.

So of course that was one element for me because I’m not really a huge fan of hierarchies. But I think the other thing was they really taught you with examples of what what they were doing. So they were doing cool methodological stuff, but why? Who are you working with and why are you doing it?

But as I say, it was a real shock to me going into that because computing was never something I was particularly interested in. You know, I’m of the generation where we did have a computer at home when I was a teenager, but I was very resistant even to writing essays or anything on it. Like I’m I am still very much a handwriting person. If I handwrite it, I know it. If I typed it somewhere, it may not exist anymore. So then I went math, again, still very chalk and talk go to statistics and suddenly start with we were using something called MiniTab. And it’s a it’s a reasonably-

Linda (12:54)
That takes me back.

Karen Lamb (12:56)
-user-friendly start-off stats program where you know it does have the drop-down menus and things. But I started learning with drop-down menus and then we suddenly were moving into more statistical computing and doing things in S plus, which was subsequently R. And I just really struggled. I mean I remember in my honours doing statistical computing and the sheer frustration that my professor had because he’s like, I was getting A’s in stats computing, but nothing I did ever worked. It’s the joy of like show me your working and I was like he’s like, Karen, you’re so close. And I’m like, and yet so far.

Linda (13:38)
Yeah.

Karen Lamb (13:38)
Because I hated that it had to be just perfect or it didn’t work and I just couldn’t be that perfect. And for the longest time I have carried that with me as like one of my things with stats that I really struggle with, except f forgetting that the fact that the personality I have, I like to jump into whatever language anyone’s using and see if I can work I I like looking at other people’s code, I like pulling it apart, I like understanding what’s going on, I just don’t like really doing it myself. So that was a long way of saying, had you asked me, even maybe a couple of years ago, I would have said this is what was missing, this is all what was missing from my background, even having studied statistics and then going on to a PhD in mathematics and statistics.

But more recently, I think, really has been the human element of being a statistician, which is just not emphasized at all. And yet, you know, the discipline that we’re in, it’s not it’s rare to have you just sitting by yourself in an office. I know there are some very, very methodological statisticians that are probably sitting doing formulaic stuff, do I don’t know, probably doing the blackboard, whiteboard stuff. this shows you the way I’m chatting how far removed I am from this.

But my job really is to bring, bring people along with me, bring the clinicians, the health researchers, get them invested in working with me and other statisticians, knowing that we’re not here to hinder their progress, we’re not here to cause them any problems or or to essentially give them answers they don’t want. We want them to work with us to make this the best it possibly can be. And that all comes from having good communication, a clear appreciation of maths anxiety. you know I have met in my career collaborating with people, applied folk come in professors in tears because the thought of coming to speak to a statistician was too much because of their feelings from high school. And I mean hell I don’t think I’m a particularly scary person but that name, that profession was what was carrying all the weight there.

So I just wished, and still wish there’s more emphasis on that side of things and the people management, probably the project management, but also helping me to get it right, helping me to, you know, I would go into meetings when I was brand new in this job and think I had to show off. Think I’d have to impress the collaborator with what I knew in statistics. when actually that’s just terrifying for anyone and doesn’t help anyone, when really the role.

As I say is bringing them along with me, learning what it is they need and making sure I’ve got that all captured well so that I can work out what I have to do next.

Linda (16:46)
My catch cry is tech is easy, people are hard. And it’s interesting, I get these themes that come up through no planning of my own. But the last episode of this podcast that I recorded, which I’m editing just now, the the my guest, Professor Crystal, she said that what was missing was the human element. She wasn’t taught enough about people.

Karen Lamb (17:09)
Yeah.

Linda (17:11)
She wasn’t specifically she said I wasn’t taught how to be a good human. And you know, the connection is there. It’s it’s that is part of you know, being a good human is not coming in and you know, smashing your expertise out on the table. Being a good human is coming in and listening and and connecting and reaching people where they are and you know, same with being a good teacher.

Karen Lamb (17:32)
I think also a reflection, there’s couple of things I’ve been thinking about recently and some of it is the the ability to realise that other people’s behaviour is not necessarily on you. I refer back to you know the science of disappointment and so on, but like ultimately, at heart I am a people pleaser. I can’t help it. I want I want people to have a nice time, but that is in stark contrast to what I’m often doing in my job where I’m like, I’m really sorry, you just can’t do that. And I think it’s taken a lot for me to is one is to to reflect and say, yes, you can’t do this, I know you really wanted to do this. Here’s what we can do, here’s what we can learn from this, and here’s how we can do it better next time, and take that to the people I’m working with. And the other thing for me is not to internalize it that if a collaborator is really disappointed or frustrated with how things are working, that it’s not my fault necessarily.

And I think that’s something I wish it’s that bit of resilience, that no one sets you up for. I mean, you start you do a maths degree, a stats degree, and it’s like, yeah, get the data, get your computer out, do the math.

Linda (18:49)
Ha ha ha

Karen Lamb (18:50)
And it’s like, that is so little of my job these days that I mean it really shocks me. I’m I’m very happy that that is less of my job as I progress because I really enjoy working with other people and learning what they want to do. But sometimes as a statistician, the frustration is taken out on you, which is hard.

Linda (19:14)
Yeah, that that’s brutal. And I have to take the opportunity now to brag on your behalf because I know that you won’t, but that you are for the second time now nominated for a Eureka prize for outstanding mentoring of researchers. and and absolutely deserved that your mentoring programmes and the structures you’ve developed at your work to to to teach people these skills that you feel were missing. I i it’s just glorious. I love it.

I will stop embarrassing you now.

Karen Lamb (19:45)
Ha ha ha

Linda (19:48)
Is there one thing that you wish everyone knew about data? Is there something that would make your life a lot easier or, you know, just make everything better if everybody understood this one thing?

Karen Lamb (19:59)
I always find it really interesting when I have to think about data like in its own right, without, as I say, all the precursor and the design. So I guess maybe that’s what it is. It’s the being mindful that the data that you’re looking at was collected for a particular purpose. And even if it was collected for a particular purpose, it doesn’t mean it was collected well for that purpose or represents the people it it truly should. So I think, you know, as a a statistician, and particularly in the the fun AI world in which we find ourselves.

I’m constantly thinking about the potential for bias and the potential of misrepresentation. And also, to be honest, even more now, is thinking about the importance. I mean, I’ve kind of always thought about this in the background, but your data provenance.

Like as I say, talking a little bit about, you know, why was your data collected, but really being able to trace that back carefully. you know, interestingly this year with some of the statistical groups that I’ve visited, there’ve been some people doing reviews on existing prediction models that are out there. And, you know, they maybe started off the review with one purpose, but actually in going through this review, they discovered that some of the big data sets that were being made available out there that people were were using to develop a prediction model, quite innocently, I don’t think there was any well, I don’t believe there was anything malicious there, but you couldn’t trace back the provenance of the data. So was this real data? Was this fictitious data? But if if data sets exist and you don’t know, you know, why then that’s a problem.

And then actually, this this PhD student that was working on this piece, who I probably should put you in touch with, Linda, I think you’d have a good chat with them. But actually in exploring the data a bit more themselves, we’re finding weird things in the distributions, things that wouldn’t make sense if these data truly were data for real people. so yeah, be worried if you don’t know much about the data, and then be worried if things appear good too good to be true because they likely are, and also just look at your data, describe your data. Does it make sense? Does it actually look like you expect it to? Are the values representative of what a real distribution would be? and if you can’t answer those questions, like don’t don’t be using the data.

Linda (22:39)
I like that. I like that a lot. And it’s it’s a theme that keeps coming up is that idea of, you know, where did your data come from? what and and the idea that data is a human artifact. We forget that. We think of it as some kind of, you know, mathematical truth. Some you know, this is this is the number of people who were in this room at this time, you know, and-

Karen Lamb (23:02)
Yeah.

Linda (23:02)
-that that’s a that’s a thing. It’s like, well, hang on even that is complicated. Are we counting people who stepped out and came back? Are we counting, you know, was it the number of people who were in there for the entire session? Was it what ha at what point? It’s like what what counts as streaming a song on a on a streaming platform? Is it the whole song? Is it all but the last ten seconds of the song? Is it thirty seconds? Is it a minute? You know, how do we count these things and every every bit of data has a definition behind it that is a human artifact.

Karen Lamb (23:38)
Exactly.

And it’s like are we asking the correct research questions or the correct questions throughout? because another thing you sometimes see if you’re you know, if you’re thinking in a clinical sense measuring something objectively like your blood pressure and or heart rate or something like that. And even even with these more objective measures, we’re aware there are some challenges like you know, I’m terrible with jargon, so I can never remember the exact term. But you when people say your blood pressure will spike when you’re actually visiting the doctor because of that doctor effect. so-

Linda (24:11)
Hmm. White coat hypertension.

Karen Lamb (she/her) (24:15)
-thank you. That’s what I was looking for. And so is that actually measuring and is that accurately measuring what you need when it is in theory objective? But I probably come at it more from I guess the subjective questionnaire measures we have or the tools. You know, I’ve worked lot with psychology researchers, behavioral scientists, people that are interested in things that are quite hard to define around well-being or mental mental health, quality of life, that kind of thing. And it’s interesting as well sometimes the perhaps lack of consideration, particularly I guess when people are earlier in their career in research and maybe just less aware of these things about the importance to look and see if there is an a tool or a measure that’s validated and reliable for the people you want to to study. And if you have created your own tool to measure something or your own question, should you have? Can’t you go and see if there’s something else out there that that would be a better measure? and actually, like to be honest, every question in a survey you have to kind of look at with that mindset, any question in a database, did that was that question asked in an appropriate way? And if not, well again should we be using maybe the other elements of the survey that are good?

But maybe that question is just not a good one to use.

Linda (25:43)
My favorite example of survey problematic survey questions is the the classic workshop evaluation question. which can be anything from how wonderful was this workshop today to how did you find this workshop today? You know, one to-

Karen Lamb (25:58)
-my word, yeah, leading questions.

Linda (26:02)
-five. Like a completely different outcome which requires an understanding of human psychology and behavioural psychology that most people who are writing surveys don’t have, which is why you need, you know, the involvement of someone who does this for a living and who has that expertise. But crafting a survey is simple, right? There’s what could go wrong?

Karen Lamb (26:28)
And it’s even that thing of like talking through people not making every question compulsory as well. So you know now if you’re using certain databases, you can just say force response for everything or you can’t progress. And you’d think that me as a statistician would be gung-ho for that. I don’t want missing data, but I am prepared to give up on some data to prioritize other parts. So there’s certain things, take income for example. you know I’ve worked a lot in social epidemiology and and often they’re trying to capture different aspects of what they would class as socio-economic status, whatever again this broad construct is.

Looking at employment, looking at education, and then there’s usually an element of reporting your income, but people don’t necessarily want to tell you their income. They might be quite prepared to tell you their educational background, but not their income. So do I want people to be forced to respond to a question about their income when it’s a desirable rather than an essential for this this particular survey?

Linda (27:44)
Yeah, and does that mean you’ll have people bail from the study who would have filled it out otherwise? And are you therefore getting a sample size of only people who are prepared to tell you your income? And is that a skewed sample of so many questions.

Karen Lamb (28:02)
And then the other fun thing with asking income is the option as well of if they do do a forced response, they might have a don’t know or prefer not to say. So that’s one way of capturing the data. But essentially for me, what do those categories represent? I still don’t know. You know, if you’re trying to capture like this ordinal measure if you like of income, what do I do with them? To me they’re still kind of missing. And then people-

Linda (28:24)
Yeah. Yep.

Karen Lamb (28:26)
-do all kind of weird and wonderful things with it. So yeah, it’s it’s tricky. Survey design, absolutely, I wouldn’t say I’m an expert at it, but I definitely know big no no’s so

Linda (28:39)
Yeah. Yeah, it’s an interesting one. There’s so so many ways it can go wrong. which leads us to mistakes. What are the worst data mistakes you’ve seen?

Karen Lamb (28:55)
Wow. Where do I start?

Linda (29:00)
So many.

Karen Lamb (29:02)
Look, I feel like I could start with the kind of common general ones without making it seem like I’m naming any names in particular, but I think, you know, the the common thing we often complain about as statisticians is the statistical significance over clinical or meaningful difference. And this comes up all the time and even though we’ve been discussing it for as long as well longer than I’ve been doing stats, but it’s this idea of your p value, your significance level of your test being the be all and end all of your research finding. And if you fail to conclude that something is significant, then that’s it. Your world is over or you can’t get this output out there, or you know, there’s no point in exploring this any further. and it’s just so incredibly frustrating that just the research environment we’ve been working in for so long has really pushed for this statistical testing, significance testing, and then that is the emphasis in reporting, rather than say.

The average difference in blood pressure between your two groups, which was the measure you were interested in. You’re still going to estimate that. Maybe you’ve found that it’s not significantly from the statistical perspective different between the groups. But what difference did you detect? What is the confidence interval that can tell you what that difference may be in the population? and while you know no difference is plausible.

It’s not the only potential difference in the population. And I think this just this misrepresentation of findings using significance is probably the one I encounter the most. It’s probably not my best example. I have some absolute howlers out there, but this is the one that I probably still have to talk the most about in my meetings with researchers.

Linda (31:07)
I guess it comes back to the point of it all, like why are we doing the stats? Why are we doing this study? What do we actually care about? And in fact-

Karen Lamb (31:16)
Mm-hmm.

Linda (31:16)
-p value is not the thing that comes to mind. It’s like you know, is the is this is this intervention worth it? p value doesn’t necessarily give you that.

Karen Lamb (31:29)
Yeah, I was actually I I reviewed an article recently and without saying anything about it at all, basically in my first review response as a statistical reviewer said, What’s all the statistical testing about? I actually think this is an interesting article, but really it should be descriptive. Just tell me what you saw. I don’t see a need for all of these p-values. And their response to me was that they were trying to make inference to the broader population and that’s like they were essentially trying to tell me what what statistics was. And I was like-

Linda (32:04)
What?

Karen Lamb (32:04)
-look, I don’t disagree that that’s why we will do like some of this statistical work. But all of your conclusions are based on p-values for the over a hundred and twenty different models that you have fitted. At no point have you told me what parameter you are wanting to look at in the population, in which-

Linda (32:25)
What?

Karen Lamb (she/her) (32:25)
-case you’ve decided that the p-value is that. So, Linda, I just I went-

Linda (32:29)
Boy.

Karen Lamb (32:30)
-I thought I was going crazy. So I was like, I’m sorry what? So I then had to push back again, going, again, I don’t think your piece of work is necessarily a bad piece of work. You’ve just focused on the wrong thing.

Linda (32:43)
Yeah.

Karen Lamb (32:43)
Like I think you’ve done this useful potentially useful piece of work, but it’s all been masked by this focus, unnecessary and incorrect focus on p-values.

Linda (32:57)
That’s disturbing. That’s really disturbing.

Karen Lamb (32:59)
Mm-hmm. Mm-hmm.

Linda (33:02)
Have you ever seen data deliberately misused and how do we spot it?

Karen Lamb (33:11)
I guess deliberately misused. I think one of the there’s some interesting articles out there about I think there was an American Stats Association survey of biostatisticians in the States around bad practice and bad requests from that early career in particular biostatisticians think were faced with. And they had some really nice examples around, you know, changing your research question after seeing the result of your original research question, which is why I’m-

Linda (33:41)
Ooh. Mm.

Karen Lamb (33:43)
-I’m all for full transparency. Because I think you know, I I’m originally from more of an observational study perspective. And it’s really tricky in an environment where there’s a lot more expectation around your presentation of trials. So if you’re conducting a randomized controlled trial, if you want this actually be considered as a good trial, it has to be registered. So in Australia, there’s Australia and New Zealand Clinical Trials Registry. So you have to register your trial, tell us who your eligible population is, tell us what your sample size you’re aiming for are, primary outcome, that kind of thing, and brief details of your analysis approach. You have to have a detailed protocol, of course, it will have gone through ethics, you likely will have published your protocol in a journal. And then when you’re getting to point of doing your analysis and reporting your findings, prior to that point, you should have developed a fully rigorous statistical analysis plan that ensures that your what you have in the trial registration is consistent with your protocol, is consistent with your plan. Any changes since the registration or protocol are clear, any major deviations are clear. And then you publish the trial result irrespective of finding, ideally, with the stats plan as a supplementary file that tells you the full detail of everything you planned and why.

Now, observational studies, on other hand, it’s a bit of a Wild West. we just don’t know w how many analyses are out there where things just don’t end up being presented. so we can push as statisticians, data scientists, for clearer rigor by if you get asked to assist with a project, making it clear in your stats plan what your exposure was, what your outcome was, who was in your eligibility population. And again, if there are changes along the way, which there sometimes are, there sometimes are in trials as well, have a record of all of that, have full transparency. But again, it comes back to this ignoring findings if they’re not statistically significant. So I think that one is particularly frustrating. Another one is like after the fact dropping people from your population after you’ve maybe not found a statistically significant result, you’ll go, actually I really meant to exclude these people. They’re not relevant for for these reasons. So if you’ve a a clearly defined-

Linda (36:23)
Just gonna look at the people it worked on.

Karen Lamb (36:27)
-exactly so. Yeah, it’s really tricky. Another thing too, again, coming back to your questions around the data, is just trying to use data for something that it just wasn’t designed to do. and so wh why are we why are we actually doing this? Like I mean a, a key example is this proliferation of these prediction models out there and you know, people will often come with data sets and they want a new prediction model, and you’re like, Well well why? There’s so many prediction models out there and they’re not used, they’re not validated, are they actually useful? This is research wastage. but it’s this feeling that it’s not gonna be considered important if they haven’t in a small study which is exploratory and still novel in its own right, it’s this feeling that descriptive might not be enough. And I think that’s really sad when you can still gain a lot of information from descriptive without going to all the effort of producing a model that is completely unnecessary.

Linda (37:36)
And not useful.

What is the first question you ask when you look at graphs in the media?

Karen Lamb (37:49)
Usually I groan when I see graphs in the media, but no, I will often-

Linda (37:53)
Me too.

Karen Lamb (37:57)
-I’m obsessed with looking at, you know, axes and where the start point of your axis is. Like, are you misrepresenting the data because you’ve not started at zero for example if you’ve got a bar chart whatever the thing that I hate is people don’t label their axes. This is not just in the media, people just don’t label axes and it’s so frustrating because like this is supposed to help you understand the information without you reading. So if you don’t label it, then well it’s no use to me. I really, really hate pie charts. Don’t get me started on how much I hate pie charts.

And I guess I think we have seen them misused a lot in the media as, you know, where they actually, you know, should sum up to a hundred percent and they’ll definitely exceed that. but just just since you’ve given me the platform to raise to get on my soapbox and raise my bugbears, I think-

Linda (38:57)
Go for it.

Karen Lamb (38:58)
-the frustration the frustration is if something can be more easily understood when you’re comparing relative heights. So a pie chart might be really useful if you’ve got a small number of categories, right? Even then I still don’t love them. They’re just not I people don’t understand differences in areas, as easily as they will height. I mean just mathematically it’s more complex. So if you’re wanting a simple descriptor to understand data, stop using pie charts, particularly when you’ve got so many categories, because I just I mean I don’t understand it.

And I’m a statistician. so I look for these thoughts, but the other thing for me, if I’m completely honest, and is overcomplex graphs that are completely impenetrable. Like people have gone gung-ho for complexity over conveying information. And that really frustrates me. As I say, I am a big advocate for simple, simple clarity. And you know, if a bar chart works, give me a bar chart, you know?

Linda (40:07)
When I first started teaching data science in high school, I was working with a maths teacher who kept saying, Why are we teaching graphs? We already do graphs in maths. This is pointless, it’s a waste of time, we’re duplicating effort. But then she saw what I was doing, which was teaching graphs as communication instruments.

So I wasn’t teaching the technicalities of the graph. I wasn’t teaching, you know, the the the mathematics of the graph. I was teaching what are you trying to say with this graph? What’s the what’s the one thing you want people to take away? What’s the best way to convey that? She was like, this is a revelation. I’m gonna use this in my teaching. It’s like why are we not doing this already? Like how can you teach graphs-

Karen Lamb (40:50)
And it’s

Linda (40:51)
-without pointing out that they’re for communication? That’s their whole purpose.

Karen Lamb (40:56)
Exactly, and I think I think I’ve heard you say something on those lines before and I found it w so bizarre because actually when I think of where I learnt best use of graphs, it was from like physics and chemistry at high school. It wasn’t maths. I actually can’t even really I mean I must have done it in maths, but it’s not where my memory is from. maybe I should have known that I wanted to be a statistician before because like, for example, in chemistry, I wasn’t the person that ever wanted to do the actual hands-on experiment. Like there was just so much potential for me to be clumsy and do something wrong and not be as accurate as I want it be. And

Linda (41:32)
I hear ya.

Karen Lamb (41:33)
and I just knew I’d be plagued with self-doubt of actually getting like everything, everything just so. so give me the the plot in your data. Give me the outlining your protocol of the experiment. And then so yeah, for me, I just remember doing everyone’s graphs in chemistry because yeah, let’s

Linda (41:50)
Ha ha ha ha

Karen Lamb (41:52)
let’s tell everyone what happened in this experiment. The graph’s the best way of of conveying this information. So yeah, definitely don’t remember it so much from maths.

Linda (42:01)
Yep. I mean that that’s because I think in in physics and chemistry you were learning it in context. You could see the point of it and in maths perhaps there wasn’t as much point. It’s it’s a bugbear of mine.

Karen Lamb (42:18)
just from the perspective of math education and the fact that I am such a practical person.

My huge frustration and the reason that I almost didn’t go and study it was because I didn’t know why I was learning what I was learning. But I was really lucky having the mathematician at home because he just loves math so much and actually is a very good communicator and a good educator. And every question I came back was, Why am I learning this? Why do I need to know this?

Linda (42:47)
Yeah.

Karen Lamb (42:47)
And my dad would be like, Well, if you’re building bridges one day, you need to know this because of this and this. And I would always

Linda (42:52)
Yeah.

Karen Lamb (42:52)
laugh at them because they’d be saying this to when I was, you know, pre-teen and I’d be when am I building a bridge? But but he had an answer for me and it brought it back into the real world and that’s what I needed. And I just-

Linda (43:05)
Mm.

Karen Lamb (43:05)
-didn’t really, if I’m honest, I didn’t really get that enough until I started studying statistics at university. Outside of, you know, my dad trying to go, you can do this, you will like it. It’s useful. Yeah, okay.

Linda (43:21)
That’s funny. I I love hearing you talk about how you just weren’t that into it in in school because

Karen Lamb (43:28)
Yeah.

Linda (43:28)
it I no I do because I mean it does make me sad as well. But like so many people feel that way now, you know, and and to think that you can you can not really be into maths or stats at school and then wind up associate professor of biostatistics at, you know, University of Melbourne, head of a research group and, you know, all this stuff and and Eureka Prize nominee and superstar of STEM and all of the things. It’s it’s it’s shocking really that you take someone who had obvious capacity and and tendency to to love this stuff and you put them through the school system and completely fail to awaken that. Like that

Karen Lamb (44:20)
Absolutely. I loved maths in primary school. I absolutely loved it. And then I just

Linda (44:25)
Mm.

Karen Lamb (44:25)
lost it going into high school. I lost my confidence with it. and yeah, really this is where again coming back to the mentoring, like people are just so important and having visible role models and backers, like if it hadn’t been for my dad’s and then I was lucky in my final year at high school in Scotland I had a math teacher who actually was a chemistry graduate who just kept going, Why are you not doing this? Like this you just really get this, it’s really useful to you. But I just didn’t know how interesting a job could be with maths because that was never how it was pitched.

Linda (45:07)
Yeah. Yep. I hear you.

I’m working on it.

Karen Lamb (45:14)
Yeah.

Linda (45:14)
this has been such a good conversation and I th there are so many bits that I’m gonna pull out and and, you know, force people to watch because they’re so, you know, on the button of of the things that are important. but we are up to the last question, which is my favorite, which is what excites you about data?

Karen Lamb (45:38)
People, always people. look, I think for me, I can I as I say coming back to my PowerPoint, the data’s one piece, but the excitement of the cool scientists, clinicians, health researchers I’m working with, the ideas they have, the innovations and the desire to do something to help people is why I do this.

And why I want to work with them to get the data they need is because they have this passion and they’re just such a joy to be around. And it’s all, as I say, about working with them and saying, you’ve got your area of expertise. I don’t know any of this stuff. Like, I’m not a clinician. Like, you need to bring me along with you, and I’ll bring you along with me, and we’ll work together and we’ll make this the best it possibly can be. and so yeah, I’m in this job because of the people.

And I’m just lucky that I happen to like the maths and the statistics that goes along with it.

Linda (46:40)
I love that. I’m putting that on the wall. Thank you so much. It’s been a wonderful conversation. It was worth the wait.

Karen Lamb (46:48)
I’m glad. I’m glad. I won’t make you wait so long next time.

Linda (46:52)
I’ll hold you to that.

Outro (46:55)
Thanks for listening to Make Me Data Literate. You can find more episodes at ADSEI.org/podcast and you can support our work at givnow.com.au/ADSEI. Have a great day.

Leave a Reply