Well it’s been a long time between episodes, but we’re back! And wow, have we got some amazing guests lined up. Starting off with, not so much a bang as a fast radio burst, we have Dr Sarah Pearce, Director of the SKA Low telescope.
“Science is a human endeavour. And no matter how kind of unbiased we try and be, it has its own sort of drivers, right? And similarly, there’s a question of what do we look for? And what do we find?”
Transcript
Linda: Welcome back to another episode of Make Me Data Literate. This one’s been brewing for some time now. I’m really excited to be able to have this conversation. I think we’ll go to some really interesting places. So welcome Dr Sarah Pearce.
Sarah: Thanks very much Linda. It’s great to chat to you.
Linda: It’s awesome to have you here. Can you tell us who you are and what do you do?
Sarah: Yeah, so I’m Sarah Pearce. I’m the SKA Low Telescope Director. So what does that mean? SKA Low will be the world’s largest low frequency radio telescope and we’re building it at the moment in Western Australia. It’s being built by a collaboration of 16 countries called the SKA Observatory who are building one telescope here in Australia and one telescope in South Africa. And perhaps for your listeners, the key thing to know about the SKA Low Telescope is that it is a big data telescope and that our data streams will be raw data, eight terabits a second, but that will have kind of 350 petabytes a year or so of processed data that we send around to astronomers around the world.
And so it’s one of the instruments that will produce the largest amounts of data. So it’s a bit different kind of data than you sometimes talk about on your podcasts, but sort of how we understand and manage that data and how we understand and manage the telescope itself is really important to us. And the reason I talk about understanding the telescope is because it will be at full scale made of 131,000 antennas. And so a really interesting challenge for us is how we monitor and understand and analyse the data that’s coming from the antennas themselves so that we know how to make the telescope work best, not just the data that we take from the sky, but the data about the instrument too.
Just to give you a little bit of background on me as well. So for the last 20 years I’ve been involved in large-scale science projects and been working in this role in the SKA for three years, but before that I was in CSIRO, Australia’s National Science Agency, working on this and with other radio telescopes. And before that I worked in a group that was doing computing for the large Hadron collider. So that’s another big data project.
So I’m a physicist by training, but for the last 20 years what I’ve been doing is helping to deliver and manage large international science infrastructure.
Linda: So this question isn’t on the list. You feel free to ditch it if you like, but it just dawned on me that there’s… Okay, now I’m going to restart. What do you think, how important is it to have a scientist managing a project like this?
Sarah: That’s a really good question. Alright, So a lot of the skills that you need for managing big projects are – and people – are quite sort of general kind of management, project management kind of skills. The advantage that you get, I think, from having somebody with at least some experience in the field is a better understanding of the kind of scientific process and the scientific method, right? And of the ways in which science instruments differ from kind of other sorts of instruments. So in some ways running our telescope is a little bit like running a mine, but in other ways it’s quite different because it’s very cutting edge and we don’t know quite how it’s going to work. And so it all has to be very much planned, of course, but delivered in stages and then we learn as we go.
And the key thing from all telescopes that are built is that you never quite know what you’re going to, what you’re going to find. So that’s the most exciting thing, but also a really big challenge for analyzing the data that comes out of the telescope because, I mean, I think what we know about AI is that it’s really good at finding things that it expects, but not necessarily very good at finding things that it doesn’t expect.
Linda: So that’s a beautiful tie-in back to the idea that this is just a monster amount of data. It’s that more data than anything else has ever produced, which as I understand it means that you’re having to develop new techniques and processes to manage the data, to process it. You can’t even store this amount of data. It’s just too much. You have to process it first and it’s a whole new science really, isn’t it?
Sarah: Yeah, look, I mean, in particular, the sort of real-time processing is and has been kind of a real challenge for these large-scale kind of instruments. So for some of the data that we take, we will have 10 billion data streams coming in, kind of 24/7 essentially, which we have to process all in parallel and then together.
And the key, the key point for us will be, obviously, we will scale up over time, you know, so I didn’t say we’re kind of in the sort of mid stages of building the telescope at the moment. We’ve got about a thousand of the antennas working. Eventually, there’ll be 130,000. So we’ll scale up as we learn how to process the data.
But it’s that sort of continuous data stream and being at the stage where, you know, once we’re at full scale, where we can really get the most out of the hardware that we’ve built on the ground through the software and through the hardware, and through the hardware in the Pawsey Supercomputing Centre. And I saw that you talked to, talked to Mark earlier.
So, yeah, so it’s a real data management challenge as much as anything for these kind of big data instruments. Yeah.
Linda: It just boggles the mind that you, you know, you think about how many laptop hard drives it would take to even approximate this amount of data. You just, you know, you can’t begin to wrap your head around it. But it’s, it’s fascinating, the potential, you know, to discover things that we don’t even know are there to discover yet. Just before I get onto the normal questions, I want to come back to – you said it’s the SKA Low and it’s a low frequency telescope. What does that mean? What are we, what are we listening to with low frequency?
Sarah: So what we’re doing is we’re observing the universe in a part of the electromagnetic spectrum, which is in the radio part of the electromagnetic spectrum. For those who speak in these terms, it’s 50 to 350 megahertz. So it’s the same kind of waves that your TV antenna, if you had one, would pick up. And then the telescope in South Africa will look at the universe in a slightly higher range of the radio, which is more the kind of radio range that your mobile phone uses. Yeah. And what we can see, we can see kind of different things. And in particular, we can see kind of back to different periods of the universe’s evolution by looking at these different frequency ranges.
Linda: Oh, that is so cool. We’re going back in history, back in time. We’ve got a little time, a big, very big time machine.
Sarah: Yes, exactly. So yes, we do, we just say we’re building a time machine when I talk to yeah, when I talk when I talk when we talk to kind of news people who’ve perhaps only got a couple of minutes to try and tell the story.
Linda: Yeah.
Sarah: Yeah, because the key, the key for the SKA low telescope is that with that low frequency range, you can see, you can see electromagnetic waves that have been stretched out over the history of the universe. And so their wavelength is much longer than it was when they were originally emitted. And so what we can do is we can map hydrogen right back to nearly the dawn of the universe, what we call the cosmic dawn about 300 million years after the Big Bang. And that’s when the first stars and galaxies started to shine. So what we’re hoping to be able to do, and I say hoping because nobody has ever built an instrument that can do this before. And so how we process this data, how we really understand the instrument and the sky in detail so we can take away all the extraneous things and see what the universe was like 300 million years after the Big Bang, that’s going to be our, that’s going to be our key challenge.
Linda: That’s just amazing. It’s, how do you wrap your head around that? I’m going to get to the questions, but I just, how do you begin to comprehend what you’re doing here?
Sarah: I don’t know. I always say to people, like, big numbers and little numbers, you don’t really understand them, you get used to them, right? It’s kind of like quantum mechanics or, you know, understanding the history of the universe. You get used to working in these spaces, whether it’s possible to actually get your head around the difference between, I don’t know, a thousand years, a hundred thousand years, a hundred million years, can you, 10 billion years, can you really understand that? I don’t know. I mean, I think it’s an interesting question, actually. You’ve probably seen online kind of graphics and videos that try and show the difference between a millionaire and a billionaire. You know, and it’s kind of similar to that, you know, in that it’s really difficult to understand the scale.
Linda: Yeah. Yeah. I just did a blog post about, you know, that partly was about understanding the difference between 80,000 and 83,000 is quite difficult, but if you put it as a percentage change, it’s something more comprehensible. We’re not really very good at comprehending these really big numbers, and 80,000 in your world is nothing.
Sarah: It is. If it’s 80,000 antennas, then it’s… And if it’s 80,000 years, yeah, no, it’s nothing.
Linda: It’s just amazing.
Sarah: One of the best kind of representations of that kind of data that I ever saw, actually, was, I don’t know, if you think of the Grand Canyon. And so along the edge of the Grand Canyon, there is a kind of, there’s a timeline that you can walk along and it shows you when humans were around and it shows you kind of all the different periods of evolution back to the kind of early stages of the earth. And really that setting that time in space, such that you can walk along it, really gives you a, you know, quite unique kind of perspective on what those long time frames really mean.
Linda: I haven’t seen that. That’s really cool. I like it. Making data comprehensible to people is just such an art. It’s fascinating to me. So you’re working, your organization is working with just, you know, problems that haven’t even been thought of before, much less solved before, which makes my question, what was missing from your formal education kind of, arguably, everything because all of this is new. But what did you have to learn to do this and how much of it was self taught?
Sarah: So, well, my qualifications are in physics and astronomy. So I did undergraduate degree in physics and then I did a PhD in instrumentation for astronomy. So I worked on a telescope called Chandra, which is a space based telescope still circulating now. But then I moved into, I’ve had quite a varied career. I worked in public policy for a while in the parliament in the UK, and then have done sort of project management outreach and now kind of organization management, I guess, if you want to put it that way.
So what did I learn? Look, so my physics degree was, you know, it was a good, it was a good degree and it taught me a lot about physics. It taught me a lot about kind of analysis and breaking a problem down and trying to come to a first pass solution, even if you don’t know exactly how things work, so that you can, you know, have a go at solving the problem and see how close you can get.
So I think it’s fascinating, but what there wasn’t, because when, you know, when I did my degree, it was quite a while ago, what there wasn’t was a great deal of focus on kind of large scale data. So obviously, we learned about errors and things like this. But we didn’t really learn about data processing. I mean, anybody kind of doing a new physics degree now, I think would learn a lot more about programming, but also about data, about statistics, about how to process big data, about how to deal with databases and things. And there really wasn’t kind of any of that. And so to the extent that I’ve needed to know that sort of thing in my career, I have kind of, it’s been self-taught.
But big data has evolved a lot over the time that I’ve been in the astronomy community. You know, when I was doing my PhD, data would be taken and it would be stored on tape and sit on somebody’s shelf for 20 years. And then eventually, they’d kind of try and work out whether they want to process it or not. You can’t do that now. The amounts of data that we have, well, I mean, still people might not process it in time, but the amounts of data we have, they can’t be stored by a person. They have to be stored by a system and then they have to be accessed. And it has to be understood how everybody can access that. And that’s a whole different paradigm to when I was a student, you know, and probably to when most of us were students.
Linda: I think I’m going to have to go and do some research now because your comment that physicists almost certainly learn about this now… I’m not sure how much they do. You know, part of the reason I started ADSI was that science degrees, not necessarily specialist physics, but at the time, which was only – hah – seven years ago, science degrees were not teaching data routinely. And that was astonishing to me because there isn’t a field of science that doesn’t rely heavily on data and data management and data analysis and processing and all the rest of it and visualization. So I’ll be interested to see how much physics degrees actually do cover that now.
Sarah: certainly it’s one of the first things that I think, you know, a sort of new PhD student in physics will learn is where to get the data from, how to process it. I’ve been quite interested to watch my daughter did a degree in biomedical science. And that’s not my field at all. But I was quite interested to see what skills and that’s quite data oriented that, what skills that that degree sort of focused on compared with when I was at university, which was quite different.
Linda: Yeah, there are some specializations, I think, that have picked it up really well and bio med and bioinformatics definitely up there. And you would think that physics would be too, but I’m going to have to go rummage now and see. Is there one thing that you wish everyone knew about data that, you know, would make your life easier or that you think would make the world run better if everybody understood this one thing?
Sarah: look, there are two, two things that I was thinking about here. And the first, I think is garbage in garbage out, right, as they say.
Linda: Yep. Like that one.
Sarah: So, you know, the first question I always ask when I look at data is, where does it come from? Like, what have we chosen to measure? Yeah, what have we kept? What have we thrown away? Yeah, because the answers you get are so dependent on what data sets you look at. And we see this with the kind of replication crisis in psychology, right? And, you know, and with the sort of in biomedical science with the requirement to kind of register what you’re actually looking at before you start looking at it.
Yeah, because it’s so easy to take kind of data sets and try and interpret them in a way that gives you the either the answer you want, or, or maybe not even, you know, something that you’re doing nefariously, but you’re just looking for interesting things. And we know that that is not the way that you get replicable good science.
And so, and so, what is the data? What is the data set? What can it tell us? And what can’t it tell us? That’s the kind of first question that I wish people would understand. What do we actually know? Yeah. And then the second is, you know, how things are shown and interpreted. What time series have you chosen to use? We see this in politics all the time, right?
Linda: Yes.
Sarah: At the moment. Yeah, that’s right. Has the market gone up 8% today or has it gone down 2% over the last week? And the story you tell, you know, depends entirely on what data you’re using to frame that. And that, that I think, you know, would be helpful if people could understand. I have a, you know, great deal of respect for people like Greg Jericho, who you, I know you interviewed earlier, you know, who did a good job of trying to explain the complexities of data sets. What, what are they really showing us, you know, and trying to, you know, sort of be independent from the desire to tell a particular story. But what does this actually mean?
Linda: And it’s not just what data you have. It’s what, it’s how it’s presented as well. And I think that that storytelling aspect is something that’s really missing from the way we deal with data. You know, it’s the idea that the data says it or it doesn’t. It’s like, well, you know, it all depends what questions you ask it and how you listen to it and how you know, how you, how you present it can tell a whole lot of different stories using the same set of data.
Sarah: And so it is, so I’m really interested in this question of kind of what data you choose and how that tells the story and how you present it, you know, and what different interpretations you can get depending on, depending on what you choose from the data, in particular time series.
One of my favorite things that are actually at the moment is watching these kind of time series that you can see on TikTok or Instagram or whatever. I don’t know if you follow the Australian Bureau of Statistics, but they have a fantastic Instagram page, which often shows you kind of how their data has evolved over time. So this ability now to show data, you know, kind of using time as an axis as well, like on video, I think is really, you know, an interesting evolution of how we try and understand data. It gives us kind of extra access that we haven’t had before.
Linda: Yeah, I find that fascinating. The change over time and sort of presenting the same data from five different census periods and things like that is, it gives you so much more insight, especially, you know, you see people saying, oh, there’s this, you know, this is terrible disparity here. And, you know, it’s getting worse. You’re like, well, actually, if you go back 10 years, to be able to actually investigate those questions and look at, that’s one of the things I love about data is it gives you the power to, you know, either explode myths or, you know, confirm the story.
Sarah: Yes, yes, exactly. And, you know, it’s so important about how you try and communicate that to people. And that’s what I wish, you know, people would think about when they see data. You know, who is telling me this? Why are they telling me this? And what are they trying to persuade me of?
Linda: Yeah, that’s those are key questions that I just blogged about absolutely this morning. So thank you for that little little hat tip to them.
Sarah: And look, it’s true in science as well as in kind of, you know, politics or public life or whatever.
Linda: 100%.
Sarah: Yeah, obviously, science is a human endeavor. And no matter how kind of unbiased we try and be or whatever, it has its own sort of drivers, right? And so, you know, similarly, there’s a question of what do we look for? And what do we find?
One of the most interesting things I found in recent years, we talked a bit earlier about finding things you didn’t expect. So radio telescopes track these, what we call transients, things that pulse on the sky. And so they don’t give off radio light all the time, but they give it off in pulses, sometimes very regularly. So there are things called pulsars, which are fast spinning neutron stars that give off pulses, thousands of pulses a second.
So we’ve been looking for many years for these very fast spinning things. But what we didn’t realize until very recently was actually there are also kind of relatively slow spinning things. And using the Murchison Widefield Array, people like Natasha Hurley Walker, and now using other telescopes to have found what they’re calling long period transients, where we see a pulse, you know, maybe every 10 minutes or every 18 minutes or whatever. And this had always been in the data. We just never looked for it. Because people hadn’t kind of worked out what the physics was there, and they didn’t expect to see that.
And so having with new search algorithms and things, we find these unexpected things in the data, which it’s really surprising to me that it was always there, but it was just in a time domain that we didn’t expect.
Linda: Yeah, that’s super interesting. You know, I think the naive perception, the naive public perception of a telescope, is that it gives you essentially a photograph that you just keep looking at it long enough and you’ll see things. But trying to find things in data is a different story.
Sarah: Yeah, so I mean, another relatively recent interesting story is about fast radio bursts. So they’re another transient. So they are extremely bright flashes of radio light that are given off in the distant universe. And the first of these was discovered with Parkes Telescope, and it was found in the Parkes Telescope data by a group which included Duncan Lorimer. So at first it was called a Lorimer burst, because they’d only found one of them, and they weren’t sure if it was real or not.
But the interesting thing is it was so much brighter than we expected that it really took a kind of analysis of the raw data to be sure that it was real. And then we found more. And now this is one of the most interesting parts of, I think, current astronomy is understanding fast radio bursts and using them to understand the universe between the fast radio bursts and us. But I think what’s likely is that if we’d had a pipeline that processed that data, and the raw data wasn’t available, that this would have been so unusual, such an outlier from what we expected. But we would never have found it if we’d been looking for it automatically, and only looking for what we expected. So this is one of the key challenges for telescopes like the SKA, where the data has to be processed as you bring it in, because there’s just way too much of it to store in its raw form, is how do we make sure that we’re not missing these kind of new and interesting things?
Linda: Not throwing away the interesting bits. That’s such a tricky problem. It fascinates me. What are some of the worst data mistakes that you’ve seen?
Sarah: This is an interesting question. I was trying to think about this in the context of astronomy at first. Look, I think we find things and then they’re shown not to have been interpreted correctly or whatever. I mean, that’s the process of science. So I don’t characterize that as a data mistake. The story of peritons is an interesting example there. So they look a bit like fast radio bursts that I was just talking about. And again, they were found with the Parkes Radio Telescope, but they had subtle differences to fast radio bursts, which made us think that maybe they weren’t actually from the distant universe. Maybe they were a local interference source. They were kind of the same, but not quite the same.
And so we kept seeing them. And then there was a great piece of analysis done by PhD student, which showed that they were occurring when staff who worked at the Parkes Telescope were opening the microwave without turning it off first.
Linda: That’s wonderful.
Sarah: It’s a really great example, I think, of looking at data and saying, look, this really doesn’t quite look what I expect it to look like. Let’s try and find out why it’s doing this. So that’s not a data mistake, but it’s a kind of correction.
Linda: I love that.
Sarah: We try and work it out. And we did. And then there was paper was published explaining what it was. And I think that’s a great story. Which is one of the reasons the SKA is right out in outback WA, right, to try to minimise those sources of interference. I know when my Pawsey friends go up there, they have to leave their devices behind and all that kind of stuff. Yes. Yes. So the telescope is about 800 kilometres outside Perth in a remote region called the Murchison, which is on the traditional land of the Wajarri Yamaji people.
And yes it’s there because all the phones and the cars and microwaves and fridges and everything that you use has electronics in it. The electronics is massively brighter than the very faint signals we’re trying to detect from space. And so the best place to do radio astronomy is somewhere where there are not many people. And the Murchison Shire before we got there was a Shire about the size of the Netherlands, but with 120 people in it.
Now, of course, now we’re there, and there’s the ASKAP telescope, which CSRO runs, Murchison Wide Field Array, which Curtin University runs. And we’ve all got people there. And so one of the key things is not to interfere ourselves, right? Because our team would like to use microwaves as well. And indeed we can, so long as we’re far away enough from the telescope.
Interestingly, we’ve recently been, so rather than microwaves, people are interested in using air fryers to heat up their lunch when they’re out near the telescope. So we’ve been in our radio frequency interference testing chamber, it’s called a reverb chamber, we’ve been testing an air fryer to see if it’s okay to take it out to the near the telescope to heat up your lunch.
Linda: That’s awesome. And is it?
Sarah: Haven’t seen the answers to that question yet, Linda.
Linda: That’s so cool. So there’s kind of corrections on that, you know, that process of science. Have you seen any deliberate misuse of data?
Sarah: Yeah, I mean, of course, I mean, we talked earlier about how people choose the data that they use in order to, in order to kind of tell the story that they want to tell. And that’s even without thinking about when people, you know, deliberately or otherwise, actually misrepresent the data that they’ve got. I mean, I think we see that all the time.
I was looking at a story the other day about kids and gaming. And it said, in the US, the average child games for eight hours a day. And I looked at it and went, that just won’t be true. I mean, it’s just physically not possible. And actually, the statistic is that the average child in the US looks at screens for eight hours a day. But that’s very different. Gaming. Yeah. I was like, where’s all the time for watching TikTok and YouTube and everything else people are worried about all the time.
So, so I think there’s that question of kind of accuracy and really understanding what the data is telling you, whether it’s being deliberately misinterpreted or not. I wouldn’t like to judge, you know, but I think there’s a tendency to try and tell the story that you want to be told rather than what’s actually there.
Linda: or the story that gets the most clicks.
Sarah: Yes, exactly. Yeah, yeah, that’s right. Which is, yeah, which is the same thing in journalism, I think.
Linda: Yes. So how do you spot things like that? You know, I mean, in that example that you gave you with like, well, hang on, that’s not with 24 hours in a day. I don’t see how that adds up.
Sarah: Yeah, kids at school, things like that. We know that they’re on social media a lot. They can’t be gaming the whole time.
Linda: What else do you look for?
Sarah: So I look for, as we talked earlier, where the data comes from. Sometimes I like to try and look at different interpretations of the data. You know, so sometimes you’ll get a news story and you’ll be like, well, I’d be interested to see what, you know, what somebody else says about this. So I mean, just this week, it’s not really a data question, but there was a new story about somebody having resurrected the dire wolf. Which wasn’t what they’d done. What they’d done was taken a normal wolf and changed 20 of the genes in it. So it looked a bit more like the dire wolf.
And so you go to see those first headlines, and then you go to, I think, what are reputable sources of, you know, scientific analysis and expertise and say, well, is that what they’ve actually done? And it becomes kind of reasonably clear that, no, that’s not quite what they’ve done. It’s still very interesting, obviously.
And so it’s the same with data, I think, is what I do. If I look at something and I say, you know, does that seem kind of reasonable? And who else, what are other people saying about this data? You know, what are the different perspectives that people have on it? I mean, when you look at all the usual things, you know, the time that they’ve chosen to use, the axes that they’ve kind of graphed that they’ve presented it on. I mean, I do think that, you know, how we communicate information is changing really rapidly. At the moment, you know, my kids watch a lot of information on YouTube, you know, it has graphs that pop up here and there, show you different things. Sometimes they send those to me and I’m like, that’s kind of right, you know, if it’s in a field that I know something about.
And I think as people who use data and you’re thinking about this a lot, I know, you know, we really think we need to think about how the sort of coming generations are receiving and understanding data and what the different ways of which we communicate with it.
Linda: Yeah, 100%. I’m seeing a lot of graphs, which is really quite alarming. I’m seeing a lot of graphs going around where that just the actual visual element of the graph bears no relationship to the data and the text, you know, you’ll have 100 and 50 and the 100 bar is smaller than the 50 bar. And you’re like, what are you doing? So, you know, increasingly the case is to teach data literacy is just go “do the numbers match the size of the bars?” I wouldn’t have thought we’d have to teach that, but we do.
Sarah: And does the graph, you know, if there is a graph, are the axes kind of linear, you know, have they have they just expanded the bit that they want to show you and the rest of the axes have been squashed, which they do quite regularly. And it’s, you know, it can be very, very misleading. And I guess we get back to that, you know, what’s the story of the data? You know, what is it actually telling you?
Linda: Yes, what does the graph tell you? And is that actually supported by the evidence that we have?
Sarah: Yeah, that’s right.
Linda: So that ties in. Perhaps we’ve already asked the question, what’s the first question you ask when you look at graphs in the media kind of answer that? Is there anything else you look for?
Sarah: I guess I look for, you know, what kind of representation have they chosen to use? You know, does it make sense that this is a bar chart? Does it make sense that this is a line graph or have they have they kind of not understood that? Or have they chosen this for a particular reason? Because it helps with the story that they’re trying to tell you if this is not a way that you would normally kind of normally present this – pie charts can be particularly bad for this, right?
Linda: Yeah.
Sarah: And so so that’s one of the things. I mean, moving away from sort of data that you see in the media and things, though like one of the interesting questions for us is, you know, with 130,000 antennas, how do you present that information so that our operators understand which parts of the telescope are working, which parts of the telescope are not working? If something’s not working, why is it not working? Right. And I was talking to some of the user interface people yesterday. In fact, that’s a really interesting challenge for us.
It’s a big instrument. We need to be able to present kind of different layers of data in a coherent way so that it’s not just the data, but the kind of knowledge that that data represents that is being communicated to the operators, so that they can really understand, well, the reason this station is not working is because the power generator failed or whatever, right?
And so this is a real, as we build, scale up the telescope, you know, at the moment, we’re kind of watching a thousand antennas at a time, and it’s more or less possible to do that one by one. But not when you’ve got 131,000. And so how that data is visualised for us and making sure that data is interpretable and flags areas that are important, that’s a real, really interesting question. And one that other kind of industries kind of manage as well.
Linda: That’s fascinating. I’d never thought before about the data about the data collection system and how important that can be. I remember and I love that you framed it as a user interface question, partly because that used to be my research area, but also because user interfaces are so pivotal sometimes in making sure systems work and there’s stories about, you know, nuclear power stations where issues with the interface meant that a problem wasn’t picked up or wasn’t, you know, fixable and it wasn’t actually the reactor that was the issue. It was the buttons and the communication with the operators that was significant.
So it’s fascinating to think about 130,000 little devices in the desert and how do you make sense of that, you know, massive field of information, without even thinking about the information that they’re recording or collecting.
Sarah: Yeah, and one of the sort of interesting things for us as well will be, you know, to what extent can we analyse that data to work out beforehand what we might need to fix, for example. You know, so to do kind of predictive maintenance, to what extent will AI be able to help with, you know, identifying stations or antennas that are likely to break soon. So these are all things that we’re kind of looking at and exploring, you know, even aside from the actual data that the telescope itself collects, which is from the sky.
Linda: It reminds me of that project at Pawsey where they’re tracking patients with head trauma and predicting increases in intracranial pressure before they happen so that they can, you know, intervene before the crisis point hits. I guess you want to do the same thing with the telescope is intervene before it falls over rather than after.
Sarah: Yeah, look, and I mean, the good thing about, you know, a very large interferometer is that, you know, a fair bit of it cannot be working without taking it down completely. But there obviously are kind of areas that are single points of failure and some failures that affect, you know, larger chunks of the telescope than others.
And so we absolutely need to learn as we go. And I mean, that’s, that’s I think one of the things that, you know, that both in sort of data and in science is really important that we don’t do particularly well in public life, is learn as we go.
Linda: Yeah.
Sarah: And actually, you know, look and say, did this do what I expected it to do? You know, what are the lessons we’ve learned from this? How can we improve that kind of continuous improvement? And so, I mean, obviously, there are people who are trying to work in that kind of space in say, public policy and things, evidence based policy.
But, but we don’t see a great deal of it.
Linda: No, we really don’t. I use that as an example in my talks all the time. We’re like, what if, every time we introduce a new policy or a new program, we actually then evaluated it to see how it worked. And everyone’s like, oh, that would be amazing. I’m like, why is that not normal? Why? I’m gobsmacked that it’s not the default.
Sarah: No, and that’s because, you know, to a certain extent, politics is a mix of, you know, achieving what you want to achieve, but also achieving your goal with some wider political aims, which might only be partly driven by what you, what the, what you set up to do, might be partly driven by whether you’re going to get reelected or whatever.
And I, there’s, I don’t know if you have read, there’s a book called Expert Political Judgment by Philip Tatlock. Have you read that?
Linda: No, I have not. So, what he did was, he took a load of experts in different fields and asked them to make predictions about, about what would happen broadly in politics and things, and in their fields sort of specifically. And then he went back and saw whether their predictions were any better than random chance. And the answer was mostly not, no.
Linda: Oh, wow. That’s amazing.
Sarah: Especially, I need to make sure I get this the right way around, but people who were quite narrow and had a particularly sort of narrow perspective and thought everything could be explained by one kind of way of thinking did particularly badly. And people who were sort of a bit broader, typically did a little bit better, but nobody did great. Right. And so that’s a really interesting, really interesting kind of analysis and great, such a great idea, right, to put together, to have this whole swathe of kind of geopolitical indicators that you ask people to predict and then see whether they were right or not. Because we all see people on chat shows or whatever saying this is going to mean this. And then next year it’s completely different. Do they ever get called up on it?
Linda: Ah, that’s so interesting. I am definitely going to have to dig out that book. I can see that making it into future talks. This has been such a great conversation. And as always, we wind up in completely unexpected spaces, which is, you know, half the joy of the podcast for me. Final question, what excites you about data?
Sarah: The ability to show us something that we didn’t know, I think, you know, whether that’s something that we haven’t looked for before, or something where we’ve been misinterpreting or misunderstanding things, you know, the double blind clinical trial and the data that results from it is one of the great inventions of the kind of modern world, right, which has stopped us saying, well, you know, stopped us using anecdote for data in medical science.
Linda: I wouldn’t say stopped…
Sarah: But at least allowed, you know, allowed an evidence based approach to: do drugs work or not? Right, which you look for the, you know, thousands of years of human history before that we didn’t have. Yeah. And so the reason that we’ve seen, you know, such great advances in life expectancy and things over the last couple of hundred years and the last hundred years in particular, is in part down to methodology, and down to our ability to kind of interrogate data and understand what it’s telling us. Yeah.
And so that’s what excites me about data is we can find fast radio bursts. We can find, you know, things that we didn’t know were there and we can, we can understand how to, how to make life better, how to understand our universe better. And we can’t do any of this if we don’t have the right data and if we can’t interrogate it in the right way. So, yeah.
Linda: What a wonderful pitch for my work. That’s a great place to end. Thank you so much, Sarah. This has been a fabulous conversation. It’s been a delight to have you on.
Sarah: Great. It’s been fantastic to chat to you Linda. Yeah, I will look forward to listening to future to the podcast.
Linda: Thank you.
Sarah: Thanks.
