Fiona Tweedie on Data Governance

ADSEI Logo with Make Me Data Literate across it
Make Me Data Literate
Fiona Tweedie on Data Governance
Loading
/

“really the heart of it is how do we organize this data in such a way that people can get the best value out of it.”

“in any situation in life if you want to make anything, having your tools in order and knowing where to find the stuff that you want to use is a really important first step.”

“the use you can get out of your data later is really going to be affected by the way you organize it up front”

“using standards to improve the interoperability of data so you can mix it up with other stuff from other places is so important to really increasing the value and the usability of that data down the line”

“I mean there are people like librarians and archivists who are absolutely amazing information scientists and have really thought about these questions of how you organize things and how you maximize discoverability. And I think their contribution is sometimes underestimated but yeah that is one discipline which has really thought about how you can organize information for for as a gift to future generations”

“there was a real disconnect between their ambitions for the data and the sorts of questions that they wanted answered and the data that they had”

“we’re going to have to talk a bit more about what you want to know and what this data can tell you”

“people don’t love surveys and you immediately have a bias in your your sample that you get responses from people who like doing surveys”

“I think it’s it’s really important to think about what you’re what are you trying to tell people with a graph and what’s what’s the question and what are you showing them and you know sometimes I think people are quite deliberate about about a distortion”

Transcript

Linda: Thank you for joining us for another episode of Make Me Data Literate.

Very excited for our guest today again we’re changing tack and looking at some new and interesting stuff that we haven’t covered before.

So welcome Fiona, who are you and what do you do?

Fiona: Well thank you very much for having me Linda.

My name is Fiona Tweedie and I work in data governance mostly these days and that means that it’s my job to try and organize data so that people can get more use out of it.

Also thinking about things like security and privacy and other laws that govern how we use data, trying to help companies and organizations set up policies to comply… meet their compliance obligations there but really the heart of it is how do we organize this data in such a way that people can get the best value out of it and people who want to do clever amazing analytics are able to do that because you know as you would know in any situation in life if you want to make anything, having your tools in order and knowing where to find the stuff that you want to use is a really important first step.

Linda: I love this topic because it’s it seems so obvious you know you just think well obviously you structure your data and it’s you know you make sure it’s easy to use but in actual practice it’s so rarely done.

The number of data sets I’ve encountered that are just chaotic and deranged and really impossible to analyze in some cases and in other cases just really really difficult because it’s very easy to to collect enormous amounts of data but to actually think about how you can organize it and how you’re gonna make it usable in future, it’s just it’s not something we teach and it’s not something we do instinctively I think, so it’s it’s really important and really neglected.

Fiona: I think that’s I think that’s very true I think that you know there are so many ways we can collect data now and people get incredibly excited about all the all the cool things they can do depending on their discipline of course. If they’re going to survey lots of people or if they’re going to put sensors all over a building but they do get really excited about that collection phase and of course the use bit you know getting to getting your hands on the data and doing the analytics and and producing some beautiful visuals is a lot of fun, but thinking about those usability issues up front, people don’t really talk about it and it’s not the easiest thing to teach. In my data career one of the the tools I’ve taught is Omeka, which is a an open-source tool that was originally developed for little galleries and libraries to be able to create online exhibitions of their of their collections but I was offering that as a tool at University of Melbourne for particularly HASS researchers who had collections of documents and images and things as a way of organizing those, and actually figuring out how to teach metadata effectively was quite complicated and it took me a few goes to work out how to how to do that and how to get people thinking that the use you can get out of your data later is really going to be affected by the way you organize it up front and it’s it’s hard for people to imagine future use cases and future questions but really trying to think about that that organization stuff the metadata using standards to improve the interoperability of data so you can mix it up with other stuff from other places is so important to really increasing the value and the usability of that data down the line

Linda: yeah it’s fundamental but as you say rarely taught and quite difficult to teach can you quickly give us an explanation of the term HASS for lessons who haven’t heard it before

Fiona: oh sorry it’s the one that’s not stem so HASS is humanities arts and social sciences so I guess people talk a lot about the stem disciplines being the science technology engineering and maths so HASS is the the humanities arts side of yeah there’s sort of two big buckets of research disciplines.

Linda: I’m increasingly of the opinion that a lot of the really interesting data is in the house space and it’s really potentially much more difficult to work with in a systematic way so it’s skills we really need to be focusing on thinking about and teaching

Fiona: yeah and I mean there are people like librarians and archivists who are absolutely amazing information scientists and have really thought about these questions of how you organize things and how you maximize discoverability. And I think their contribution is sometimes underestimated but yeah that is one discipline which has really thought about how you can organize information for for as a gift to future generations

Linda: I like that expression that’s great what did you have to learn to do your job what was missing from your formal education

Fiona: well my formal education I did an undergrad with majors in ancient history and Latin and then went on to do a PhD in ancient history, so I didn’t think very much about data and information at the time and there are things which I realize now would have made my life even as a research student in history much easier so we didn’t think very much about how we were organizing our information and even something like version control. I think it would have been useful maybe if someone had pointed out to me that giving files consistent names and incrementing a number on the end or just having some sort of a system makes you so much less likely to send the wrong version to somebody.

Linda: yep

Fiona: because I have done that but I think yeah so I didn’t think about things like version control, I didn’t think a whole lot about the best ways to organize information to make it easy for myself to find, and we had this we had this document that the University University of Sydney where I was studying sent out called the seven steps to completion for a research project and it had obviously been written for …really for researchers in the natural sciences because it had these this advice like you know you finish your literature review you know in your first six months

Linda:

Fiona: which is just not how well yes probably doesn’t work for science and certainly doesn’t work for the humanities where you’re you’re reading new stuff right up until the last minute and you know my and it talked about you know how to organize your data and you know I you do what data collection are you doing and you know my reaction was what do you mean I don’t have data I have books go away! but of course I did have data and so just thinking about the yeah you know are there ways in which I could have organized this information to be a bit more user-friendly down the line I think would have been a benefit.

I was lucky enough to get to go and do some field work in Italy and I had a lot of photographs when I came back of various sites that I’d visited but it turns out that as time goes past one pile of Roman bricks starts to look a lot like another in the photo reel and I hadn’t documented which photos belong to which site and as part of the process of importing my photos from my camera to my computer the photo program overwrote all of the metadata that was in the camera and so I wrote the dates.

Linda: Oh no!

Fiona: yes yes people always flinch when I tell this story I wrote those dates with the date of the import so I couldn’t go you know it was hard to go back and you know I couldn’t even match the dates to my travel itinerary yeah so the yeah the magic of metadata was something I really I think would have been helpful to learn a bit more about and just my subsequent career doing things like data governance and data analytics I again I’ve often thought maybe if I’d done a unit of an information science type subject or maybe database administration would have been really useful figuring out how databases work and how you can make data relate has been something that I’ve really needed to figure out.

You know what does it mean to have a primary key what does it mean to have a foreign key what does it mean to be able to join tables? Some of those questions which seem very sort of abstract but are actually very I think if you’re going to work with data in any volume are concepts that you need to understand and I mean one of the really useful units that I did do actually was symbolic logic which I actually did through the philosophy department but if you’ve done a bit of symbolic logic it means that it meant that when I encountered programming languages it made a lot more sense to me because all of those concepts like ands, ors and ifs and nots were already really familiar to me from the from doing that that logic and you know there’s that point where philosophy and maths really intersect in that that sort of formal logic space and so that was a massive help to me and I would you know I think it I think it would be really beneficial actually if everybody did that as a subject just because it teaches you so much about how to organise thoughts and how to think about how information relates and how you construct a query. so yeah more symbolic logic for everyone.

Linda: That’s really interesting I’ve never heard anyone recommend philosophy as a path to data science before but it makes perfect sense to me when you when you talk about it like that. I did symbolic logic from within a computer science degree but to sort of step it back from the technology and think about the the pure logic and reasoning part is super useful.

Fiona: Yeah and it can transfer to any programming language or concept that you want but it’s just getting that that way of thinking into your head.

Linda: Yeah I remember when I was teaching secondary school we did an assignment around the time it was a privacy assignment was around the time the government was trying to to make laws about metadata and revealing that it had no idea what metadata actually was.

It occurs to me perhaps we should step back for a second go well what is metadata and why does it matter?

Fiona: Metadata is the data about the data.

I think in his attempts to explain it Senator Brandis described it as the envelope that goes around the data. okay so the metadata is is the is the way in which you describe and organize your data so it provides the context which lets you interpret the data and one of the really important things about it I think is learning to use as much as possible standardized metadata so there are loads of metadata standards out there for different different things but it means that if you can describe whatever you’re talking about in a consistent way then it makes your information comparable with other sources and it makes it comprehensible to other people.

so you know and metadata standards are designed you know they can be very subject specific. one of the ones I’ve worked with quite a lot is Dublin Core which came out of that’s out of the Glam sector that’s galleries libraries archives and museums and that’s as a way that’s a way of describing objects in a collection and so things like the author and the provenance and what what something is made out of and you know the date it was produced and all of that sort of information which when you’ve got I don’t know a book or an artefact you know it tells you about it.

Linda: that’s really cool and you can think you can immediately see how that could be useful if you wanted to go okay where are all of the bronze items in the in the museum or all of the things that come from this particular time period or this particular artist

Fiona: absolutely and so having having that data and having it organized in a consistent way means that you can query across your collection and ask those sorts of questions and that was actually what I ended up doing when I was teaching Omeka and trying to get people to think about organizing information was this exercise called bag of thingies where I would put them into groups and just give them a pile of crap and say sort it. and most times people would kind of shuffle it into two piles and then kind of look at me. and so and then I would ask them those sorts of questions I was like okay can you find me the biggest thing can you find me all the blue things? can you find me you know all the things made out of natural materials? and you know so the better you can organize those piles and the more detail you can get in the better the questions you can ask of your collection later.

Linda: that’s awesome I am totally nicking that. that’s brilliant

Fiona: I did that once with a bunch of archivists and they you know of course being you know that’s their bread and butter and then practically constructed a relational database out of pipe cleaners to as part of the exercise, but generally you know people shuffle it into a couple of piles and are like now what?

Linda: yeah that’s that’s wonderful I love that I mean I could I can see that going straight into the classroom when you’re trying to teach kids this stuff because it you know we we used to have such drama with kids submitting electronically because you try to persuade them not to not to submit assignment dot py for their you know Python code because I was gonna get 50 assignment dot py’s and persuading them to put their name and the class or the teacher you know or the the question number whatever in the in the in the file name and getting 50 kids to adhere to the same standard was almost impossible so even getting people to think about the idea that there could be a standard and that you could do things in a systematic way would be I think a big step forward

Fiona: well yeah and I mean if it’s submitting an assignment or a CV or a job application I think it’s a really helpful idea to put your name in the file name just to help those recruiters to know who you are

Linda: yeah yeah everything that makes their life easier is as good news for you as the applicant is there one thing that you wish everybody knew about data one thing that would change the world for you if everyone understood it?

Fiona: that the answers you get out are only going to be as good as the data that you’ve got in the first place you know people talk about data science and everybody is terribly excited about AI at the moment but if the underlying data is not of sufficient quality you know it doesn’t matter what shiny tools you throw at it your results are going to be limited in their quality and their usefulness.

Linda: I really like that that’s that’s something that’s that I think is very poorly understood the idea that you know oh we have all this data we can extract amazing insights from us like well how good is that data how meaningful

Fiona: and I had a job I was the one and only data scientist employed by the Australian Ballet and they had had a review by a consulting firm who’d said you have a lot of data you should probably get someone to science that. so they so they hired me and none of us really quite knew what we were doing. and yeah the the problem was I think that they they just hadn’t … there was a real disconnect between their ambitions for the data and the sorts of questions that they wanted answered and the data that they had I mean they had a subscriber database that was great but then you know they would say to me things like “how do we future-proof the business of ballet?” I don’t think that answer is not directly in your subscriber database we’re going to have to break that down we’re going to have to talk a bit more about what you want to know and what this data can tell you and you know they were quite averse to surveys and subscriber surveys you know people don’t love surveys and you immediately have a bias in your your sample that you get responses from people who like doing surveys but it did mean it did mean that what they had was very thin and the sorts of things which they would have liked to be able to do just weren’t really possible and I think you know now that I’m a bit more experienced I could see that what they really needed to start with was a fundamental information management program and I think may in retrospect maybe insisting on that, because I mean they wanted to do some analysis on the success of their subscriber their subscription sales program and work out whether it was better to employ a third-party call center to do these sales or whether in-house was more effective, and the CFO asked me to look into that, and one of the things I found was that for some reason they had deleted the soft copies of the invoices that the call center had supplied to them and then archived the hard copies so they were in a box or warehouse.

Linda: Oh god.

Fiona: yeah right and so then you know trying to do some analysis on the data that just wasn’t there. and that the the information about what it cost to to run the internal call center was in one place and you know this other information was in another place and so even that work of trying to gather up enough information to be able to make a stab at an analysis took a long time and it’s a it’s a constant complaint of data professionals everywhere is that they can’t get the data that they need and yeah it’s because you know I think people just haven’t thought what if we what if we need this what what would we want to do with this and so yeah you know sending hard copies of invoices away to a warehouse was an absolutely reasonable thing for the accounting department to do because in their world they’d finished with that stuff. it’s like can we you know the idea that that is also information someone might want was kind of new to them

Linda: yeah not the way things used to happen or you know not the way people expected things to be done and your comment early on about how can we how can we future proof ballet reminded me of the the section and the head checkers guide to the galaxy where they ask this big computer to figure out the answer to life the universe and everything and it turns out to be 42 and like but that’s useless and the computer says well yes I don’t think you ever actually knew what the question was. so I think Douglas Adams may have foreseen the the world of data science because there seems to be an awful lot of 42 and not a lot of knowing what the question was.

Fiona: yeah exactly and yeah people get terribly excited about this idea of you know the amazing insights that we’re going to be able to generate but yeah I think without necessarily understanding what the question is or what they’re hoping to find out.

Linda: and what what questions the data set can actually reasonably answer. it’s one of the first things we look at when we’re not I teach data science and it’s something that doesn’t appear to be in any curriculum anywhere. what what can we actually get out of this data rather than what do we want to get they’re not necessarily the same.

Fiona: yeah and I think it can be a little deflating for students in a way to start with that question because you’re thinking about limitations and you’re suddenly saying well you know we’d like to know all of this stuff but the data probably can’t tell us that. but I think it’s incredibly important to be honest about that and not give people this false idea that they will be able to magic up answers if if the if the data set simply doesn’t support it.

Linda: yeah now there does seem to be a lot of that in the data science industry the idea that you know we can just we can just magic up stuff and it’s magic and give me some AI give me some data science magic things will happen a little more skepticism and an informed approach would be would be would be useful I think. what are the worst data mistakes you’ve seen?
Fiona: for an absolutely criminal use of data that still makes me incredibly cross robodebt in Australia. yeah that was an appalling abuse of data and one of the things which really came out of the Royal Commission’s report was the fact that people had been doing things with data and again that the data quality simply didn’t support and you know this had a consequence of ruining people’s lives. the department responsible had in fact tried to engage a data science team some external consultants and said can you AI this data for us? and I think it was I think it was actually data 61, and their response which was the absolutely honest thing to say was that your data is not robust enough for for what you’re hoping to do we cannot in good conscience AI that data because it’s just not up to it.

And people talk about you know people were talking about that program as you know AI run wild and algorithms run wild and it wasn’t. it was just crummy data analysis of yeah insufficiently robust data.

Linda: People doing things they shouldn’t have done ‘

Fiona. absolutely and I think people not thinking about consequences or people in positions of power refusing to listen to the consequences. I think one of the things that came out of the Royal Commission was the frontline staff were reporting these problems and that higher-ups were refusing to listen. but yeah the fact that you just couldn’t identify people sufficiently consistently across datasets, so people were getting debts which didn’t belong to them because they happen to have a similar name to somebody else.

Linda: oh wow.

FionaL yeah and in Australia there are some quite tough laws around what you can do with a tax file number and the ways in which you can use it. and a tax file number would have been a really good identifier to use for a program like this but because of the laws they didn’t use that, and so they were reliant on other identifiers. and so just even trying to data match people across sets that was done not terribly well.

you know the fact that they were asking people to comment on things from years and years ago when you’re not even required in Australia as an individual to keep your tax records for more than five years and yet people were being asked to account for you know their wages from a casual job they had for three months for a longer time period yeah there was yeah data abuses all the way down in that program.

Linda: yeah and it’s difficult to know whether that was as much mistake as deliberate and willful misuse but that is my next question have you ever seen data deliberately misused. Do you have a different example or do we talk more about robo debt?

Fiona: yeah I mean I don’t know how much of it was deliberate but whether at what point someone said I don’t think you know the data is of sufficient quality to for us to be confident in doing this. and I think thinking about yeah thinking really thinking about consequences that you know my work now I think a lot about data ethics and when people in in my company have got a bright bright idea for something I would like to do with data one of the things I get to do is sit down with them and say okay what’s the worst thing that could possibly happen you know who are the most vulnerable people who might be affected and what does that look like for them? and you know happily the work that we do is fairly low stakes and so it’s unlikely that anything really terrible is going to happen but yeah the higher the stakes if you’re if you’re raising debts against people you really do have to be confident that you’re right.

Linda: I was just gonna say that that is one of the things that we don’t typically teach when we teach with textbook examples and textbook data sets that they don’t have these kind of implications which is why I like to get the kids doing real projects so that you can actually, in the evaluation phase, go who is helped by this and who is harmed.

Fiona: absolutely and you know textbooks you know they give you these beautiful clean data sets for a reason. that you know they want you to focus on the on the analysis and trying to clean up a messy spreadsheet is less fun if you’re just wanting to teach the analysis. but you know as you’ve always emphasized data comes from the real world and it’s it’s messy stuff and you’ve kind of got to confront that.

Linda: yeah how do you spot deliberate misuses of data? what do you look for?

Fiona: I think as you said at the beginning you know or earlier, being a bit skeptical and the asking those questions about who’s benefiting? who paid for this? is a is a question to ask. asking questions about where data came from, you know what what data was collected? what’s the you know if it’s a claim about a population what was the sampling like? when was this done? I think are useful questions to start asking.

you know there was a when I was in primary school a magazine published an article saying that left-handeds don’t live as long as right-handeds which was you know distressing news to the left-handed and it turned out that the the study was actually of US baseball players which is a very specific population, and there are maybe some other variables which contribute to how long baseball players live.

you know for a bunch of Australian primary school students to extrapolate that to themselves and start worrying

Linda: that’s really interesting because I hear that bandied about all the time with you know particularly having a couple of left-handers in my in my immediate family and I keep hearing that stat I didn’t realize it was so very specific

Fiona: then there may be a representative, there may be other studies but this particular one was yeah of baseballers

Linda: so I would have to go look now what’s the first question you ask when you look at graphs in the media?

Fiona: I want to see the X’s and I want to see the scale. are the axes labeled? do they start at zero or do they start somewhere else? what scale are we using? are the scale on the X and Y axis comparable? is this an appropriate form of visualization for the data?

Linda: Oh I like that one

Fiona: when you get a line graph which bubbles up and down but it’s not the data isn’t actually a series it can be very misleading, because it implies a relationship that something is happeningm when you know if the data is categorical then no it’s just not actually a feature of the data and you should have a bar graph in my opinion in a situation like that.

Linda: can you explain the difference between the data being in a series and the data being categorical

Fiona: yeah so data being categorical means that it’s you can sort the characteristic into kind of into buckets I guess so you might have a data set that’s about pet ownership in Melbourne and your categories are going to be cats dogs birds fish

Linda and Fiona: sugar gliders

Fiona: whatever people are keeping and each of those and those are types of animals those are categories, whereas series data is something like time or temperature so if we have a we can have a graph that shows temperature during the day and our x-axis is going to be the the time of day and we’ll put temperature on the y-axis and the resulting line you know we can see that it goes it starts cooler in the morning, and then it goes up and then it goes down again in the evening, and it’s meaningful to put a line there because time is is something that’s continuous and so we’re looking at change over time.

whereas if we’re going to graph pets kept in Melbourne each of those categories of animal is is distinct and so I think we should put those into columns and we can see that there are more cats than fish being kept, and to use a line graph there and join up those points I think can imply a relationship or a change that isn’t actually there it’s not like the number of cats and the number of fish is influencing each other. well it might be but yeah yeah generally it’s not that yeah there’s not a relationship between those categories, and I think it’s important that you use an appropriate type of visualization so you’re not implying a relationship that doesn’t exist in the data.

I mean similarly you know everybody loves to bag pie graphs but pie graphs are useful for showing if you’ve only got a couple of categories, you know they can show proportion quite nicelym you can say that you know we can present that two thirds of survey respondents think X and one third of them think Y, and that’s a that’s something you can show quite nicely in a pie graph. but again if it’s not a if it’s not about proportion, if we were to put that that pet cat pet data into a pie graph for instance, I don’t think it would be as useful because it’s not like there’s a zero-sum game happening here, and that the number of cats influences the number of dogs it’s not that you could have one or the other, it’s you know some people have both, so it’s it’s not as it’s not as useful, so thinking about is this the right type of graph? yeah have you labelled your axes? are you telling me what’s going? on yeah have you what’s the scale?

because we’ve we’ve seen quite famous examples of where someone is deliberately trying to exaggerate a feature of data and they don’t start the axis at zero or they sometimes even break the axis or break break the count in order to make a difference look much bigger or much smaller than it really is and yeah so those are the sorts of things that I want to look at with a graph and then also thinking about where did that data even come from?

Linda: yeah yeah and one thing that that constantly worries me with those particularly with axes issues is that a lot of those issues are exacerbated by the defaults in various software packages and that you just get by default most graphing software gives you the range of the data on the y-axis rather than starting from zero so if the data is sort of you know from 98 to 104 it’ll look like there’s big differences between 99 and 101 look very different, whereas if you start from zero they you can see that actually they’re all quite quite similar. so the defaults are a problem and I think that’s you know that’s part of the education process around data science is what are the appropriate defaults and should you just be accepting the defaults from your from your software do you need to actually think more about the message that you’re trying to communicate with this data and and is it is it meaningful or are you giving a distorted sense of how your data really looks?

Fiona: yeah I think it’s it’s really important to think about what you’re what are you trying to tell people with a graph and what’s what’s the question and what are you showing them and you know sometimes I think people are quite deliberate about about a distortion

Linda: what excites you about data?

Fiona: that’s a great question. what excites me is how powerful it can be and the stuff you can do with it and the way in which I mean I really came to data I was working in policy for the Office of the Australian Information Commissioner back in the heady days of Kevin 07 and government 2.0 and there was a big push on then to open up more government data. that you know public sector collects huge amounts of information and the argument was made, rightly I think, that if this is collected with public money it’s public property and people should have access to it health data and things not withstanding. but so it was I was in that message really resonated with me and I was incredibly excited by this idea that by putting information like raw data into people’s hands that they could make things and find out things and really be able to to be much more informed about the world in which they live and I was involved for some years with the Open Knowledge Foundation and they talk about a vision for a world where information benefits the many and that’s really what I love about data is that its potential to sort of democratize access to information. to enable people to ask questions and to you know hopefully hold powerful people to account. to be able to call out some of the obfuscations and not entirely true statements that come from our leaders in politics and business.

Linda: yeah yeah I love that idea that you know claims get made and we can actually investigate the truth of those claims and you know go back and look and go well hang on a minute. and that was the great thing in some ways during COVID a lot of the data that we had was open and immediately open so the case numbers that we collected, for all of the issues around what was collected and how it was collected, and and and all of those things, what information we did have for the most part was available that day or the next day and that I found that really interesting as a you know a potential for the future. To say well what if we did that all the time? what if we made this stuff really open and gave people the capacity to to look for themselves?

Fiona: yeah and it meant that the conversations we were having were based off hopefully the same information. that you know if we had you know an agreed source of truth on COVID case numbers for instance, that you know sometimes I think our public conversations end up kind of going past each other because people are looking at different sources of information and prioritizing different things, but at least if we have a common data set we can argue about how to interpret it but we should hopefully be talking about the same thing.

Linda: yeah yeah that’s awesome thank you so much this has been a wonderful conversation and there’s so many bits that I think I’m gonna be stealing and putting in my workshops. Telling people: you should listen to this! yeah it was super interesting

Fiona: oh well thank you it’s been it’s been a lot of fun and I mean data governance is not a topic that gets a lot of love. people yeah switch off when they hear it, but it really is the makes a huge amount of difference to what you’re able to do with data and not just avoiding big regulator fines

Linda: yes there’s there’s what can you do what are you allowed to do and what should you do and they don’t necessarily line up.

Fiona: yeah yeah those are all interesting axes to explore, but if the data is if the data is crap then that’s then the “can” is very limited.

Linda: yes yeah really important message thank you so much!
Fiona: thank you Linda

Leave a Reply