Antony Green on the changing election data landscape

A fascinating conversation with Psephologist Antony Green on how the collapse of the two party vote, and the rise of independents and One Nation, impacts election night analysis. Plus the usefulness of polling, and why polling results get misinterpreted.

Transcript

Linda (00:00)
Welcome back to another episode of Make Me Data Literate. once again I’ve got a repeat guest this time, which we don’t do very often, but I’m excited because the voting landscape has changed significantly and I think it’s made some really challenging changes to the data. So we’re very exciting to welcome back Antony Green. Thanks for coming.

Antony (00:19)
Thank you.

Linda (00:21)
This is gonna be fun. so I wanted to jump straight into it. I know that you’re retired. you don’t seem entirely retired as far as far as I can tell. How how how does this sort of voting landscape make you feel? Like how w why are you still doing this stuff?

Antony (00:44)
I mean, I’ve just stepped away from doing television and the sort of having to do all the stuff I used to do at the ABC I can sort of pick and choose what I do now. but I’m still still have a big interest in the whole process and it is becoming more complex. Maybe it’s a good time to step away from it because it’s become much more complex. That the old models we used to use where you could predict with pretty good accuracy from early on, those are getting more complex to do because voting patterns are changing.

Linda (01:02)
Yeah.

So that’s exactly what I wanted to talk about because you used to be able to rely on the fact that it would be in any given seat almost certainly a two two person race. you’d you’d have two front runners, they might not always be Labor and LNP, but but you’d have you’d only have to really work with two main contenders. that has broken down significantly almost across the board. How does that change the the work of psephology?

Antony (01:49)
Well, it used to be that in every seat, you you had, you knew the two candidates beforehand. The two major party candidates would get eighty, eighty-five percent of the vote. There would be preferences from a small number of smaller votes, fifteen, twenty percent. But that’s what the model was. And everything was very one-dimensional. You could do everything based on a swing from Labor to coalition. Every seat could be lined up as safe Labor, safe coalition. or gradiated in between. And when there was a swing, they tended to be the same all across the country, they tended to be consistent within within electorates by polling place.

But what’s happened since is in the old days they would count those primary votes, then count the preferences and they’d be phoned through to us. Nowadays there’s a lot more votes in that third party pool. So it takes longer to count them if you’ve got the right pairing. And if you’ve got the right pair wrong pairing, so say they’d pick to count Liberal versus Labor and it turns out to be Liberal versus an independent. Then the Labor in the labour/Liberal preference count you get is useless. It doesn’t tell you anything. and it takes longer for that to arrive as well. So there’s a lot of where there is a big delay now in that more accurate preference count arriving if it’s the right pairing. So that delays the really reliable data till later in the evening. Secondly we’ve got bigger pre-poll centres they come much later in the evening, large numbers of votes, they take longer to count, and because they’re not cast on the day, they cast beforehand, they behave a bit differently than on the day. And not as not it’s not consistent from election to election. Sometimes there’s a big gap between the major party vote, Labor Party vote.

Linda (03:23)
Yeah.

Antony (03:41)
We tend to talk Labor versus non Labor just in terms of modelling. It’s not a political view. It’s just historically that’s the way it tends to be the Labor Party has been there for over a hundred years. The coalition has been tended to be different names over the years. So we tend to talk Labor non-Labor. Once upon a time that swing and variability of Labor non-Labor used to occur between the primary votes, and that would be reflected in the two candidate preferred. So if Labour’s primary vote went down, the coalition’s primary vote went up. And that was reflected also in the two candidate preferred.

But what’s happening nowadays is there’s multiple distributions going on in an electorate. It’s not just Labor versus Liberal. You may have an independent and their vote will be different across an electorate. The Green vote will be different across an electorate. And there’s a much larger pool of third party votes. And if they’re not flowing as you’d expect previously, it affects your prediction. And then much later in the evening, those pre-poll votes that arrive, which are cast beforehand, sometimes they’re more strongly very strongly towards the coalition, leaning towards the coalition, sometimes less so. But it varies from election to election in a way which doesn’t seem to correspond exactly to what’s happening on the date. And that’s making the whole thing much more difficult. And then you get the issue of if you’ve got the wrong pairing, then you can’t use a preference count. But what if you’ve got three candidates with roughly equal votes and at one point one will be ahead and then another? It means that instead of just picking who’s past fifty percent after preferences the way we used to do all the mathematics, nowadays you’ve got to look at well, who’s going to finish first and second here and if in you know what what’s the way to predict Who’s gonna finish first and second? We do information on primary votes, we look at the polling place patterns. But as I said, it’s not just one dimensional in terms of the swing. The the when I talk about distributions, I mean I’m talking about a mathematical and normal distribution. there’s the distribution of votes across an electorate for it used to just be Labor and Liberal, they used to be inversely related.

Nowadays as I said, you’ve got the Greens or you’ve got independents and they’ve got their own distribution across electorate. So feeding up to the preference count, you’ve got this difference between Labor and Liberal, but you’ve also got differences from area to area in these other candidates and their preferences. So the the pattern of votes at the top level when you get to two candidates preferred isn’t as reliable. The numbers are more variable. They’ve got a bigger variance over them than they used to have. now We use polling places to try and counteract for statistical bias that early figures are not reliable for later. That’s being undermined by pre poll votes because we don’t always know what those big final booths coming in are going to look like. but we’re also just getting a bigger variance on the numbers and that just makes it more difficult to build a confidence interval to predict who’s winning.

Linda (06:21)
Can I come back to that idea of getting the right pairing to start with? I think it’s possible that some of our listeners don’t understand this idea of setting up a pairing to start with. I certainly don’t. Can you can you explain that for us a bit?

Antony (06:37)
Okay, historically in Australia, preferences were only ever distributed in seats where a candidate where no-one reached fifty percent on first preferences. Now starting in 1984, at the end of the election they started to do a full preference distribution in every seat. So we had that data available. In 1990 we had a very close election and we had the com- electoral commissioner started to do this work to compare first preference votes. But come election night we suddenly had twenty percent of the vote with third parties, much higher than previously, and there were no preferences available. So it was unclear who had won. In in fact at that election, Democrat preferences flowed much more strongly to Labor. The Labor Party actually campaigned and advertised to get those preferences, so that was one of the reasons why there was a shift. So starting in 1993, they started to do what’s called an indicative preference count on the night. They they count all the votes in a polling place, allocate them to candidates, you know, by what’s count the first preference votes. And then beforehand they nominate two candidates as likely final pairing. And they look at all the votes for the other candidates and work out for those other ballot papers which are the final two candidates they go to and they tally that preference destination, that preference flow data, add it to the primaries and phone that through as a two candidate preferred. It’s worked perfectly well for a very long time, but it’s getting harder and harder to be certain who the final pairing is.

And the commission has to make that decision beforehand. If you’ve got fifty polling places in a district and you’ve got pre-polls and you’ve got postals, you can’t count some of the votes and make the decision. Otherwise you’ve got people sat around in polling places all night waiting for everyone to decide what preference counts they should do. You have to nominate it before there’s no way around it. You you have to nominate them before. If you pick the wrong pairing, that’s it. You’ll just do it again in the days after the election.

Because after the election all the ballot papers come in from the polling places, they’re all at one central centre. You can actually then count the preferences again. And that’s what they do. But you’ve got to make that decision on the night. And as we saw in South Australia, two thirds of the decisions they made were wrong. They basically they basically picked Labor versus Liberal in nearly every seat across the state. Because historically that’s what you’ve had. And then One Nation finished second in two thirds of the electorate.

Linda (08:44)
Wow.

Antony (08:56)
So half of those preference counts done between Labor and Liberal were pretty useless for trying to predict Labor versus One Nation. Now, it wasn’t a close election, so it didn’t really impact any. But it did mean the preference counts we got on the night were wrong and they were all redone on the Monday and Tuesday after the election. people can say, well the Commission should have done that. One Nation was polling strongly. Well One Nation had polled strongly, but they had never actually on election day got those sorts of votes. So they went with the more historic pattern.

Now we’re coming up to Victoria later this year and it’s going be just as wild. most recent polls have all three parties between twenty-six and twenty-seven percent, which makes it very hard to make a decision beforehand. They’re gonna have to make a decision in every seat who they’re gonna count to. and to not count towards one of at least one of the two parties in the previous election, especially if they’ve got a sitting member, is quite a decision by an electoral commission.

Linda (09:37)
Wow.

Antony (09:54)
I must say these these decisions are made in secret. They are not released before 6PM on the night. some of us sign a a a secrecy under undertaking with the electoral commission. So the media with tally rooms tend to have some information on this. But and some states don’t tell us anymore. But it means it you don’t that decision has to be made beforehand and it has to be kept secret because if it gets out people say, no, a a candidate would say, ‘Well, there’s no point voting for so and so the Commission’s not counting preferences to them’. I remember a decision one of those in 2013, which is the election of Indi, where Cath McGowan was the independent who won the seat. And the commission made the decision to count Liberal Labor, which meant it was useless for us on election night. I remember ringing the Commissioner for Victoria and saying you’ve picked the wrong pairing but it was too late to change.

Linda (10:28)
Right.

Antony (10:52)
And if that decision had been public, the Liberal Party would have been telling voters there’s no point voting for the independent the Commission’s not counting preferences. I mean that’s that’s misinformation because it’s irrelevant. They get the wrong pairing and they count them all the next day.

Linda (11:06)
That’s not the way it works.

Antony (11:07)
So the information has to be kept secret so that that decision, which is just for information purposes, isn’t used to influence how voters change. When we vote, it’s a preferential voting system. We want voters to list the candidates in the order they would like to see them elected.

And that’s the basic principle of the electoral system and the counting method. The counting method assumes that the votes you have got are by each of the voters freely choosing to select the candidates in the order they would like to see them elected. And so the least preferred candidates get knocked out and their preferences get distributed and they get the final preferred candidate. That’s the basis.

Now I can give you an example of where that’s completely breached as a counting method. People remember that old ticket voting system in the upper house where you have the above and the below the line votes. Until recently it used to be, until 2016 federally, if you voted above the line your preferences went to by went by the ballot the the lodged ticket of party. And the parties would do deals. So all the deals, all the preference tickets, ninety five percent of the votes were under control of these tickets. And the tickets weren’t done on the basis of a preferred ordering. They were done on the basis of strategic patterns, alliances, keeping votes with small parties ahead of large parties. And so what happened is if you, if you’re counting the votes and you presume that these are in a preferred ordering and actually all the votes are in a strategic ordering, if someone makes a wrong choice in their preference ticket, if you get a very close contest, then a preference decision which if if one candidate was one vote ahead, the preferences would all go off to the right. And if the other candidate was ahead they’d all go off to the left. So you get completely different outcomes. Which wouldn’t occur if voters filled in their own preferences because you wouldn’t get such strong patterns.

And people remember back in 2013, in WA Senate election, they lost some ballot papers about twelve hundred. Now, if voters had filled in the preferences, it wouldn’t have mattered, that wasn’t enough to change the results. But in fact, because of these tickets, there was a very close call, there were perpendicular preferences, preferences going in opposite directions, and a two or three vote difference changed two final seats in that election.

Linda (13:14)
Wow.

Antony (13:14)
And the difference between candidates 9th and 10th. Now you couldn’t have modeled that. It was a chroni…, it was an unstable system because the tickets that were involved were strategic. They weren’t showing preferences. And that’s why you don’t normally get that in our system in the lower house where voters feel their own preferences. You do get situations where you’ve got a three-way split. You might have a Labor candidate, a Liberal candidate, a Green, for instance. And the Labor Party would preference the Greens, the Greens would preference Labor.

But the Liberals would preference Labor. So you get the situation where the Greens might be first. there was actually which was the one? There was a case in Brisbane, the twenty twenty two election. The Liberal Party were ahead with the high thirties, and the Greens and Labor were roughly equal, about twenty nine percent, and there was about eight to ten percent off with some other candidates. If Labor finished third, the Green would win. If the Green finished third, Labor would win.

If they’d been even closer and the L N P had finished third, Labor would have won. So in a close three-way contest like that, the ordering of the candidates matters. It’s not common for that to happen. And preferential voting allows a candidate to come to mind. And if the Labor and the Green votes coalesce against e towards each other, that makes sense because they do preference each other. So you’re looking at in that case the LNP versus Labor and the Greens, and you get the right result because most voters didn’t vote for the LNP. But you get odd do you can get the odd situation where it just is a, it’s a different result depending on who finishes.

Linda (14:43)
That’s that’s wild. I I didn’t realise how how complicated it could be, you know, based on decisions made before the election. but

Antony (14:56)
I mean those decisions are made before the election but they are and they they are indicative counts. They are indicative of the result. If they have picked the wrong pairing, it’s all redone on the Monday. and that’s how it should be. but it’s a difficulty for the Electrical Commission, they have to make a decision. They don’t have any polling data. and the closer the contest the harder it is to be certain of that polling data.

Linda (15:02)
Yeah.

Hmm. Hmm. Absolutely.

Yeah.

Yeah.

Antony (15:26)
If the South Australian Electrical Commission have picked One Nation as the second candidate in twenty seats and it hadn’t finished that way, they would look like fools. So the commission does tend to err on the side of what they have seen before. So if you have a seat where the Greens have done well in the past and they’re in race, there’s there’s also a decision made if you’ve got a seat where the Greens are polling strongly and it’s like a three way split, then they may make the decision that they’ll pick a pairing, say like Liberal versus Labor, and if the Greens finish second, because those Labor and Green preferences swap, you can still have a rough idea what the result is. It’s where you’ve got a preference role for an independent that’s come on the scene since the previous election, and you’ve got no information on how preferences flow to them. That can be difficult if you’ve got the wrong pairing. A couple of seats in South Australia which were closer as two party contests, and you really need to see that second One Nation count to be certain.

And in the case of South Australia, I remember there was one seat which One Nation did win. And the One Nation, on election night, the One Nation candidate was first, the Liberal candidate second, and Labor third. And on that basis, you would say Labor preferences would flow to the Liberal and the Liberal would’ve won. But the minute I was looking at the results on the night, I could see there was one big pre poll centre to come from Gawler which is an urban area.

It was a mixed rural urban seat. And I knew from past patterns and I could see from looking at the comparison of polling place right b places by polling place by party. You could see that when that pre-poll centre would come in, Labor would move ahead of the Liberal. And the Liberal preferences would then elect One Nation. So on the night, the numbers didn’t show that that Labor would finish in the final pairing and One Nation would win. The numbers we had said the Liberal would win on Labor preferences, but we knew the order of those candidates was almost certain to change when that pre-poll came in. It came in on the Monday, and that proved to be correct. So there’s a lot of data you look at, and you set up the data and the polling place’s history and the and and just to clarify to people who are listening, we operate on swing. We operate on change in votes since the last election. If you just look at the raw numbers, they will go up and down as the votes come in. but if you operate on swing.

The swing from polling place to polling place, which is the change in vote. If you’ve got a result last time and it’s sixty percent for one party, if you’re trying to predict, you don’t start from zero, you what you operate on is comparing the figure by polling plates from last time. Say they’ve got three polling places in and you say, well, the swing is five percent. You can say, Well, it was sixty last time, there’s a five percent swing against them, so our prediction is fifty-five. That’s more accurate than just using the raw numbers because swing is less biased, it tends to be consistent across polling places where just the raw polling places, the size of the booth and the labour non-labor vote is correlated. So, you know, country electorates come in early in the night so the labour vote is low and it rises through the evening. If we operate on swing, that doesn’t happen. It’s much smaller. Secondly is it’s got so it’s got less bias, it’s got smaller variance, so it’s a better thing to make your predictor of the final result. but increasingly as I said if we haven’t got the preference counts, we’re having to use other decisions. We can still compare on primaries, so the primaries, the first preference votes will s and the swing on first preferences will still give us a guide. But if you’ve got a contest where last time you had Labor and Liberal and you’ve got history for both of them, and this time you’ve got a One Nation candidate coming to the field, if you operate on swing, you’ve got two figures which are matched and one which is a raw number from zip.

Linda (18:59)
Mm.

Antony (19:11)
Because they weren’t there last time. And that one will have a different variance pattern than the two where you’ve got history. And if and and technically if you operate operate on swing and you apply to the swing from last time, your predicted numbers don’t on primaries don’t add to a hundred percent. And that’s just simply because of the basis of the numbers, it settles down for the evening. When you’ve got the same pairing of candidates, the swings always match each other on either side and they add to a hundred percent.

But when you’ve got four and five candidates on primaries, this method of using swings and projecting doesn’t always guarantee the projected percentages add to a hundred percent. we also we we some of the numbers are reliable and some of them are not. It’s just you just have to know what you’re talking about and try and explain it to the viewers if you’re doing it on television. So it’s so just to clarify that. So if I’m coming up to the Victorian election, I can be relatively confident that if I talk about the Labor vote statewide corrected for swing using swing.

Linda (19:54)
Yeah.

So that you talked about

Antony (20:09)
I mean relatively confident that that’s correct, but the One Nation figure will be all over the place ’cause I don’t have any history for it. so the statewide figures won’t add to a hundred percent. But that’s a modelling issue. It’s it’s not an error. It’s just you have to be aware of the limitations of the day you’re looking at.

Linda (20:24)
Yeah, you have to know how the system works. you’ve mentioned polling a few times as sort of the opinion polls in in advance that give you some idea of how it’s going to to shape up. What we’ve seen it was really obvious in 2016 in the US that the polls were predicting a completely different outcome to the one we got. Do you feel like opinion polls have become less useful or less reliable?

Antony (20:59)
Well in the case of 2016 in the United States, the issue was people just tend to take the statewide poll, but the election is decided by the electoral college, by who wins each state and gets all the votes from that state. And there’s a series of very close calls at that election. I’m trying to remember the guy who does all this stuff in America. If if you did a modelling exercise where you predicted this turn the statewide poll into predictions in each state and then projected the who would win each state and what the electoral college would be.

If you did modeling, it showed that I think Hillary Clinton had a probability of about seventy percent of winning. now, if you use statistical probability, seventy percent’s not good enough. If you are doing something like gambling and you’re p and you have something that says 70% chance that of knowing that this number will pop up rather than fifty fifty, and it’s a repeatable event, seventy percent’s good enough. If you’re gambling, you can bet on a seventy percent turnout that most of the time you’ll get your money back if you keep betting on that.

In the old the old argument about coin tossing, you know, what’s if you cross toss got eight heads in a row, what’s the probability of the next one being a head? Well fifty fifty. It just it’s a sample space you’re talking about. It’s but if you’ve got a seventy, seventy thirty split you can get bet on that. But if you’re talking event which is not repeatable, like a presidential election, seventy percent’s not good enough. You need to get into a much higher probability. And so even in 2019 in Australia when the polls appeared to get it wrong.

Yes, they did. They’re on the I think they’re all about fifty one and a half label and it turned out to be about fifty one and a half label. that’s just on the edge of the confidence interval, the probability ninety nine ninety five percent probability s range. So th even that wasn’t terrible. but the fact that all of them were wrong is a great suspicion of a lot of herding going on there because opinion polls are not random samples, they’re weighted samples.

What they tend to report is the error margin for traditionally they tend to report a probability margin or an error margin for a real random sample. But they’re not random samples, they’re weighted samples. And when you apply weights, you in fact you actually increase the error margin. And nowadays they sometimes report that as well.

Linda (23:18)
How are they weighted samples?

Antony (23:24)
Because they they have difficulty getting, traditionally it was age, which is always a big variable. But when they do their sample and they look at the types of voters in the sample and reflect the national vote, like university education has always been used as a a proxy for education level. There’s a number of things like that they do to try and get the sample accurate.

Linda (23:44)
So they they try to, try to change the, the results that they have in the opinion poll to reflect better reflect the population of the country as a whole rather than the population of the sample.

Antony (23:58)
Yeah, when, when, when we did phone calls to houses twenty years ago, they did them all on Thursday, Friday, Saturday. it’s rather hard to they they would ask questions to randomize who answered the phone. Like who answers a phone in a house? There’s a bias on who answers a phone when it rings in house. So what they would do they would ask for the person whose birthday was most recent in the household.

Linda (24:03)
Mm.

Antony (24:27)
Or if, if they were doing stratified samples, you know, you’d you’d always fill up the sample of voters over the age of sixty when you were polling. It was very hard to get the po voters under thirty. So as as they were doing their sampling, there’d come a point where they stopped taking anyone over the age of sixty anymore. Because they’d already filled enough of them. So there’s lots of things they used to do to try and get the sample good good that way.

Linda (24:27)
Mm.

Antony (24:53)
But there was always the issue with phone polling, they were underestimating young people and they were getting the sort of young people who are home on a Friday and Saturday evening rather than the ones who are going out raving. So I mean it’s it’s there there’s things like that. And nowadays there’s there’s very few of them do phone polling. Most of them are operating off samples. there’s big samples of people who do surveys and they do random samples from this big big sample and then they compare it with the population.

Linda (24:57)
Yeah.

Antony (25:20)
But they have got better but, you know, there’s still a bit of a fudge factor I think in some of these polls and if someone produces a poll which is substantially different from everyone else, someone goes and looks at their sample and goes, look, I think we’ve got too many young people in this sample or Yeah, there’s there’s they do their best. But I mean they’re just they’re just guides. I think they’re useful. They give you a rough sample. But I mean to say you get an opinion poll which in the last federal election was interesting. There were plenty of polls showing the Albanese government would lose its majority at the start of the year.

There were very few that ever showed the coalition would win. But the ones which showed the coalition ahead were being constantly reinterpreted as the coalition would win. And I don’t think that that wasn’t as it proved a correct assessment. In the end, the polls did shift, the public mood shifted as we got closer to the election. So the polls at the start of the year turned out not to be a good guide for what happened on election day. the polls the polls even as late as election the last week of the election still underestimated the actual results.

There’s one pollster that got closer and I think there’s I think there’s some suspicion that some of them just you know, pulled it back a bit.

Linda (26:28)
Yeah. Is there is there something that you think would it would be better or or the world would change if if more people understood this thing about about voting data, or is there one thing that you think, you know, a lot of people misunderstand about voting data that that is problematic?

Antony (26:42)
When you told that yeah,

We have a preferential voting system. But you know it was a constant whinge after the election. How did Labor get a landslide victory with only thirty six percent of the vote? Well, it’s ’cause we have a preferential system. And that’s a state that’s a nationwide vote, it’s thirty, thirty five, thirty six percent nationwide. Labor had two advantages. One we’re a preferential system, they got more preferences, way more. the second one is that Labour’s vote was more evenly was better distributed.

Labor polls very poorly in country electorates it doesn’t win. Its vote evaporated in seats it didn’t win or where it went finished third or lower and was excluded. But it got a good vote in all the seats it win. So it was very good, a very effective at turning its vote. I mean even even though I mentioned that opinion poll in Victoria which said Labor and Liberal on twenty-six and One Nation on twenty-seven, there was also thirteen percent with the Greens. We have a preferential voting system.

So Labor on twenty-six during the distribution of preferences statewide it turns out to be more in the low to mid thirties because the Green vote’s there. There’ll be some seats where the Greens will finish in the top two. But essentially, I think in a number of opinion polls at the moment, there’s a reporting of the primary votes, which is good, and the the two party preferreds being estimated are rubbish. But when you see an opinion poll, eighty five percent of Green preferences flow to Labor. We we know that election after election.

So if you add the Labor and the Green vote together, you get a something which is a bit more reliable. If you had the Liberal and and One Nation vote together, Coalition and One Nation vote together, they have preferences for about seventy to seventy five percent. So you can’t add them together, you have to discount the fact that there’s there’s preferences that leak. So I think there’s I don’t think the two party preferred that they used to report is not very useful these days. But if you remember how preferences work, that you can get a bit more information about a poll in the primaries and I think people should do a little more of that, adding the Labor and Green vote together to see what that looks like and it’s a bit more reliable about what the opinion polls are saying.

Linda (28:56)
That makes sense.

It’s it’s changed a lot just in the last sort of ten to twenty years. Do you think it’s gonna change more?

Antony (28:59)
Well it depends on the next next two years, I suspect. I mean we’ve seen this real rise of One Nation. Is that a temporary phenomenon? I mean, a Victorian election’s interesting. If, if they end up with no party having a clear majority and the Liberal Party end forms government as the one most likely to be able to govern in that parliament and has to rely on One Nation, the way One Nation behave in the situation where they hold the balance of power for government for four years in a parliament.

That would impact how how people view they can be trusted at the next federal election, for instance. because there’s nothing that makes exposes a party’s ability than actually getting an actual handle on on power. It shows whether they’re trustworthy or not. so that will be interesting to see.

At the moment they’re an anti anti incumbent party. they’re against things. They’re just they’re they’re they’re feeding on people’s resentment that people feel they’re working hard and they’re not getting anywhere and there’s a real resentment going in. Some of this flows on from COVID and distrust of government and a lot of conspiracy theory stuff coming out of the States. That’s all feeding in there. we saw at last year’s federal election that, you know, the Trump stuff seemed to be working for the Liberal Party at the end of twenty twenty four. But then once Trump was in power again and was putting tariffs on left, right and centre was doing all sorts of strange things, it didn’t quite look as safe. And people who are, parties that were close to Trump seemed to not do as well. So I think it’s

It’s an interesting exercise how much American politics is influencing us.

And, Victoria will be a big test. It’s the most consequential state election in my lifetime for its implications for federal politics. Because if One Nation do do well, they get somewhere, you know, they play a role in government. If they’re incapable of doing that sensibly, then that will reflect how they’re viewed for future elections. Because remember, the party does have a bad record of keeping its MPs in line in Parliament. It’s lost a lot of its MPs to the crossbench, a lot of them have been kicked out of the party. If the One Nation gets part of the control of the Victorian Parliament and it falls apart, that’s a bad sign for the party in the future.

Linda (31:02)
Yeah. Just

Yep. Speaking as a Victorian, I’d rather not go through that pr that finding out process.

Antony (31:24)
Can I also say we’ve been talking about data? I mean, data shouldn’t affect how people vote. People should when they vote, number the ballot paper in the order you would like to see the candidate selected. That’s voting. I think one of the biggest problems is peop people trying to explain the voting system as a way of trying to explain how to vote, and that’s not the way to do it. You vote, it’s a preferred voting system. You list the you vote for your preferred candidate, your second preferred candidate, your third preferred candidate. That’s what you do.

Linda (31:29)
Yeah.

Yep. Yep.

Antony (31:53)
And and if that if everybody does that the counting system works fine.

It’s it’s it’s like you haven’t listed the candidates in a preferred order, you’ve listed them in a strategic order to engineer a result. And of course you’re only one of a hundred thousand votes, so it’s waste of time. Don’t vote strategically.

Linda (32:10)
So that’s that’s the the message. That’ll be the the key talking point for this podcast. It’s a waste of time. Don’t vote strategically.

Antony (32:18)
Yeah, I mean you get you get some arguments in pla things like Hare Clark, the proportional representation system about what’s the most effective way to have your vote count. But I mean people who go in there trying to fill in the ballot paper so the vote the ballot paper goes from candidate to candidate to candidate. wasting your time. You can’t pick that. You just just vote for the candidates you want elected.

Linda (32:37)
Yep. That’s solid advice. I think it’s nicely ethical advice too. Are are there misleading stories that that the data tells in in elections? Are there things that you’ve sort of that the data seems to say that that don’t bear out or that change?

Antony (33:02)
I’m not I’m not sure you need anything something specific or-

Linda (33:06)
No, I was just I I sometimes in my workshops I teach that the story the data tells you isn’t or th the story the data seems to tell isn’t necessarily accurate and you have to keep sort of testing and exploring to see if you’re actually getting the right idea from the data. We we tend to assume that that you know, what’s in the data is clear and obvious and and makes sense. And I just wonder whether there are times when what the data seems to be saying is is confusing or misleading in some way.

Antony (33:40)
Well, it’s another example, I mean, voting is easier than polling. doing election analysis is easier polling because your polling places, your results on the night, you treat them as a sample, but the sample keeps getting bigger. so the question is, I mean, you know, if you plot a result in electorate, a two party preferred line, it might jiggle around and then settle down, ’cause it’s a progressive count. You know, as your s as your sample gets bigger, it settled down, glides towards the final figure. an electorate which say is mixed country rural electorate will start if you just plot the same graph of labor vote it will tend to start low and then rise and then go to the level. At what point does it settle down? we did something like the referendum. We knew that the referendum would start with a low, the voice referendum would start with a low yes vote and it would rise. So we were plotting it on the night visually, because we didn’t have real good statistical tools. We could plot we the minute it started to level off you knew knew that was it.

Linda (34:35)
Mm.

Antony (34:37)
While I say that there is one change nowadays, is again this thing with pre poll voting, they come in much later and they may behave differently. And as big data spots, they can sort of shift the graph away from the initial trend. because they they they are a much big lumpier data that comes. And the other thing is, while I I draw a nice plot, I always remind people that if you p plot the vote received against the time, it comes as a something that most like the a curve which starts off slowly, then speeds up quicker and then s and levels off. So it’s an S shape. An extended S shape. Okay. and that’s means that all though all that jiggling around with early figures is occurring over a long period of time when you’re trying to figure out what’s going on. so often I draw the graph of of of results on the night, I draw it once with as against the percentage counted and it looks smoother. And then you draw against time and of course there’s all that jiggling at the start extends over a longer period and then the rest of it is flat. So it’s it’s depends on how you look at the data also.

Linda (35:48)
Thank you so much. This has been a fantastic conversation. I really appreciate you coming back on and sharing your expertise. I think hopefully this will help people understand the voting system better and and get some inspiration for for what data can tell us and how we can understand it.

Antony (36:07)
Yeah, it’s good. I mean it’s there is more data about Australian elections published than most other countries in the world. And yeah, I the sheer volume of data which is published down to the polling level the polling place level is enormous. that’s interesting. And even now, say with the Senate system, the ballot papers are all scanned. So there’s actually ballot paper data you can look at nowadays and look at you can compare it with things like how to vote, did people follow the how to vote and stuff?

Linda (36:15)
Really? I didn’t know that.

Yeah.

Antony (36:36)
That’s that’s something which isn’t a lot of work done on. But it’s interesting. It seems that most people I always say from my observation and looking at the data, when people put in preferences, it’s the first, second, third preferences is the most important. If they vote for a minor party, they’re likely to go to their eventual destination with a major party at second or third. So that’s the most important. There are very few people who go all round the ballot paper before put before they say put the last two major parties last. So it it

Linda (36:56)
Yeah.

Yeah.

Antony (37:05)
It’s it’s always the top of the ballot paper, what people do the first, second, and third. And that’s got more important as the number of candidates. So if you’ve got eight, nine, ten candidates, there’s a point where people will have three or four preferences in mind, but the rest of them are just like random. They’ll often just end up drifting down the ballot paper. They’ll go one, two, three, four, and then go five, six, seven, eight, nine, ten. You know, fill in the rest of the numbers. And you can see that in the distribution of preferences. You can always see this slow drift towards the top of the ballot paper.

Linda (37:25)
Yeah.

Antony (37:32)
And when the candidate in number one gets excluded, there’s always this drift down the ballot paper. you can always what they call a donkey vote. You can always see it. It’s not just a donkey vote in terms of the candidate votes one at the top and straight down, which is the classic donkey vote. We’re incre increasingly seeing a preferential donkey vote, which is once people have filled in the numbers they knew, they know, they tend to go to the top of the ballot paper and just number down, which means then you can sometimes see this drift down the ballot paper.

Linda (37:37)
That’s really cool.

Mm.

That’s really interesting.

It the the amount of data that that we have, I I just assumed that was normal. I didn’t realise it was unusual, but I actually used it for a year ten project with my students back in twenty sixteen. It was one of the things that triggered the the foundation of my organization actually, where we just downloaded the Senate votes for Victoria. And so we had this three million plus line CSV file and every line in the file was a vote and you could see what happened in every box on that ballot paper and that’s just magic! I love you know the level of exploration you can do with that kind of data is fascinating.

Antony (38:35)
Yeah. And then there’s only has the formal preferences, the ones which it becomes informal aren’t included. And the second thing is because it’s got both eight above and below the line ballot papers, if you see one with an above and below the line option, you have to actually sieve it first to figure out which one’s formal. I cause I’ve written code to do it and it’s an absolute pain trying to do it.

Linda (38:56)
Yes.

Yes.

Antony (39:07)
Yeah, it’s, it’s, it’s, just I’ve I’ve got some lovely tables of the Senate preference flows and there’s something wrong somewhere in my data. I think it’s an error of about one percent, which doesn’t doesn’t invalidate my conclusions. But it does make it very difficult for me to use precise numbers when I’m down to sort of looking how many ballot papers ended up with the last two candidates. You know, it’s it’s it’s it’s it’s a bit annoying, but when you’ve got three million records, it’s damn difficult to know whether your code’s getting every one of them right or not.

Linda (39:33)
Yeah, yeah. It’s that’s one of the challenges of working with that scale of data is, you know, how do you how do you verify? How do you you sense check? How do you spot or detect where things have gone wrong? It’s really difficult.

Antony (39:47)
I’ll I’ll finish with one anecdote. There’s a very embarrassing incident in the two thousand and three New South Wales election where for for reasons which you won’t go into, they when they do a distribution of someone who’s got more than a quota of votes, they have to random sample the the ballot papers to determine preferences. Every other state counts them out, but in law New South Wales says you have to random sample. So they had to program the computer to random sample.

Linda (40:11)
Wow.

Antony (40:16)
So the commissioner was afraid to test the count on the actual data because it didn’t necessarily guarantee you get the same result anyway. So he didn’t test it. But then there was a problem where in the transfer of votes from the data entry system to the counting system. This is one of these examples of you don’t think about what people write on the ballot. People putting ticks and crosses and stuff and some of them are formal, some are informal.

People were just zero, what do you put do with a zero? And anyway, they’ve got some terrible problem, they got a muddle between spaces, zeros and nulls. Anyone who’s worked computers knows what happens when you get a null in a data file. It just just stuffs you up. Anyway, so when all the data was transferred from the data entry system to the to the counting system and they hit the button to the distribution of preferences, it stopped. And it was this null issue. It’s just it’s just just felt very sorry for them. but-

Linda (40:49)
no.

Antony (41:09)
-if you’ve ever done any data work, it’s very annoying when there’s something in your data and you can’t find where it’s wrong and you spend forever chasing it and so I’ve hit that with a Senate data problem. I I think I’ve got it mostly right but there’s something which isn’t quite right but I can’t find it.

Linda (41:23)
Well, maybe it’s not yours, that’s not quite right. Maybe it’s the AEC.

Antony (41:29)
Yes. All right. It was good fun.

Linda (41:31)
Thank you so much, that was a fantastic conversation.

Antony (41:34)
Thanks very much.

Leave a Reply