Posts written by Thomas Lumley (2645)

avatar

Thomas Lumley (@tslumley) is Professor of Biostatistics at the University of Auckland. His research interests include semiparametric models, survey sampling, statistical computing, foundations of statistics, and whatever methodological problems his medical collaborators come up with. He also blogs at Biased and Inefficient

October 29, 2016

Uncertainty and symmetry in the US elections

Nate Silver’s predictions at 538 give Donald Trump a much higher chance of winning the election than anyone else’s: at the time of writing, 20% vs 8% from the Upshot, 5% from Daily Kos, or 1% from Sam Wang at the Princeton Election Consortium.

That’s mostly not because Nate Silver thinks Trump is doing much better: 538 estimates 326 Electoral College votes for Clinton; Daily Kos has 334; the Princeton folks have 335.  The popular vote margin is estimated as 5.7% by 538 and about 8.4% by Princeton (their ‘meta-margin’ is 4.2%).

Everyone also pretty much agrees that the uncertainty in the votes is symmetric: if the polls are wrong, the estimated support for Clinton could as easily be too high as too low.  But that’s the uncertainty in the margin, not in the chance of winning.  Probabilities can’t go above 100% or below 0%, and when they get close to these limits, a symmetric uncertainty in the vote margin has to turn into an asymmetric uncertainty in the probability prediction, and a larger uncertainty has to pull the probability further away from the boundaries.

Nate Silver’s model thinks that opinion polls can be off by 6 or 7 percent in either direction even this close to the elections; the others don’t. It’s question that history can’t definitively answer, because there isn’t enough  history to work with. If Silver is wrong, we won’t know even after the election; even if he’s right, the most likely outcome is for the results to look pretty much like everyone predicts.

October 28, 2016

False positives

Before a medical diagnostic test is introduced, it is supposed to be evaluated carefully for accuracy. In particular, if the test is going to be used on the whole population, it’s important to know the false positive rate: of the people who test positive, what proportion really have a problem?  Part of  this process is to make sure that the test works as a biological or chemical assay: is it accurately measuring, say, carbon monoxide or glucose in the blood.  But that’s only part of the process.  You also need to worry about what threshold to use — how high is ‘high’ — and whether people could have high carbon monoxide levels without being smokers, or high glucose levels without being diabetic.

I haven’t heard any suggestion that the tests for methamphetamine contamination in houses fail the first step. There’s meth present when they find it. But Housing NZ were treating the high assay value as evidence that (a) the house was dangerous to live in, and (b) that the tenant was responsible. The false positive rates for (a) and (b) were not established, and appear to be shockingly high given the consequences.

The Ministry of Health has now released new guidelines on meth contamination, with concentration thresholds based on evidence (though towards the low end of what their evidence would support).  They claim to have repeatedly warned Housing NZ. Russell Brown has an excellent summary of the situation at Public Address.

While this is all a step forward, it’s not addressing the question of (b) above: if there’s methamphetamine present at above the new action threshold, it appears that this is still going to be taken as evidence of the tenant’s culpability. That would only make sense if, contrary to the advertising from the meth-testing companies, low-level meth contamination were very rare in rented NZ houses.

October 25, 2016

Oversampling

From the election on the other side of the Pacific.

Wikileaks also shows how John Podesta rigged the polls by oversampling democrats, a voter suppression technique.

Now, as Josh Marshall at Talking Points Memo goes on to point out, the email in question is not to or from John Podesta, is eight years old, and refers to the Democrats internal polls not to public polls. So it’s kind of uninteresting. Except to me.  I’m a professional sampling nerd. I do research on oversampling; I publish papers on it; I write software about it: ways to do it and ways to correct for it. And just like a sailing nerd who has heard Bermuda rigging described as a threat to democracy, I’m going to explain more than you ever needed to know about oversampling.

The most basic form of oversampling in medical research has been widely used for over sixty years. If you want to study whether, say, smoking causes lung cancer, it’s very inefficient to take a representative sample of the population because most people, fortunately, don’t have lung cancer. You need to sample maybe 1000 people to get two people with lung cancer. If you have access to hospital records you could find maybe 200 people with lung cancer and 800 healthy control people.   Your case-control sample would have about the same cost as a representative sample of 1000 people, but nearly 100 times more information.  And there are more complex versions of the same idea.

Your case-control sample isn’t representative, but you can still learn things from it.  At a simple level, if the lung cancer cases are more likely to smoke than the controls in the sample, that will also be true in the population. The relationship won’t be the same as in the population, but it will be in the same direction.  For more detailed analysis we can undo the oversampling. Suppose we want to estimate the proportion of smokers in the population. The proportion of smokers in the sample is going to be too high, because we’ve oversampled lung-cancer patients, who are more likely to smoke. To be precise, we’ve got one hundred times too many lung cancer patients in the sample. We can fix that by giving each of them one hundred times less weight in estimating the population total. If 180 of the 200 lung cancer patients smoked, and 100 of the 800 controls did, you’d have a weighted numerator of 180×(1/100)+100×1, and a weighted denominator of 200×(1/100)+800×1, for an unbiased estimate of  12.7%, compared to the unweighted, biased (180+200)/1000 = 38%.

In polling, your question might be what issues are important to swing voters. You’d try to oversample swing voters to ask them, and not waste time and money annoying people whose minds were made up.  Obviously that would make your sample un-representative of the whole population. That’s the point; you want to talk to swing voters, not to a representative sample.  Or you might want to compare the thinking of (generally pro-trump) evangelical Christians and (often anti-Trump) Mormons. Again, if you oversampled conservative religious groups you’d end up with an unrepresentative sample; again, that would be the point. Oversampling isn’t the best strategy when your primary purpose is finding out what a representative sample thinks; it often is the best strategy when you want to know more about some smaller group of people.

However, if you also wanted an estimate of the overall popular vote you could easily undo the oversampling and downweight the swing voters in your sample to get an unbiased estimate as we did with the smoking rates.  You have to do that anyway;  even if you try to get a representative sample it probably won’t work because some groups of people are less likely to answer their phones and agree to talk to you.  The weighting you use to fix up accidental over- and under- sampling is exactly the same as the weighting you use when it’s deliberate.

 

October 24, 2016

Why so negative?

My StatsChat posts, and especially the ‘Briefly’ links, tend to be pretty negative about big data and algorithmic decision-making. I’m a statistician, and I work with large-scale personal genomic data, so you’d expect me to be more positive. This post is about why.

The phrase “devil’s advocate” has come to mean a guy on the internet arguing insincerely, or pretending to argue insincerely, just for the sake of being a dick. That’s not what it once meant. In the early eighteenth century, Pope Clement XI created the position of “Promoter of the Faith” to provide a skeptical examination of cases for sainthood. By the time a case for sainthood got to the Vatican, there would be a lot of support behind it, and one wouldn’t have to be too cynical to suspect there had been a bit of polishing of the evidence. The idea was to have someone whose actual job it was to ask the awkward questions — “devil’s advocate” was the nickname.  Most non-Catholics and many Catholics would argue that the position obviously didn’t achieve what it aimed to do, but the idea was important.

In the research world, statisticians are often regarded this way. We’re seen as killjoys: people who look at your study and find ways to undermine your conclusions. And we do. In principle you could imagine statisticians looking at a study and explaining why the results were much stronger than the investigators thought, but since people are really good at finding favourable interpretations without help, that doesn’t happen so much.

Machine learning includes some spectacular achievements, and has huge potential for improving our lives. It also has a lot of built-in support both because it scales well to making a few people very rich, and because it fits in with the human desire to know things about the world and about other people.

It’s important to consider the risks and harms of algorithmic decision making as well as the very real benefits. And it’s important that this isn’t left to people who can be dismissed as not understanding the technical issues.  That’s why Cathy O’Neil’s book Weapons of Math Destruction is important, and on a much smaller scale it’s why you’ll keep seeing stories about privacy or algorithmic prejudice here on StatsChat. As Section 162 (4) (a) (v) of the Education Act indicates, it’s my actual job.

 

Briefly

  • I would never have guessed this was a problem, but “Data from three national surveys indicated that people are unaware that age is a risk factor for cancer. Moreover, those who were least aware perceived the highest risk of cancer regardless of age.” (free abstract but paywalled paper, via @RolfDegen)
  • Useful graph of uncertainty in vote margin and winner from Nate Silver on Twitter.
    us-uncertainty
  • There’s a computer-personalised education system supported by Facebook that seems to be getting good results. On the other hand, the evidence for the effectiveness isn’t very good quality, and the handling of data privacy is weak. There’s going to be a lot of this sort of issue coming up in the data-based policy world. (Washington Post)
October 23, 2016

Psychic meerkats and Halloween masks

Prediction is hard — especially,  as the Danish proverb says, when it comes to the future. In the Rugby World Cup we had psychic meerkats. For the US elections the new bogus prediction trend is Halloween masks: allegedly, more masks are sold with the face of the candidate who goes on to win.

The first question with a claim like this one, especially given some of the people making it, is whether the historical claim is true.  In this case it’s true-ish.  The claim was made before the 2012 election, and while the data aren’t comprehensive, they are from the same big chain of stores each year. From 1980 to 2012, the mask rule has predicted the eventual winner of the presidency.  That’s actually an argument against it.

If there’s more to the mask sales than there is to psychic meerkats, it would have to be as a prediction of the popular vote — you’d need data from individual states to predict the weird US Electoral College. But if the mask rule got the 2000 election right, it must have got the popular vote wrong that year — George W. Bush won the electoral college, but lost the popular vote to Al Gore. From that point of view, we’re looking at 8 out of 9.

More importantly, 9 out of 9 isn’t all that impressive. Suppose you got your predictions by flipping a coin.  Your chance of getting either all heads for the Republican wins or all heads for the Democratic wins is 1 in 256, increasing to 1 in 128 if you’re allowed to choose which way to treat the 2000 election.  The chance of getting 8 of 9 agreement is much better: about 1 in 13.  If only one in a million people in the US had tried coming up with just one prediction rule each, you’d expect someone to get it perfect and dozens to get it nearly right.

Given these odds, it wouldn’t be surprising if, say, a US professional sports team had results agreeing with the Presidential results — and in fact, there was a rule based on the results for the Washington Redskins football team that worked from 1940 to 2000, was fudged to work in 2004, and then failed completely in 2012.    That’s 17/19 correct, but since the rule was first publicised in the run-up to the 2000 election, it’s 2/4 correct in actual use.

If you’re allowed to combine multiple variables it gets even easier to find rules. With anything from basic linear regression to a neural network you’d expect to get perfect prediction from five unrelated variables. Even restricting the models to be simple doesn’t help much.  I downloaded some OECD data on national GDP for various countries, and found that since 1980 the Republicans have won the popular vote precisely in years when the GDP of Sweden increased more than the GDP of Norway.

My advice is to stick with the psychic meerkats for entertainment and the opinion poll aggregators or the betting markets for prediction.

October 22, 2016

Stat of the Week fixed

Because of changes at WordPress, the Stat of the Week competition has been eating the URLs you submitted.

Um.

Sorry.

 

We’ve fixed it now.

Cheese addiction hoax again

Three more sites have fallen for the cheese addiction hoax

As you may remember, this story is very very loosely based on real research from the University of Michigan. However, the hoax version misrepresents which foods were most addictive and makes up an explanation based on the milk protein casein that isn’t mentioned in the real research at all.

The reason I’m calling this a hoax is that it wasn’t the fault of the researchers, their institution, or the journal, and it’s obvious to anyone who makes any attempt to scan the research paper that it doesn’t support the story. It isn’t an innocent mistake, and it isn’t a simple exaggeration like most misleading health science stories.

There’s a good post at Science News describing what was actually found.

October 20, 2016

Brute force and ignorance

At a conference earlier this week, a research team from Microsoft described a computer system for speech transcription. For the first time ever, this system did better than humans on a standard set of recordings.

What’s more impressive — and StatsChat relevant — is that this computer system does not understand anything about the conversations it writes down. The system does not know English, or any other human language, even in the sense that Siri does.

It has some preconceived notions about what tends to follow a particular word, pair of words, or triple of words, and about what sequences of sounds tend to follow each other, but nothing about nouns or verbs or how colorless green ideas sleep. As with modern image recognition, the system is just based on heaps and heaps of data and powerful computers.  It’s computing and statistics, not linguistics.

In a comment to a post at Language Log, the linguist Geoffrey Pullum says

I must confess that I never thought I would see this day. In the 1980s, I judged fully automated recognition of connected speech (listening to connected conversational speech and writing down accurately what was said) to be too difficult for machines, far more difficult than syntactic and semantic processing (taking an error-free written sentence as input, recognizing which sentence it was, analysing it into its structural parts, and using them to figure out its literal meaning). I thought the former would never be accomplished without reliance on the latter.

There are many problems where enough data is not available to construct a model with no understanding of the problem. There won’t be a shortage of work for human statisticians or linguists any time soon. But there are problems where brute force and ignorance works, and they aren’t always the ones we expect.

October 18, 2016

Evidence-based policy chants

An old one, seen at the ‘Rally To Restore Sanity and/or Fear”

What do we want?
EVIDENCE-BASED CHANGE!

When do we want it?
AFTER PEER REVIEW!

 

A new one, from @zentree and @bex_stevenson on Twitter

What do we want?
RELIABLE NUMBERS!

When do we want them?
STAT!

(this sort of thing is why we have a ‘Silly’ tag on StatsChat)