Posts written by Thomas Lumley (2645)

avatar

Thomas Lumley (@tslumley) is Professor of Biostatistics at the University of Auckland. His research interests include semiparametric models, survey sampling, statistical computing, foundations of statistics, and whatever methodological problems his medical collaborators come up with. He also blogs at Biased and Inefficient

July 17, 2012

You’re all individuals

The Herald  is at least showing some scepticism about Italian-style patisserie that is supposed to make you lose weight (they include green tea and guarana, ie, caffeine).  The manufacturer isn’t willing to give any numbers

But she said it was not possible to measure how much eating the treats would help boost the metabolism because ‘everyone is different’.

Of course, this is a pretty transparent excuse.  If the fact that everyone is different made measurements impossible, medical science would be in a bad way.  We can measure the average effect.  We can measure the variability in the effect.  We can measure the proportion of people helped.  And we do all these things.  For example,  we’ve known for years that angiotensin converting enzyme inhibitors reduce blood pressure on average by about 10mmHg.  More recently, some Sydney researchers reanalyzed the data from the randomized trials to look at how much person-to-person variation in effect there was, and found it was extremely small.

The story goes on to say

Registered public health nutritionist Charlotte Stirling-Reed said that consumers should always look for evidence before making purchases based on health claims.

True, but that would spoil all the fun.

Margin of error yet again

In my last post I more-or-less assumed that the design of the opinion polls was handed down on tablets of stone.  Of course, if you really need more accuracy for month-to-month differences, you can get it.   The Household Labour Force Survey gives us the official estimates of unemployment rate.  We need to be able to measure changes in unemployment that are much smaller than a few percentage points, so StatsNZ doesn’t just use independent random samples of 1000 people.

The HLFS sample contains about 15,000 private households and about 30,000 individuals each quarter. We sample households on a statistically representative basis from areas throughout New Zealand, and obtain information for each member of the household. The sample is stratified by geographic region, urban and rural areas, ethnic density, and socio-economic characteristics. 

Households stay in the survey for two years. Each quarter, one-eighth of the households in the sample are rotated out and replaced by a new set of households. Therefore, up to seven-eighths of the same people are surveyed in adjacent quarters. This overlap improves the reliability of quarterly change estimates.

That is, StatsNZ uses a much larger sample, which reduces the sampling error at any single time point, and samples the same households more than once, which reduces the sampling error when estimating changes over time.   The example they give on that web page shows that the margin of error  for annual change in the employment rate is on the order of 1 percentage point.  StatsNZ calculates sampling errors for all the employment numbers they publish, but I can’t find where they publish the sampling errors.

[Update: as has just been pointed out to me, StatsNZ publish the sampling errors at the bottom of each column of the Excel version of their table,  for all the tables that aren’t seasonally adjusted]

July 16, 2012

When a dog bites a man, that’s not news

A question on my recent post about political opinion polls asks

– at what point does the trend become relevant?

– and how do you calculate the margin of error between two polls?

Those are good questions, and the reply was getting long enough that I decided to promote it to a post of its own. The issue is that proportions will fluctuate up and down slightly from poll to poll even if nothing is changing, and we want to distinguish this from real changes in voter attitudes — otherwise there will be a different finding every month and it will look as if public opinion is bouncing around all over the place.  I don’t think you want to base a headline on a difference that’s much below the margin of error, though reporting the differences is fine if you don’t think people can find the press release on their own.

The (maximum) margin of error, which reputable polls usually quote, gives an estimate of uncertainty that’s designed to be fairly conservative. If the poll is well-designed and well-conducted, the difference between the poll estimate and the truth will be less than the maximum margin of error 95% of the time for true proportions near one-half, and more often than 95% for smaller proportions.  The difference will be less than half the margin of error about two-thirds of the time, so being less conservative doesn’t let you shrink the margin very much.   In this case the difference was well under half the margin of error.  In fact, if there were no changes in public opinion you would still see month-to-month differences this big about half the time.

For trends based on just two polls, the margin of error is larger than for a single poll, because it could happen by chance that one poll was a bit too low and the other was a bit too high: the difference between the two polls can easily be larger than the difference between either poll and the truth.

The best way to overcome the random fluctuations to pick up small trends is to do some sort of averaging of polls, either over time, or over competing polling organisations.  In the US, the website fivethirtyeight.com combines all the published polls to get estimates and probabilities of winning the election, and they do very well in short-term predictions.  Here’s a plot for Australian (2007) elections, by Simon Jackman, of  Stanford, where you can see individual poll results (with large fluctuations) around the average curve (which has much smaller uncertainties).  KiwiPollGuy  has apparently done something similar for NZ elections (though I’d be happier if their identity or their methodology was public).

So, how are these numbers computed?  If the poll was a uniform random sample of N people, and the true proportion was P, the margin of error would be 2 * square root(P*(1-P)/N).  The problem then is that we don’t know P — that’s why we’re doing the poll. The maximum margin of error takes P=0.5, which gives the largest margin of error, and one that’s pretty reasonable for a range of P from, say, 15% to 85%. The formula then simplifies to 1/square root of N.   If N is 1000, that’s 3.16%, for N=948 as in the previous post, it is 3.24%.

Why is it  2 * square root(P*(1-P)/N)?  Well, that takes more maths than I’m willing to type in this format so I’m just going to mutter “Bernoulli” at you and refer you to Wikipedia.

For trends based on two polls, as opposed to single polls, it turns out that the squared uncertainties add, so the square of the margin of error for the difference is twice the square of the margin of error for a single poll.  Converting back to actual percentages, that means the margin of error for a difference based on two polls is 1.4 times large than for a single poll.

In reality, the margins of error computed this way are an underestimate, because of non-response and other imperfections in the sampling, but they don’t do too badly.

July 14, 2012

BBC radio equivalent of StatsChat

If you don’t like StatsChat you will probably not enjoy “More or Less”,  a BBC radio show/podcast/blog on statistics in the British media, presented by Tim Harford.  Their new season starts this week.

Poll shows not much

According to the Herald

The latest Roy Morgan Poll shows support for the National Party has fallen two per cent since early June.

 The poll is based on 948 people, so the maximum margin of error (which is a good approximation for numbers near 50%) is about 3.2%, and the margin of error for a change between two polls is about 1.4 times larger: 4.6%.

July 13, 2012

When randomized trials don’t help

The Herald reports on a study of weight gain after quitting smoking, which is based on analyzing the results of 62 randomized trials of treatments to help quitting.  Ordinarily, data from randomized trials is what we want, so why is Professor Simon Chapman quoted as complaining the results are unreliable?

Well, it’s partly because he doesn’t like anything that might be construed as favorable about smoking, but he has a good point in this case.  Randomized trials give us fair and trustworthy comparisons between two treatments: in this case the 62 trials tell us something about which ways to quit actually lead to the most quitting.   The information on weight gain, on the other hand, isn’t a comparison of two randomized treatments, it’s a before-after comparison on the people who managed to quit.  The fact that the treatments were randomized is of no help at all, since the analysis lumps all quitters together.

In fact, it’s worse than that. A lot of smokers manage to quit with only moderate difficulty.  They tend not to end up in randomized trials.  Some smokers find quitting much harder, and they are much more likely to end up in randomized trials.  So the research actually has found that a group of people who probably found quitting hard have gained 4-5kg after quitting.   It’s quite likely that people who find quitting hard are also going to gain more weight than those who quit without major difficulties, so we may well be overestimating the impact of quitting on weight.

 

Our new robot overlords

Since I regularly complain about the lack of randomised trials in education, I really have to mention a recent US study.  At six public universities in the US, introductory statistics students who consented were randomised between the usual sort of teaching by real live instructors or a format with one hour per week of face-to-face instruction augmented by independent computer-guided instruction.  Within each campus, the students were assessed in the same way regardless of their instruction method, and across all campuses they also took a standardised test of statistics competence.   Statistics is a good target for this sort of experiment, because it is a widely required course, and the median introductory statistics course is not very good.

The results were interesting.  The students using the hybrid computer-guided approach found the course less interesting than those with live instructors, but their performance in the course and in the standardised tests was the same.   If you ignore the cost of developing the software (which in this case already existed), the computer-guided approach would allow more students to be taught by the same number of instructors, saving money in the long run.

This doesn’t mean instructors are obsolete — people like face-to-face classes, and we do actually care if students end up interested in statistics –but it does mean that we need to think about the most efficient ways to use class contact time.  There’s an old joke about lectures as a method of transferring information from the lecturer’s notes into the students’ notebooks without it passing through the brains of either.  We’ve got the internet for that, now.

July 9, 2012

Earthquake maps

Stuff is linking to a map of earthquakes by John Nelson of IDV Solutions.  Long-term readers may recall my earthquake map, which uses just the earthquakes since 1973, where the data is more complete.   John Nelson’s map is certainly prettier, but I think mine is clearer.

Book review: Thinking, Fast and Slow

Daniel Kahneman and Amos Tversky made huge contributions to our understanding of why we are so bad at prediction.  Kahneman won a Nobel Prize[*] for this in 2002 (Tversky failed to satisfy the secondary requirement of still being alive).  Kahneman has now written a book, Thinking, Fast and Slow about their research.  Unlike some of his previous writing, this book is designed to be shelved in the Business/Management section of bookshops and read by people who might otherwise be  looking for their cheese.

The “Fast” and “Slow” of the title are two systems of thought: the rapid preconscious judgement that we use for most of our decision-making, and the conscious and deliberate evaluation of alternatives and probabilities that we like to believe we use.   The “Fast” system relies very heavily on stereotyping — finding the best match for a situation in a library of stories — and so is subject to predictable and exploitable biases.  The “Slow” system can be trained to do much better, but only if we can force it to be used.

A dramatic example of the sort of mischief the “fast” system can get up to is anchoring bias.  Suppose you ask a bunch of people how many UN-member countries are in Africa.  You will get a range of guesses, probably not very accurate, and perhaps a few people who actually know the answer.  Suppose you had first asked people to write down the last two digits of their telephone number, or to spin a roulette wheel and write down the number that is chosen, and then to guess how many countries there are in Africa.  Empirically, across a range of situations like this, there is a strong correlation between the obviously irrelevant first number and the guess.   This is an outrageous finding, but it is very well confirmed.   It’s one of the reasons that bogus polls are harmful even if you know they are bogus.

Kahneman gives many other examples of cognitive illusions generated by the ‘fast’ system of the mind.  As with optical illusions, they don’t lose their intuitive force when you understand them, but you can learn not to trust your intuition in situations where it’s going to be biased.

One minor omission of the book is that there’s not much explanation of why we are so stupid: Kahneman points out, and documents, that thinking uses up blood sugar and is biologically expensive, but that doesn’t explain why the mistakes we make are so simple.  Research in computer science and philosophy, by people actually trying to implement thinking, gives one possibility, under the general name of “the frame problem“.  We know an enormous number of facts and relationships between them, and we cannot afford to investigate the logical consequences of all these facts when trying to make a decision.  The price of tea in China really is irrelevant to most decisions, but not to decisions about tea purchases, or about souvenir purchases when in Beijing, or to living-wage levels in Fujian.  We need some way of ignoring the price of tea in China, and millions of other facts, except very occasionally when they are relevant, without having to deduce their irrelevance each time.  Not surprisingly, it sometimes misfires and treats information as important when it is actually irrelevant.

Read this book.  It might help you think better, and at least will give you better excuses for your mistakes.

 

* to quote Daniel Davies: “blah blah blah Sveriges Riksbank. Nobody cares, you know.”

Kiwi workers say “don’t know” to more migrants?

Or perhaps not. It’s hard to tell.

The Herald’s headline is “Kiwi workers say ‘no’ to more migrants”, with the reported data apparently being on ethnic diversity in the workplace, rather than migration

  • 27% want more
  • 33% want less
  • 40% not sure

Now, the difference between 27% and 33% is smaller than the margin of sampling error based on 200 NZ respondents (and much smaller than the usual ‘maximum margin of error’ calculation), but that’s not the main issue.

Another problem is that “More migrants” is not the same as “more ethnic diversity”, and it’s certainly not the same as “more non-English-speaking background”, which the story also mentions.  (I’m a migrant, I’m from an English-speaking background, and I don’t think preferring Aussie Rules to rugby is the sort of ethnic diversity they had in mind).

More important, though, is the question of whether this is a real survey or a bogus poll.   The story doesn’t say.  If you ask the Google, it points you to a webpage where you can participate in the survey,

In order to continue to provide the most current insights into our modern workplace we need your valuable input.

which is certainly an indicator of bogosity.

On the other hand, Leadership Management Australasia, who run the survey, also give some summary reports.  One report says

The survey design and implementation is overseen by an experienced, independent research practitioner and the systems and process used to conduct the survey ensure valid, reliable and representative samples.

which seems to argue for a real survey (though not as convincingly as if they’d actually named the independent research practitioner).   So perhaps the self-selected part of the sample isn’t all of it, and perhaps they do some sensible reweighting?

If you look at the demographic profile of the survey, though, at least two-thirds of the participants are male, even at the non-managerial level.  Now, in both NZ and Australia, male employment is higher than female, but it’s not twice as high.  The gender profiles are definitely not representative.  So even if the survey is making some efforts to be a representative sample, it isn’t succeeding.

 

[Updated to add: in case it’s not clear, in the last paragraph, I’m talking about the summary report for second quarter 2011]