Posts written by Thomas Lumley (2645)

avatar

Thomas Lumley (@tslumley) is Professor of Biostatistics at the University of Auckland. His research interests include semiparametric models, survey sampling, statistical computing, foundations of statistics, and whatever methodological problems his medical collaborators come up with. He also blogs at Biased and Inefficient

April 17, 2013

Drawing the wrong conclusions

A few years ago, economists Carmen Reinhart and Kenneth Rogoff wrote a paper on national debt, where they found that there wasn’t much relationship to economic growth as long as debt was less than 90% of GDP, but that above this level economic growth was lower.  The paper was widely cited as support for economic strategies of `austerity’.

Some economists at the University of Massachusetts attempted to repeat their analysis, and didn’t get the same result.  Reinhart and Rogoff sent them the data and spreadsheets they had used, and it turns out that the analysis they had done didn’t quite match the description in the paper.  Part of the discrepancy was an error in an Excel formula that accidentally excluded a bunch of countries, but Reinhart and Rogoff also deliberately excluded some countries and times that had high growth and high debt (including Australia and NZ immediately post-WWII), and gave each country the same weight in the analysis regardless of the number of years of data included. (paper — currently slow to load, summary by Mike Konczal)

Some points:

  • The ease of making this sort of error in Excel is exactly why a lot of statisticians don’t like Excel (despite its other virtues), so that has received a lot of publicity.
  • Reinhart and Rogoff point out that they only claimed to find an association, not a causal relationship, but they certainly knew how the paper was being used, and if they didn’t think provided evidence of a causal relationship they should have said something a lot earlier. (I think Dan Davies on Twitter put it best)
  • Chris Clary, who is a PhD student at MIT, points out that the first author (Thomas Herndon) on the paper demonstrating the failure to replicate is also a grad student, and notes that replicating things is job often left to grad students.
  • The Reinhart and Rogoff paper wasn’t the primary motivation for, say,  the UK Conservative Party to want to cut taxes and government spending. The Conservatives have always wanted to cut taxes and government spending. Cutting taxes and spending is a significant part of their basic platform. The paper, at most, provided a bit of extra intellectual cover.
  • The fact that the researchers handed over their spreadsheet pretty much proves they weren’t deliberately deceptive — but it’s a lot easy to convince yourself to spend a lot of time checking all the details of a calculation when you don’t like the answer than when you do.

Roger Peng, at  Johns Hopkins, has also written about this incident. It would, in various ways, have been tactless for him to point out some relevant history, so I will.

The Johns Hopkins air pollution research group conducted the largest and most comprehensive study of health effects of particulate air pollution, looking at deaths and hospital admissions in the 90 largest US cities.  This was a significant part of the evidence used in setting new, stricter, air pollution standards — an important and politically sensitive topic, though a few order of magnitude less so than austerity economics.  One of Roger’s early jobs at Johns Hopkins was to set up a system that made it easy for anyone to download their data and reproduce or vary their analyses. The size of the data and the complexity of some of the analyses meant just emailing a spreadsheet to people was not even close to acceptable.

Their research group became obsessive (in a good way) about reproducibility long before other researchers in epidemiology.  One likely reason is a traumatic experience in 2002, when they realised that the default settings for the software they were using had led to incorrect results for a lot of their published air pollution time series analyses.  They reported the problem to the EPA and their sponsors, fixed the problem, and reran all the analyses in a couple of weeks; the qualitative conclusions fortunately did not change.  You could make all sorts of comparisons with the economists’ error, but that is left as an exercise for the reader.

 

Open data on the West Island

If you want to get Australian census summary data, you can download it from the Australian Bureau of Statistics, or buy a DVD for A$250.

An article in iTNews explains why someone might pay rather than downloading

“You have to click to download each pack individually, and they’ve set the site up deliberately to make it difficult to use a browser plugin to download everything that is contained on the released DVD image,” Bowland told iTNews.

That’s not hyperbole: Grahame Bowland quotes JavaScript code comments that actually say they are trying to make automatic downloading difficult.

Or, the data release is now available using bittorrent, thanks to Bowland, who bought the DVD (this is perfectly legit: the data are Creative Commons licenced).

(via @keith_ng)

April 16, 2013

Bogus poll news again

Stuff has a story “Dishonest Kiwi Travellers”, based on a survey press release from Hotels.com.  The survey asked people if they had stolen things from hotels and used the responses to rank countries by honesty, with the travellers who denied taking things being rated more honest than the ones who admitted it (rather than the other way around).

Fortunately it doesn’t really matter how honest the responses were, since if you follow a few links you can find a press release for the Canadian part of the survey, which admits

In a recent survey hotels.com® asked its Canadian email subscribers , including those in Quebec, about what they look for in hotel accommodations, and you might be surprised at what they had to say.

Or in other words, it’s a bogus poll.

April 15, 2013

Two good local pieces

  • Martin Johnston in the Herald on a nationwide blood pressure survey. Blood pressures are up.  This is not good.
  • Nikki Macdonald in Stuff, on recreational genotyping (or as the story more properly calls it, “direct to consumer” genotyping).
April 14, 2013

Infectious disease science communcation

Today we have two examples of the important issue of communicating scientific knowledge about infectious disease epidemics.

The first is the WHO, which is doing an excellent job of describing the limited information about the new H7N9 influenza outbreak in China. Their media release is here, and they’ve had someone answering questions on Twitter as well as more traditional venues.  There’s currently evidence of a small amount of human-to-human transmission, but not enough to sustain a pandemic. On the other hand, the virus does appear to have mutated to live more successfully in people, and this could continue.  They don’t advise actually doing anything specific at the moment.

 

The second is the UK measles epidemic, where The Independent, has as its top front-page headline “MMR scare doctor: this outbreak proves I was right”.  Of course,  it does nothing of the sort, as the story admits later . He’s claiming that the MMR vaccine should have been replaced by three single vaccines, and even if you believe that anti-vaccination campaigners would then suddenly have stopped their misrepresentations, having three single vaccinations is actually more dangerous than one combined one.

The Independent is maintaining that its story is accurate if you read the whole thing. Even if that were so, it’s still hard to imagine why the opinion of a discredited researcher and struck-off former doctor is the single most important piece of information they have about epidemic and the world today. And, as  Martin Robins writes in New Statesman

 It would be a great example of the false balance inherent in ‘he-said, she-said’ reporting, except that it isn’t even balanced – Laurance provides a generous abundance of space for Wakefield to get his claims and conspiracy theories across, and appends a brief response from a real scientist at the end. 

April 13, 2013

Briefly

“Cox, when you are a bit older, you will not quote Indian statistics with that assurance. The Government are very keen on amassing statistics – they collect them, add them, raise them to the nth power, take the cube root and prepare wonderful diagrams. But what you must never forget is that every one of these figures comes in the first place from the  [village watchman], who just puts down what he damn pleases”

April 12, 2013

Metrics and multivariate data

If you have a large collection of measurements there isn’t going to be a unique way to put them together into a single ordering: what you get out depends to some extent on your criteria. That doesn’t mean the measurements, or even the rankings, are meaningless, but it does mean that you should work out what your criteria are before you see the data, and that criticism of the choice of criteria is perfectly reasonable.

An illustration is yesterday’s PBRF research evaluation results.  For those of you playing along at home, PBRF is one of the mechanisms the country uses to allocate research funding. Unlike individual research grants, which are based on competition between individual proposals, PBRF funding is allocated to large groups of researchers for long periods of time based on a aggregated results from a single standardised evaluation.

Though it’s not the point of the system, the existence of a large number of grades invariably tempts university management into coming up with ways to combine them to make their institution look good. And they always succeed:

  • In first place, we have the University of Auckland: “secured the largest share of the fund, $80.4m or 30.6% of the national total… This is due in part to the University’s impressive 288 international quality (A-rated) researchers – the greatest number of leading researchers anywhere in the country.”
  • And in first place, the University of Otago: “Otago was ranked first among New Zealand universities in the measure of research quality weighted by its postgraduate roll (AQS (P)) and second in the measure weighted by degree-level enrolments and higher (AQS (E)). The University is the only TEO to be ranked in the top four in all four AQS measures.”
  • In first place, also, Victoria University of Wellington: ” the latest PBRF Evaluation ranks Victoria as number one in New Zealand….With 678 staff actively involved in research, and 70 percent of them operating at the highest levels (ranked as either an A or B), we now have external confirmation of our status as New Zealand’s most research intensive university.”
  • And finally, in first place (special subject), Lincoln University:  “confirms Lincoln University’s position as New Zealand’s specialist land-based university. … Lincoln University also has the highest amount of external funding for research (measured as income/staff member), demonstrating close links with industry and relevance of the University’s research.”

Other institutions aren’t claiming first place, but are still saying that the PBRF demonstrate how successful they have been. This has been another illustration that anyone saying “the data speak for themselves” is not to be trusted. The question matters.

Briefly

  • “With less than two months to go until I graduate from the UC Berkeley Graduate School of Journalism, I’ve been looking back at my experience over the past two years. I’m among a handful of students at the school who are really interested in data journalism and making pretty and functional online news packages. It’s made me think about how J-schools need a more structured and thorough track for us computer-assisted reporters, for lack of a better term.John Osborn

 

April 11, 2013

Power failure threatens neuroscience

A new research paper with the cheeky title “Power failure: why small sample size undermines the reliability of neuroscience” has come out in a neuroscience journal. The basic idea isn’t novel, but it’s one of these statistical points that makes your life more difficult (if more productive) when you understand it.  Small research studies, as everyone knows, are less likely to detect differences between groups.  What is less widely appreciated is that even if a small study sees a difference between groups, it’s more likely not to be real.

The ‘power’ of a statistical test is the probability that you will detect a difference if there really is a difference of the size you are looking for.  If the power is 90%, say, then you are pretty sure to see a difference if there is one, and based on standard statistical techniques, pretty sure not to see a difference if there isn’t one. Either way, the results are informative.

Often you can’t afford to do a study with 90% power given the current funding system. If you do a study with low power, and the difference you are looking for really is there, you still have to be pretty lucky to see it — the data have to, by chance, be more favorable to your hypothesis than they should be.   But if you’re relying on the  data being more favorable to your hypothesis than they should be, you can see a difference even if there isn’t one there.

Combine this with publication bias: if you find what you are looking for, you get enthusiastic and send it off to high-impact research journals.  If you don’t see anything, you won’t be as enthusiastic, and the results might well not be published.  After all, who is going to want to look at a study that couldn’t have found anything, and didn’t.  The result is that we get lots of exciting neuroscience news, often with very pretty pictures, that isn’t true.

The same is true for nutrition: I have a student doing a Honours project looking at replicability (in a large survey database) of the sort of nutrition and health stories that make it to the local papers. So far, as you’d expect, the associations are a lot weaker when you look in a separate data set.

Clinical trials went through this problem a while ago, and while they often have lower power than one would ideally like, there’s at least no way you’re going to run a clinical trial in the modern world without explicitly working out the power.

Other people’s reactions

April 10, 2013

Health claims not berry well supported

I don’t usually bother with general nutrition stories that don’t contain any direct reference to research, but the Herald story about berries was irresistible. There are lots of biologically active compounds in berries, and many of them have been shown to have interesting properties in test-tubes or mice. As you know by now,  this sort of interesting biochemistry is important because it occasionally translates to genuine health benefits, so you should be asking what the human clinical research shows.

If you go to the Cochrane Library (which is free to everyone in New Zealand), and look for clinical research in humans involving blueberries or cranberries you don’t find much. The only topic with enough information to draw any sort of conclusion is on cranberry juice to prevent urinary tract infections. Which it basically doesn’t. The plain-language summary says

Cranberries (usually as cranberry juice) have been used to prevent urinary tract infections (UTIs). Cranberries contain a substance that can prevent bacteria from sticking on the walls of the bladder. This may help prevent bladder and other UTIs. This review identified 24 studies (4473 participants) comparing cranberry products with control or alternative treatments. There was a small trend towards fewer UTIs in people taking cranberry product compared to placebo or no treatment but this was not a significant finding. Many people in the studies stopped drinking the juice, suggesting it may not be a acceptable intervention. Cranberry juice does not appear to have a significant benefit in preventing UTIs and may be unacceptable to consume in the long term. 

As with many fruits and vegetables, eating more of them instead of other stuff is both enjoyable and probably healthy. As with pretty much any food, there might be some specific additional benefits (or harms), but if so we don’t yet have much evidence for them.