Posts written by Thomas Lumley (2645)

avatar

Thomas Lumley (@tslumley) is Professor of Biostatistics at the University of Auckland. His research interests include semiparametric models, survey sampling, statistical computing, foundations of statistics, and whatever methodological problems his medical collaborators come up with. He also blogs at Biased and Inefficient

May 4, 2017

Summarising a trend

Keith Ng drew my attention on Twitter to an ad from Labour saying “Under National, the number of young people not earning or learning has increased by 41%”.

When you see this sort of claim, you should usually expect two things: first, that the claim will be true in the sense that there will be two numbers that differ by 41%; second, that it will not be the most informative summary of the data in question.

If you look on Infoshare, in the Household Labour Force Survey, you can find data on NEET (not in education, employment, or training).  The number was 64100 in the fourth quarter of 2008, when Labour lost the election.  It’s now (Q1, 2017) 90800, which is, indeed, 41% higher.  Let’s represent the ad by a graph:

neet1

 

We can fill in the data points in between:
neet2
Now, the straight line doesn’t look as convincing.

Also, why are we looking at the number, when population has changed over this time period. We really should care about the rate (percentage)
neet3
Measuring in terms of rates the increase is smaller — 27%.  More importantly, though, the rate was even higher at the end of the first quarter of National’s administration than it is now.

The next thing to notice is the spikes every four quarters or so: NEET is higher in the summer and lower in the winter because of the school  year.  You might wonder if StatsNZ had produced a seasonally adjusted version, and whether it was also conveniently on Infoshare…
need4
The increase is now 17%

But for long-term comparisons of policy, you’d probably want a smoothed version that incorporates more than one quarter of data. It turns out that StatsNZ have done this, too, and it’s on Infoshare.
neet5
The increase is, again 17%. Taking out the seasonal variation, short-term variation, and sampling noise makes the underlying pattern clearer.  NEET increased dramatically in 2009, decreased, and has recently spiked. The early spike may well have been the recession, which can’t reasonably be blamed on any NZ party.  The recent increase is worrying, but thinking of it as trend over 9 years isn’t all that helpful.

May 3, 2017

A century of immigration

Given the discussions of immigration in the past weeks, I decided to look for some historical data.  Stats NZ has a report “A Century of Censuses”, with a page on ‘proportion of population born overseas.” Here’s the graph

nz-oseas-born

The proportion of immigrants has never been very low, but it fell from about 1 in 2 in the late 19th century to about 1 in 6 in the middle of the 2oth century, and has risen to about 1 in 4 now. The increase has been going on for the entire lifetime of any NZ member of Parliament; the oldest was born roughly at Peak Kiwi in the mid-1940s.

Seeing that immigrants have been a large minority of New Zealand for over a century doesn’t necessarily imply anything about modern immigration policy — Hume’s Guillotine, “no ought deducible from is,” cuts that off.  But I still think some people would find it surprising.

 

April 28, 2017

Trends and pauses

There’s a story at the Guardian about whether there has been a ‘pause’ and an ‘acceleration’ in global warming.  The underlying research paper actually puts the question more clearly

While it is clear and undisputed that the global temperature data show short periods of greater and smaller warming trends or even short periods of cooling, the key question is: is this just due to the ever-present noise, i.e. short-term variability in temperature? Or does it signify a change in behavior, e.g. in the underlying warming trend?

Models for climate change predict that annual mean surface temperature should be going up fairly smoothly, so that the trend over a decade or so looks like a straight line. A deviation from this trend might indicate important factors have been left out of the model, or might indicate that the background processes are changing (eg as ice sheets retreat).  If you look at the recent past, compared to a straight-line trend, the observed data dipped below the straight line for a few years and have now caught up.  This raises the question of whether either of these indicated an important change in the underlying processes or a new inadequacy of the models.

To start with, let’s establish that we’re not talking about measurement error here.  The variability of annual mean temperatures around the straight-line trend isn’t like the variability of opinion poll results around a trend. The observed data are the truth.  The world really did warm less for a couple of years; it really has warmed more since then.  The straight line trend omits many factors that we know are relevant: events such as volcanic eruptions that affect the incoming sunlight, and events such as El Niño that affect the balance between air and ocean warming.

The question is whether the straight line trend is changing (fast enough to worry about).  You might reasonably object that the annual mean temperatures are far too crude to make that sort of decision; that you need much more details and more sophisticated modelling. As it turns out, you’d be right. However, the crude appearance of a slowdown and speedup in the annual means has been the fuel for a lot of discussion, so it’s worth evaluating.

What the research paper did was to model the deviations from the straight line trend as a simple random process, ignoring any year-to-year correlation.  The researchers could then evaluate mathematically how likely we would be to see an apparent pause or acceleration in warming with that amount of random variation, if in fact the trend was a perfect straight line. The deviations we have seen in the recent past are no larger than you’d expect just from the variation around a constant trend.

To be clear, this doesn’t mean there have been no changes in the trend.  In fact, we know that El Niño does cause systematic changes.  What it means is that the annual mean temperatures alone aren’t enough information to tell us about changes over a period as short as a few years. You shouldn’t change your beliefs (in any direction) over data like that. If the ‘hiatus’ had gone on for a decade, it would have meant something. If the acceleration goes on for a decade, it will mean something. But two or three years isn’t long enough to say anything.  It’s like looking at a month of data on road deaths: you can’t — or at least shouldn’t  — say much.

April 27, 2017

On debates about data

On Wednesday, the NZ Herald website featured a story and graphics by Harkanwal Singh and Lincoln Tan on immigration. This story was based on permanent and long-term migration data from Statistics New Zealand. The graphics allowed readers to explore the data for themselves. The data source was accurately described and was well targeted to the current political discussion about changing immigration policies.

The specific data set and visualisation used are not the only possible ones, and reasoned criticism of the data and analyses is entirely legitimate. StatsChat encourages that sort of thing. We have done it ourselves, and we have published links when other people do it.

Winston Peters, however, claimed that the Herald story was “fake news” and attributed the conclusions to the reporters being Asian immigrants themselves. The first claim is factually incorrect; the second (in the absence of convincing evidence) is outrageous.

James Curran (Professor of Statistics)
Thomas Lumley (Professor of Statistics)
Chris Triggs (Professor of Statistics)

April 26, 2017

Simplifying to make a picture

1. Ancestry.com has maps of the ancestry structure of North America, based on people who sent DNA samples in for their genotype service (click to embiggen)ncomms14238-f3

To make these maps, they looked for pairs of people whose DNA showed they were distant relatives, then simplified the resulting network into relatively stable clusters. They then drew the clusters on a map and coloured them according to what part of the world those people’s distant ancestors probably came from.  In theory, this should give something like a map of immigration into the US (and to a lesser extent, of remaining Native populations).  The map is a massive oversimplification, but that’s more or less the point: it simplifies the data to highlight particular patterns (and, necessarily, to hide others).  There’s a research paper, too.

 

2. In a satire on predictive policing, The New Inquiry has an app showing high-risk neighbourhoods for financial crime. There’s also a story at Buzzfeed.

sub-buzz-24605-1493145131-7

The app uses data from the US Financial Regulatory Authority (FINRA), and models the risk of financial crime using the usual sort of neighbourhood characteristics (eg number of liquor licenses, number of investment advisers).

 

3. The Sydney Morning Herald had a social/political quiz “What Kind of Aussie Are You?”.

1486745652102

They also have a discussion of how they designed the 7 groups.  Again, the groups aren’t entirely real, but are a set of stories told about complicated, multi-dimensional data.

 

The challenge in any display of this type is to remove enough information that the stories are visible, but not so much that they aren’t true– and not everyone will agree on whether you’ve succeeded.

April 25, 2017

Electioneering and statistics

In New Zealand, the Government Statistician reports to the Minister of Statistics, currently Mark Mitchell.  For about a decade, the UK has had a different system, where the National Statistician reports to the UK Statistics Authority, which is responsible directly to Parliament. The system is intended to make official statistics more clearly independent of the government of the day.

An additional role of the UK Statistics Authority is as a sort of statistics ombudsman when official statistics are misused.  There’s a new letter from the Chair to the UK political parties

The UK Statistics Authority has the statutory objective to promote and safeguard the production and publication of official statistics that serve the public good.

My predecessors Sir Michael Scholar and Sir Andrew Dilnot have in the past been obliged to write publicly about the misuse of official statistics in other pre-election periods and during the EU referendum campaign. Misuse at any time damages the integrity of statistics, causes confusion and undermines trust.

I write now to ask for your support and leadership to ensure that official statistics are used throughout this General Election period and beyond, in the public interest and in accordance with the principles of the Code of Practice for Official Statistics. In particular, the statistical sources should be clear and accessible to all; any caveats or limitations in the statistics should be respected; and campaigns should not pick out single numbers that differ from the picture painted by the statistics as a whole.

I am sending identical letters to the leaders of the main political parties, with a copy to Sir Jeremy Heywood, Cabinet Secretary.

We don’t have anyone whose job it is to write that sort of letter here, but it would be nice if the political parties (and their partisans) still followed this advice.

April 24, 2017

Briefly

  • The Herald (from the Daily Mail) recommends drinking beetroot juice, based on a study of brain waves: “This finding could help people who are at-risk of brain deterioration to remain functionally independent, such as those with a family history of dementia“.  The NHS Choices blog commented on a similar study by the same research group in 2010; their comments still apply.
  • Testimonials and motivational speakers tell you “I did this and look how it turned out”.  As XKCD illustrates, results may not be typical 
  • “Data made available for reanalysis, a journal that promptly responded to the outcomes of that reanalysis, and a finding that could save lives.” (from Stat). Another moral to the story: don’t edit data by copy-and-paste.
  • The company says it has studies that back up its claims, but refused to release them on the grounds that they are commercial-in-confidence.” It appears that Johnson & Johnson would rather pull their ad than let people look at the evidence. (from The Age)
  • it’s not acceptable if you’ve got the information readily available to leave it to the last minute for release, that’s not what the Act says you can do”  The Chief Ombudsman interviewed by Newsroom  about the Official Information Act.

And finally

If you give a mouse a strawberry…

 

So, the Herald (from the Daily Mailhas a headline Why women should eat a punnet of strawberries a day. That seems a little extreme, especially as punnets of strawberries are fairly seasonal.

The story leads off with

Eating just 15 strawberries a day protected mice from aggressive breast cancer in a new medical study.

So, first of all, mice, not women.  Also, when you go to the open-access research paper, it didn’t exactly ‘protect’ the mice.  The mice had cells from a breast cancer cell culture implanted under their skins, and the study looked at the change in size of those implanted tumours, not at spread within the mouse or health of the mouse or anything like that.  It’s a useful approach to learning about cancer cell biology, but not all that close to preventing or treating human cancers.

More surprisingly, though, “15 strawberries a day” seems quite a lot for a mouse — several times its body weight. The story changes a bit later:

In total, the strawberries made up 15 percent of the mice’s diet. That is just shy of the recommended daily amount of fruit we should eat each day, and would be equivalent to a punnet of strawberries, reported the Daily Mail.

A figure of 15% seems more plausible than 15 strawberries, though it’s still not quite true, since actually the mice were given concentrated strawberry extract in their food rather than strawberries.  Using the standard (lowish) estimate of 2000 kcal/day, 15% of calories would  be 300 kcal/day  which would take nearly a kilogram of strawberries.

Previous studies have already shown that eating between 10 and 15 strawberries a day can make arteries healthier by reducing blood cholesterol levels.

There isn’t a reference, but the same researcher has studied strawberries and cholesterol (this time even in humans). The ‘between 10 and 15 strawberries a day’ was actually 500g per day.

[via Sam Warburton]

April 17, 2017

Slow on the uptake

Q: Did you see gin can increase your metabolism?

A: Um…

Q: Here, in the Herald, new research from Latvia!

A:  Not really convincing.

Q: Why? Is it in mice?

A: Up to a point.

Q: <reads> Yes, it’s in mice: “In fact, the mice who were fed regular doses of the spirit saw a 17 percent increase in their metabolic rate”.  That’s a lot, isn’t it?

A: Indeed. One might almost say an incredible amount.

Q:  Ok, were these some special sort of mutant mouse with a weird metabolism?

A: The story doesn’t seem to say.

Q: Of course it doesn’t, but can’t you find the original research paper? The story says it’s in Food & Nature. Doesn’t University of Auckland subscribe to it?

A: No.

Q: That’s usually a bad sign, isn’t it?

A: Especially in this case. The journal doesn’t exist, the university doesn’t exist, and Professor Thisa Lye is, apparently a lie.

Q:  😕?

A: The story is two weeks old. It was an April Fool’s hoax. Thanks to Elle Hunt I was saved potentially quite a bit of time looking for the journal. She tweeted a link from Latvian Public Broadcasting, who have tracked the story down.

Q: So the Herald got it from the Daily Mail who got it from Yahoo who got it from Prima. And none of them checked that the research existed? I mean, ok, checking science isn’t what journalists are trained to do, but checking that sources actually exist? With Google?

A:  On the positive side, no mice were harmed in conducting the research.

April 14, 2017

Cyclone uncertainty

Cyclone Cook ended up a bit east of where it was expected, and so Auckland had very little damage.  That’s obviously a good thing for Auckland, but it would be even better if we’d had no actual cyclone and no forecast cyclone.  Whether the precautions Auckland took were necessary (at the time) or a waste  depends on how much uncertainty there was at the time, which is something we didn’t get a good idea of.

In the southeastern USA, where they get a lot of tropical storms, there’s more need for forecasters to communicate uncertainty and also more opportunity for the public to get to understand what the forecasters mean.  There’s scientific research into getting better forecasts, but also into explaining them better. Here’s a good article at Scientific American

Here’s an example (research page):

hurricane

On the left is the ‘cone’ graphic currently used by the National Hurricane Center. The idea is that the current forecast puts the eye of the hurricane on the black line, but it could reasonably be anywhere in the cone. It’s like the little blue GPS uncertainty circles for maps on your phone — except that it also could give the impression of the storm growing in size.  On the right is a new proposal, where the blue lines show a random sample of possible hurricane tracks taking the uncertainty into account — but not giving any idea of the area of damage around each track.

There’s also uncertainty in the predicted rainfall.  NIWA gave us maps of the current best-guess predictions, but no idea of uncertainty.  The US National Weather Service has a new experimental idea: instead of giving maps of the best-guess amount, give maps of the lower and upper estimates, titled: “Expect at least this much” and “Potential for this much”.

In New Zealand, uncertainty in rainfall amount would be a good place to start, since it’s relevant a lot more often than cyclone tracks.

Update: I’m told that the Met Service do produce cyclone track forecasts with uncertainty, so we need to get better at using them.  It’s still likely more useful to experiment with rainfall uncertainty displays, since we get heavy rain a lot more often than cyclones.