Posts written by Thomas Lumley (2645)

avatar

Thomas Lumley (@tslumley) is Professor of Biostatistics at the University of Auckland. His research interests include semiparametric models, survey sampling, statistical computing, foundations of statistics, and whatever methodological problems his medical collaborators come up with. He also blogs at Biased and Inefficient

August 9, 2017

Briefly

  • From econ blog “Worthwhile Canadian Initiative”:  “The fraction of children earning more than their parents fell from approximately 90% for children born in 1940 to around 50% for children entering the labor market today. Not children. Boys, perhaps, but not children. “
  • From North and South, a story on what direct-to-consumer genetic testing might be good for.
  • There are lots of websites with useful and interesting data out there, but you need to worry about what the data mean. Kaiser Fung has an example from a Kaggle challenge involving Hollywood movies “Huge alarm bells should be going off in the analyst’s head right around now. There were only eleven movies about vampires? Only eleven martial arts movies? Only twelve movies involving superheroes?” (via Andrew Gelman)
  • Wired magazine reprints a Harper’s story about that 1984 revolution in numerical computing, the spreadsheet. “It is not far-fetched to imagine that the introduction of the electronic spreadsheet will have an effect like that brought about by the development during the Renaissance of double-entry bookkeeping. ” If anything, an underestimate.
August 8, 2017

Breast cancer alcohol twitter

Twitter is not an ideal format for science communication, because of the 140-character limitations: it’s easy to inadvertently leave something out.  Here’s one I was referred to this morning (link, so you can see if it is retracted)

latta

Usually I’d think it was a bit unfair to go after this sort of thing on StatsChat.  The reason I’m making an exception here is the hashtag: this is a political statement by a person of mana.

There’s one gross inaccuracy (which I missed on first reading) and one sub-optimal presentation of risk.  To start off, though, there’s nothing wrong with the underlying number: unlike many of its ilk it isn’t an extrapolation from high levels of drinking and it isn’t obviously confounded, because moderate drinkers are otherwise in better health than non-drinkers on average.  The underlying number is that for each standard drink per day, the rate of breast cancer increases by a factor of about 1.1.

The gross inaccuracy is the lack of a per day qualifier, making the statement inaccurate by a factor of several thousand.  An average of one standard drink per day is not a huge amount, but it’s probably more than the average for women in NZ (given the  2007/08 New Zealand Alcohol and Drug Use Survey finding that about half of women drank alcohol less than weekly).

Relative rates are what the research produces, but people tend to think in absolute risks, despite the explicit “relative risk” in the tweet.  The rate of breast cancer in middle age (what the data are about) is fairly low. The lifetime risk for a 45 year old woman (if you don’t die of anything else before age 90) is about 12%.  A 10% increase in that is 13.2%, not 22%. It would take about 7 drinks per day to roughly double your risk (1.17=1.94)  — and you’d have other problems as well as breast cancer risk.

 

August 7, 2017

Millennials and their pink wine

From Stuff, under  the headline Millennials love rose so much they’ve warped the traditional wine market

Millennials dominate Kiwi rose drinking, according to the report. Seventeen per cent of the still wine drunk by under-24s is rose . At 25-35 years it is about 11 per cent and at 35-44 years it drops to 6 per cent. 

Even if that’s true, younger people are less likely to be drinking wine than older people. Here are two graphs of the probability that the most recent alcoholic drink was wine, for women on the left and men on the right, by age groups (source).  The blue is under-24, then 25-44, 45-64, and 65+.

women-winemen-wine

A slightly larger proportion of the wine drunk by millennials is pink, but that’s not the same as saying they drink a large proportion of all the pink wine.

 

August 5, 2017

Just a temporary inconvenience

From Radio NZ

The book explores the widely held view that farm livestock are responsible for an enormous net production of new global warming gases.

“Once you take into account the entire cycle of the life of a cow, it’s actually impossible for the cow to omit even one extra atom of carbon to the atmosphere that wasn’t there already there, they are carbon neural in the end.” he says.

As you’d expect, there’s a sense in which this is completely true. It’s just not a sense that contradicts the standard views of methane and global warming.

What’s going on is easier to see if you consider carbon outputs from the other end of the cow.  Some of the carbon a cow takes in comes out as cowshit. This carbon doesn’t lie around for ever; it returns to the skies and the soil as part of the Great Circle of Life. Hakuna Matata. This doesn’t happen instantaneously, though. In the short term, you still need to wear sensible footwear or watch your step when you cross the field.

There’s an equilibrium between the production and decay of cowshit. When you increase the number of cows, the ambient cowshit level increases, and settles in at a new, higher equilibrium. When you decrease the number of cows, it decreases towards a new, lower equilibrium. The time this takes is governed by how long cowshit takes to decay, so it’s pretty fast.

In a similar, but more serious way, some of the carbon that goes into a cow comes out the front end as methane.  The methane doesn’t hang around in the air for ever; it turns back into carbon dioxide and water. As with cowshit, this doesn’t happen instantaneously.

There’s an equilibrium between the production and decay of methane. When you increase the number of cows, the ambient cow-derived methane level increases, and settles in at a new, higher equilibrium. When you decrease the number of cows, it decreases towards a new, lower equilibrium. The time this takes is governed by how long methane takes to decay: over each passing decade about half of it goes away.

Carbon emitted as methane, unlike carbon emitted in cowshit, is more than a local nuisance.  Per atom of carbon, methane has 24 times the greenhouse warming effect of CO2, and while it doesn’t last for ever, it lasts long enough to make an important contribution to climate change.  There’s more than twice as much methane in the atmosphere now as there was two centuries ago.

Cows are long-term carbon-neutral: that means reducing cow numbers (or finding ways to reduce their methane production) would, in mere decades, roll back the increases they’ve caused in an important greenhouse gas.

August 2, 2017

Briefly

  • Graphics: there’s a solar eclipse soon in the US. Washington Post‘s WonkBlog shows Google Trends search interest in iteclipse
  • Persuasive Cartography: 800 historical maps “intended primarily to influence opinions or beliefs – to send a message – rather than to communicate geographic information.”
  • Should there even be an app for this?”  and other tech questions from a workshop on design ethics. (Subquestion: should there be a prediction of this?)
  • “I never knew until very recently that the standard National Readership Survey socio-demographic classifications – ABC1, C2DE etc – deal with pensioners by classifying them all as working-class unless they are rich enough to be considered independently wealthy and therefore bucketed in with the As. ” Alex Harrowell on social class assessment and the politics of data
August 1, 2017

Holiday travel trends

The Herald has a story and video graphic, and a nice interactive graphic on international travel by Kiwis since 1979.  The story is basically good (and even quotes a price corrected for inflation).

Here’s one frame of the video graphic
escape

First, a lot of the world isn’t coloured. There are New Zealanders who have visited say, Germany or Turkey or Egypt, even though these countries never make it into the 1-24,999 colour category. It looks as if the video picks a set of 16 countries and follows just those forward in time: we’re not told how these were picked.

Second, there’s the usual map problem of big things looking big (exacerbated by the Mercator projection). In 1999, more people went to Fiji than the US; more to Samoa than France. A map isn’t good at making these differences visually obvious, though the animation helps. And, tangentially, if you’re going to use almost a third of the map real estate on the region north of 60°, you should notice that Alaska is part of the USA.

The other, more important, issue that’s common to the whole presentation (and which I understand is being updated at the moment) is what the country data actually mean. It seems that it really is holiday data, excluding both business and visiting friends/relatives (comparing the video to this from Figure.NZ), but it’s by “country of main destination”.  If you go to more than one country, only one is counted.  That’s why the interactive shows zero Kiwis travelling to the Vatican City, and it may help explain numbers like 300 for Belgium.

Official statistics usually measure something fairly precise, but it’s not always the thing that you want them to measure.

July 30, 2017

Coffee news?

In 2015, the Herald said

Drinking the caffeine equivalent of more than four espressos a day is harmful to health, especially for minors and pregnant women, the European Union food safety agency has said.

“It is the first time that the risks from caffeine from all dietary sources have been assessed at EU level,” the EFSA said, recommending that an adult’s daily caffeine intake remain below 400mg a day.

(I quoted it at the time: the link seems to be dead now).

Now we have, under the headline Good news for coffee lovers: Caffeine is harmless, says research

A review of 44 trials dispelled the widespread myth that caffeine, found in tea, coffee and fizzy drinks, is bad for the body.

It found that sticking to the recommended daily amount of 400mg – the equivalent four cups of coffee or eight cups of tea – has no lasting damage on the body.

The recommendation that 400mg/day is generally safe was described as ‘caffeine is dangerous’ in 2015 and ‘caffeine is harmless’ now.

Other not-news about this is the not-new research. Obviously the Daily Mail (the only link) isn’t a research source. The research was published in Complete Nutrition, a professional magazine for UK dieticians. As their website says

Each issue of CN is packed with articles which are practical, educational and topical, and all are written by independent, well-respected authors from across the profession.

That’s a valuable mission for a journal, but it would be surprising if an expert opinion article in a journal like that contained new research worth international headlines.

What are election polls trying to estimate? And is Stuff different?

Stuff has a new election ‘poll of polls’.

The Stuff poll of polls is an average of the most recent of each of the public political polls in New Zealand. Currently, there are only three: Roy Morgan, Colmar Brunton and Reid Research. 

When these companies release a new poll it replaces their previous one in the average.

The Stuff poll of polls differs from others by giving weight to each poll based on how recent it is.

All polls less than 36 days old get equal weight. Any poll 36-70 days old carries a weight of 0.67, 70-105 days old a weight 0.33 and polls greater than 105 days old carry no weight in the average.

In thinking about whether this is a good idea, we’d need to first think about what the poll is trying to estimate and about the reasons it doesn’t get that target quantity exactly right.

Officially, polls are trying to estimate what would happen “if an election were held tomorrow”, and there’s no interest in prediction for dates further forward in time than that. If that were strictly true, no-one would care about polls, since the results would refer only to the past two weeks when the surveys were done.

A poll taken over a two-week period is potentially relevant because there’s an underlying truth that, most of the time, changes more slowly than this.  It will occasionally change faster — eg, Donald Trump’s support in the US polls seems to have increased after James Comey’s claims about Clinton’s emails in the US, and Labour’s support in the UK polls increased after the election was called — but it will mostly change slower. In my view, that’s the thing people are trying to estimate, and they’re trying to estimate it because it has some medium-term predictive value.

In addition to changes in the underlying truth, there is the idealised sampling variability that pollsters quote as the ‘margin of error’. There’s also larger sampling variability that comes because polling isn’t mathematically perfect. And there are ‘house effects’, where polls from different companies have consistent differences in the medium to long term, and none of them perfectly match voting intentions as expressed at actual elections.

Most of the time, in New Zealand — when we’re not about to have an election — the only recent poll is a Roy Morgan poll, because  Roy Morgan polls more much often than anyone else.  That means the Stuff poll of polls will be dominated by the most recent Roy Morgan poll.  This would be a good idea if you thought that changes in underlying voting intention were large compared to sampling variability and house effects. If you thought sampling variability was larger, you’d want multiple polls from a single company (perhaps downweighted by time).  If you thought house effects were non-negligible, you wouldn’t want to downweight other companies’ older polls as aggressively.

Near an election, there are lots more polls, so the most recent poll from each company is likely to be recent enough to get reasonably high weight. The Stuff poll is then distinctive in that it complete drops all but the most recent poll from each company.

Recency weighting, however, isn’t at all unique to the Stuff poll of polls. For example, the pundit.co.nz poll of polls downweights older polls, but doesn’t drop the weight to zero once another poll comes out. Peter Ellis’s two summaries both downweight older polls in a more complicated and less arbitrary way; the same was true of Peter Green’s poll aggregation when he was doing it.  Curia’s average downweights even more aggressively than Stuff’s, but does not otherwise discard older polls by the same company. RadioNZ averages the only the four most recent available results (regardless of company) — they don’t do any other weighting for recency, but that’s plenty.

However, another thing recent elections have shown us is that uncertainty estimates are important: that’s what Nate Silver and almost no-one else got right in the US. The big limitation of simple, transparent poll of poll aggregators is that they say nothing useful about uncertainty.

July 29, 2017

Anything goes

According to a story in the Herald, based on what looks like it might be a bogus poll (press release), you need $5.3 million in Australia now to be considered rich.  If we assumed the number did actually measure something, how surprising would it be?

Before “Who wants to be a millionaire?” was a quiz show franchise, it was a Cole Porter song, from the  1956 movie “High Society”, so that seems a reasonable comparison period. The Australian CPI has gone up by a factor of 15.6 since 1956 (and while Australia didn’t have dollars until 1966, US and Australian dollars were roughly comparable then).

On top of pure currency conversion, though, Australia is richer now than in 1956.  Australia’s GDP in current purchasing-power adjusted dollars is nearly 8 times what it was in 1956. The population has gone from 9.4 million to 24.1 million, so real GDP per capita is up by a factor of about 3.5.

So, a 1956 million would be 15.6 current millions just from inflation, and over $50 million as a share of Australia’s economy: a millionaire in those days was not just rich, but Big Rich — as the song says: “flashy flunkies everywhere… a gigantic yacht… liveried chauffeur.”

We’re not given any real reason to believe the $5.3 million figure — there’s no reason you should rely on it more than your own guess. And ‘millionaire’ isn’t a useful comparison without a lot of additional qualification.

July 27, 2017

Will we ever use this in real life?

From deep in the archives at Language Log

The Pirahã language and culture seem to lack not only the words but also the concepts for numbers, using instead less precise terms like “small size”, “large size” and “collection”. And the Pirahã people themselves seem to be suprisingly uninterested in learning about numbers, and even actively resistant to doing so, despite the fact that in their frequent dealings with traders they have a practical need to evaluate and compare numerical expressions. A similar situation seems to obtain among some other groups in Amazonia, and a lack of indigenous words for numbers has been reported elsewhere in the world.

Many people find this hard to believe. These are simple and natural concepts, of great practical importance: how could rational people resist learning to understand and use them? I don’t know the answer. But I do know that we can investigate a strictly comparable case, equally puzzling to me, right here in the U.S. of A.

From context, you can probably guess where he’s heading