NZ Census updates
Two press releases from Stats NZ today
2018 Census data release delayed
Why the 2018 Census data release is delayed
Thomas Lumley (@tslumley) is Professor of Biostatistics at the University of Auckland. His research interests include semiparametric models, survey sampling, statistical computing, foundations of statistics, and whatever methodological problems his medical collaborators come up with. He also blogs at Biased and Inefficient
Two press releases from Stats NZ today
2018 Census data release delayed
Why the 2018 Census data release is delayed
Brian Wansink, a prominent food researcher from Cornell, was forced to retire earlier this year. Andrew Gelman has some good perspectives. Wansink’s research was on contextual effects on eating — eg, the impact of plate size — and a bunch of these papers have now been retracted.
This week, Wired has a Thanksgiving-themed article about his research. It includes this quote from an email he sent to colleagues ahead of his retirement
“We may believe that our papers have been unfairly retracted. But what they can’t retract is the impact these have had on people’s lives and the impact they will continue to have.”
That’s a perfect summary of the problems with studies that over-promise and are over-publicised. Whether the research was done well or not, it’s going to stick. Subsequent developments — whether modifications, replication failures, or retractions — never get the impact of the original claim.
The Herald has a story from the Washington Post on an “AI” screening tool for babysitters, that allegedly uses both computer vision and text processing to screen social media for risk factors. Here are three quotes from it:
1. A company co-founder says
Parents, he said, should see the ratings as a companion that “may or may not reflect the sitter’s actual attributes.”
But the danger of hiring a problematic or violent babysitter, he added, makes the AI a necessary tool for any parent hoping to keep his or her child safe.
The first thing to note about this is that you could make the same claims about astrology or handwriting analysis or a tarot reading. There’s no quantitative information about accuracy given, and it’s hard to see how the company could even know much how accurate its ratings were, or how biased. It’s not even for sure that the risk rating is positively correlated with risk to kids; the company seems careful not to make even a claim this weak.
2.
Parents could, presumably, look at their sitters’ public social media accounts themselves. But the computer-generated reports promise an in-depth inspection of years of online activity, boiled down to a single digit: an intoxicatingly simple solution to an impractical task.
If the algorithms actually predicted risk better and were less biased than typical employers there might be an advantage to this: your social media would be shared with a faceless US company rather than your potential employer, so there might be less actual privacy invasion — after all, some faceless US companies already have your social media. It might be less embarrassing than your boss knowing what your favourite member of the appropriate sex calls you. The computer could also be set up to ignore irrelevant information like whether you talk about your sexual orientation online. With the setup as it is, that’s not the case, and one of the biggest risks is the completely unfounded appearance of both accuracy and objectivity — “mathwashing” as the jargon puts it.
3.
Where she lives, “100 per cent of the parents are going to want to use this,” she added. “We all want the perfect babysitter.”
One of the significant risks of automating human judgement is that it can go viral. There’s a limit to how well ordinary human prejudices can scale — you’ve got some chance of finding someone who has different biases. The prejudices of one computer checklist, though, can keep someone completely out of an employment sector.
Yesterday and today we had talks by our postgraduate students about their short research projects
From Axios
The wealth of the world’s billionaires rose by $1.4 trillion in 2017, the largest annual increase ever.
The details: Nearly all of that increase was driven by the Asia-Pacific region, and specifically China, where billionaire wealth rose 39%.
The graph is not well designed for illustrating the claim about billionaires: in a stacked bar chart like this it is easy to compare the first level, the sum of the first two, and the sum of all three, but not the top level. If the chart had been stacked the other way up it would be more obvious that it doesn’t seem to go along with the claim.
Measuring the height of the bars in MacOS Preview (because it’s not possible to read the graph to high enough precision) I get 35 pixels increase in the Asia-Pacific region and 46 in the rest of the world, so the growth in the Asia-Pacific region actually is less than half the total growth. Estimating the total growth by counting pixels does agree with the $1.4 trillion total, so what has gone wrong?
Clicking through to the UBS/PwC press release gives a bit more detail:
Chinese billionaires increased in number to 373 in 2017 from 318 in 2016 and their wealth rose by 39 percent to USD 1.12 trillion
So, the net increase in total wealth for Chinese billionaires is about $0.44 trillion. That hardly qualifies as “nearly all” of the increase for the Asia-Pacific region, let alone for the world
The original source is the the UBS/PwC Billionaires report. Their graph is better (apart maybe from the second y-axis). Not only is it the right way up for looking at increases in Asia vs the rest of the world, but it’s got tick marks on the right-hand axis where they’re actually useful. And it shows some historical context.

Normally I wouldn’t go to this much effort for a business news item — but the Axios Edge newsletter where I read this is edited by Felix Salmon, who is usually better at the difference between “nearly all” and “nearly half”
P.S. Yes, I did think of titling this post Crazy Rich Asians. No, I’m not sorry I didn’t.



I see from Twitter that it’s World Restart A Heart Day, encouraging people to learn CPR. Which you should do. It’s not hard. Nowadays there are even Spotify playlists of songs with the right beat for chest compressions.
However, two statistical points:
Recent comments on Thomas Lumley’s posts