Posts written by Thomas Lumley (2645)

avatar

Thomas Lumley (@tslumley) is Professor of Biostatistics at the University of Auckland. His research interests include semiparametric models, survey sampling, statistical computing, foundations of statistics, and whatever methodological problems his medical collaborators come up with. He also blogs at Biased and Inefficient

April 11, 2012

Tooth nuking

The Herald (and media sources worldwide) is covering a research paper on brain tumours and dental x-rays.  The paper asked roughly 1500 people with meningioma, and the same number of healthy people, about their histories of dental X-rays.  The people with meningiomas were more likely than the controls to report having X-rays at least annually, and the researchers estimated a relative risk of 1.5.

Now, meningioma is pretty rare, so this increase works out to an extra lifetime risk of maybe 5 cases for each 10,000 people.  Also, if you are going to have a brain tumour, meningioma is the one to have — some are not even diagnosed, and most diagnosed ones are treated successfully.  On the other hand, brain tumours are usually something you’d like to avoid, so is the risk real?

There are at least two issues that make the relative risk of 1.5 less plausible

  • Self-report of risk factors for cancer is notoriously unreliable
  • Since meningiomas can be relatively minor, the time of diagnosis varies, there might be some tendency for the sort of people who have regular dental x-rays to also be the sort of people who get earlier diagnoses, which would show up as a higher rate

The Science Media Centre also has a good summary, with quotes from experts.

It’s interesting to work out whether the risk increase is in the right ballpark given general knowledge about radiation.  A 1991 paper looked at the dose from different sorts of bite-wing dental X-ray setups, and found a range from 2 microSievert to 20 microSievert.   (XKCD shows what a microSievert means).   The same paper quotes an estimated risk of 0.73 health events including cancers per Sievert of dose to a population.   We don’t know what the population size was, but we can get a rough idea from the original paper.  They found 1500 meningiomas in 5 years, so at a rate of 3 per 100,000 people per year, that means about ten million people.  Roughly a third of the controls (and so roughly a third of the population) had at least yearly X-rays, so let’s suppose we are looking at 20 x-rays exposure on average for this third of people.Multiplying all the numbers together gives about 150 extra health events at 20 microSieverts per X-ray, or about 40 at the more-typical modern value of 5 microSieverts per X-ray.  The 1.5 relative risk that the researchers found is larger than this crude extrapolation would predict, but the order of magnitude is right.

So, there may well be a small increase in risk of a rare, mostly treatable brain tumour from having yearly dental x-rays.  It’s uncertain how big the risk is, and there are reasons to expect it might be less than a 1.5-fold increase, but that increase is at least of a plausible order of magnitude.   The radiation exposure from a dental x-ray is quite a bit less than from a trans-Tasman flight, and hugely less than from a CT scan, but it’s not zero.

April 9, 2012

‘Causal’ is not enough

Yesterday’s post about crime rates and liquor stores was tagged ‘correlation vs causation’, but it’s more complicated than that.  It’s not even clear what sort of causation is at stake.

I think we can all agree that being drunk, like being young and male, is a causal factor in violent crime. But that’s not the question.  There are two possible causal stories behind higher crime rates near liquor stores, or, more precisely, alcohol licenses.   These are truly causal alternatives to the skeptical argument that it’s actually (demand for) drinking that leads to alcohol licenses.

The weaker causal story is that people get drunk, and when they do, they are more likely to do it nearer to alcohol licenses.  That’s certainly the case for pubs and restaurants — if you buy beer from a pub, you are going to be drinking it at the pub  — and could be true for liquor stores as well.   This story would say that if you moved an alcohol license the crime would move, and if you shut down one place, the drunkenness and crime would relocate among the available options.  If this is true, it’s useful to local community groups wanting to improve local conditions, but it’s pretty much useless from a public health and safety viewpoint.

The stronger story is that people won’t drink if they have to go further to get alcohol, so that reducing the number of licenses will reduce drinking.   On this theory, reducing licenses could have a health and safety impact beyond just local redistribution of crime.

It’s not possible to distinguish these using the available data.  There’s good evidence that something like the first story holds for CCTV installation — it pushes crime out of the surveillance zone but doesn’t stop it.  And there’s some evidence that something like the second explanation works for stopping kids from smoking — adding inconvenience and cost has much more of an impact on them than on adults.

The future needs statisticians

The current issue of the journal Science has an editorial on the importance of statistics, and on the increased demand for statisticians in the `Big Data’ future.  The writers, Marie Davidian and Tom Louis, call out the need for increased funding in graduate programs — it hasn’t kept up with inflation, let alone with demand.

They also note

The future demands that scientists, policy-makers, and the public be able to interpret increasingly complex information and recognize both the benefi ts and pitfalls of statistical analysis. It is a good sign that the new U.S. Common Core K-12 Mathematics Standards introduce statistics as a key component in precollege education, requiring that students be skilled in describing data, developing statistical models,making inferences, and evaluating the consequences of decisions.

Here, at least, New Zealand is ahead of the game.

April 8, 2012

Statistical crimes double near liquor stories[updated]

Stuff has the  headline “Crime doubles close to liquor outlets”, based on an analysis from the University of Canterbury.  Now, can we think of possible non-headline explanations for this?  Indeed we can. As the story admits, near the end

The areas with the most serious violent crime had more Maori and young males, over-represented in crime statistics, and the highest population densities.

and

The three spikes with the highest numbers of liquor outlets were Auckland central (447 alcohol licences), Wellington central (423) and Christchurch central (394), all of which had high crime rates.

These numbers raise the question of what sort of alcohol licenses were included.  I’d be surprised if there were 447 liquor stores in Auckland Central, but if you include pubs and licensed restaurants the numbers look more plausible. If so, we’re not talking about liquor stores at all.  The fact that the three CBD areas (all places with bans on alcohol consumption in the street)  top the list also suggests that there’s a problem with denominators: since many of the people in the CBD don’t live there, rates of crime per 1000 population  will tend to be inflated.

What is really infuriating is that the researchers actually did a better version of the analysis, but we don’t get to see it. In the last paragraph of the story, we get

Day said the correlation was weaker, but still held, when those factors were statistically removed from the equation.

So why don’t we get told the numbers that at least have a chance of meaning something, rather than the “crime doubles”?

Updated to add:  A commenter on a later post gave a link to the published paper,  and the adjustment brings relative rates of 2.4, 2.0, and 2.4 for any license, on-license and off-license, respectively, to 1.5, 1.6, and 1.4.   Also, without adjustment there is a much higher rate in for the areas closest to off-license stores, but after adjustment the elevated rate is constant out to 5km, which seems much less plausible.

April 6, 2012

Looking under the lamppost

Stuff is reporting on new drug tests being pushed by NZDDA

Hardy said hair testing was more accurate and effective method of detecting drug use, and it gave a history of drug or alcohol use over the previous 90 days….With urine tests more drugs were undetectable if urinalysis was carried out more than three days after use.

Since the advertised purpose of employee drug testing is to catch people who are impaired on the job, expanding the history from 3 days to 90 days surely makes the test less accurate, not more accurate.  It’s more accurate only if you don’t distinguish between on-the-job and off-the-job drug and alcohol use — like the drunk looking for his keys under the lamppost because he could see better there.

One of the key contributions of statistics to evidence-based medicine has been in forcing medical researchers to measure what they really want to affect, not what is convenient and plausibly related to it.  Drug use in the past 90 days is not the same as on-the-job impairment, and it’s probably a pretty lousy surrogate.

The Assistant Privacy Commissioner is quoted in the article as saying

“Employers should only use it where there is a genuine business need. For example, drug testing has been allowed where there are safety issues with operating machinery.

An interesting approach used by some US companies is drug testing after accidents.  A study from Princeton (PDF) found that this did reduce accidents by about 10%, though some of the reduction may have been due to under-reporting of accidents — an important tradeoff to consider.

There isn’t going to be a quick technological fix, however, and we do need some sort of regulation. As Stanford’s Keith Humphreys puts it

…we use public policy to pick the particular sort of drug problem society will have. For example, different policy environments can make it a human rights problem, an addiction problem, a crime problem, an AIDS problem, a public disorder problem etc., but no policy will produce a true ending of all of society’s problems with drugs. There are some policies that ameliorate multiple aspects of the problem, but in most cases we are faced with hard choices about what sort of problem we will have rather than a problem-free alternative.

When in doubt, randomise.

This week, John Key announced a package of mental-health funding, including some new treatment initiatives.  For example, Whanau Ora will be piloting a whanau-based approach, initially on 40 Maori and Pacific young people.

It’s a pity that the opportunity wasn’t taken to get reliable evidence of whether the new approaches are beneficial, and by how much.  For example, there must be a lot more than 40 Maori and Pacific youth who could potentially benefit from Whanau Ora’s approach, if it is indeed better.  Rather than picking the 40 test patients by hand from the many potential participants, a lottery system would ensure that the 40 were initially comparable to those receiving the current treatment strategies.  If the youth in whanau-based care did better we would then know for sure that the approach worked, and could compare its cost and effectiveness, and decide how far to expand it.   Without a random allocation, we won’t ever be sure, and it will be a lot easier for future government cuts to remove expensive but genuinely useful programs, and leave ones that are cheaper but don’t actually work.

In some cases it’s hard to argue for randomisation, because it seems better at least to try to treat everyone.  But if we can’t treat everyone and have to ration a new treatment approach in some way, a fair and random selection is no worse than other rationing approaches and has the enormous benefit of telling us whether the treatment works.

Admittedly, statisticians are just as bad as everyone else on this issue.   As Andrew Gelman points out in the American Statistical Association’s magazine “Chance”, when we have good ideas about teaching we typically just start using them on an ad hoc selection of courses. We have, over fifty years, convinced the medical community that it is possible, and therefore important, to know whether things really work.  It would be nice if the idea spread a bit further.

April 2, 2012

Big data and Downton Abbey

The hit British TV series Downton Abbey has drawn some fire for alleged anachronisms: phrases that just don’t fit Georgian-era Britain.

Ben Schmidt has unleashed gigabytes of data on this problem, with the Google Books n-grams.  When Google digitized lots of books, it also tabulated the frequencies of words, pairs of words, triples of words, and so on, by year of publication. In two posts, Ben compares word pairs from the TV script with the Google frequencies for books published in the 1910s and the 1990s.   The comparison shows up several two-word phrases that were much less common in Downton Abbey’s historical period than they are now, but still appear in the script.  In some cases these phrases were not observed at all in written English until much later; in other cases they existed but were rare.

As a check on the process, he also looks at a genuine play from the period, George Bernard Shaw’s Heartbreak House, which passes the phrase test with flying colors.

April 1, 2012

Heart disease vaccine?

Prime News tonight (I don’t see how to link to an individual story there) reported on a ‘vaccine for heart disease’.  This is really exciting research from the Karolinska Insitute in Sweden, studying the role of the immune system in coronary artery plaque.  The previous belief was that damaged (oxidised) LDL cholesterol, which is a consequence of plaque, triggered the immune responses; the researchers showed that the immune response was to normal LDL cholesterol.  They also showed that a vaccine blocking the immune response led to reduction of white blood cell involvement and shrinkage of plaques in transgenic mice.

Prime News went on to say that a vaccine might be available within five years.  I hope this isn’t realistic. The initial human studies to show an effect on plaque could easily be done in that time, but not a trial that actually demonstrates reductions in heart attack rates.

There is lots of depressing experience in cardiovascular research with good ideas for treatments that affect a biological measurement related to heart disease, but don’t actually reduce the risk of heart attack or death, because something goes wrong.   The US FDA, who will be the primary group that researchers have in mind when designing trials, is fairly insistent on having actual evidence of clinical benefit from new treatments.  Their Cardiovascular & Renal advisory committee is one of the most tough-minded, ever since it relied on mere biological surrogates of benefit to make what was probably the worst drug approval decision in history: approving drugs to regulate heart rhythm without evidence that they actually prevented cardiac arrest.  They didn’t.

 

March 30, 2012

Powerball and the Kelly criterion

A popular decision rule for investment and other forms of gambling is the Kelly criterion, named after mathematician (and successful investor) John Kelly.  In the long run, following this rule will maximise long-run expected wealth.

If we assume that the story in Stuff about Wairarapa bettors having had a 2:1 return in the past year can be applied to tomorrow’s Powerball (which it can’t), we can look at what that would imply about rational betting.

The Kelly criterion specifies what fraction of your total wealth you should spend on an investment opportunity.  The fraction is always less than your probability of winning.  With 2:1 expected payoff and large odds, the recommended fraction is about half the probability of winning

The chance of the top Powerball prize (since this isn’t a ‘must win’ week) is 1 in 38 million for a $1 bet, so you should bet less than 1 dollar for each 76 million dollars of your current disposable wealth.   For most of us, that’s less than one dollar.

It’s worth noting that while not everyone supports the Kelly criterion, most of the critics suggest that you should bet less than the criterion recommends, not more.

(via a commenter at Cornell physics blog The Virtuosi)

Traffic congestion and data science

Recently, I mentioned the possibility of using bus timing data to probe congestion on Auckland roads.  This idea has been bypassed by Google, who now provide real-time congestion maps of New Zealand using smartphone location data.

If you run Google Maps or Google Navigation, you have the option of sending anonymous GPS-based location data to Google, so they know the locations of lots of phones.  By tracking the speed of phones that are moving along roads, they can work out the traffic speed, and measure congestion.   This is harder than it sounds — GPS accuracy on its own is not enough to distinguish phones in cars from phones carried by pedestrians — but using combined location and speed data they can even give separate congestion information in each direction on many roads.

For example, if you were coming to our public lecture on Tuesday, you might look on Google Maps and click on the “Traffic” label, and see that Symonds St is totally clogged, and decide to come up Grafton Rd instead.