Posts written by Thomas Lumley (2645)

avatar

Thomas Lumley (@tslumley) is Professor of Biostatistics at the University of Auckland. His research interests include semiparametric models, survey sampling, statistical computing, foundations of statistics, and whatever methodological problems his medical collaborators come up with. He also blogs at Biased and Inefficient

August 15, 2011

Why useful genetic testing is hard.

The NZ Herald reports (from a Cancer UK press release): A single genetic fault in a gene that normally helps the body to repair its DNA increases a woman’s risk of ovarian cancer six-fold, a study has found.   That sounds as if it might be useful as a way of detecting women at high risk, until you look at the numbers in more detail.

In the UK, where the study was conducted, about 6500 women per year are diagnosed with ovarian cancer, and 40-50 of these women will have the genetic fault.  New Zealand has about 15 times fewer people than the UK, so that would be about 3 women per year in the whole country who develop ovarian cancer because of this genetic variant.

 

More NZ government data

From the Dominion Post, via stuff.co.nz: The Cabinet will today issue an instruction to government departments that they should make all data they hold available for free or at a reasonable-price in accessible file formats for reuse by businesses, unless there are good reasons not to.

This policy has worked very well in the US, where the website data.gov, after just over two years’ of operation, has nearly 400,000 data sets, and there are 236 apps available on the web that make use of the data to do things that would otherwise be expensive or impossible.

 

August 12, 2011

Genetics and intelligence

There’s a new genetic study (stuff.co.nz has the Associated Press article) claiming that variation in intelligence is about 50% genetic.  That claim is not new, but previous studies got estimates like that by studying the IQ of close relatives,  and this study actually measures genes.  The researchers studied 3500 unrelated people in Scotland and England,   measuring half a million genetic variants on each person and relating them to two different types of intelligence test, and found that while they couldn’t identify specific genetic variants that affected intelligence, they did find evidence that there were hundreds of variants with some effect.  (more…)

August 10, 2011

Census time

The West Island had their census last night, so the Australian Bureau of Statistics is in the process of collecting millions of little bits of paper from around the country. It’s a good occasion to think about what the census is actually good for, because this goes further than you might think. (more…)

How do cohort studies work?

A very good radio documentary (BBC Radio 4) from Ben Goldacre of Bad Science fame.

August 8, 2011

The moon and earthquakes

A recurrent theme of this blog is that your taxes (or, even better, other people’s taxes) pay for the collection of a lot of high-quality data on an amazing range of topics, and you can just go and look it up.

For example, the US Geological Survey has a database of earthquakes that anyone can search.  It includes pretty much everything since 1973, and a more limited coverage back into prehistory.   John Walker’s web site Fourmilab has a range of interesting astronomical pages, including an earth-moon distance calculator.  

Combining these, we can look at the distance from the earth to the moon for all 62667 earthquakes of magnitude 5.0 or greater from 1973 up to early this morning.  Using a complete census this way it’s easy to avoid cherry-picking results that support whatever conclusion you want.  For convenience I translated the earth-moon distance calculator into R, so I could run it more easily on large batches of data. (more…)

August 7, 2011

Anatomy of a hoax

Last week, many newspaper websites (though apparently not any Kiwi ones) reported a study purporting to that users of Internet Explorer had lower IQs than users of other browsers, with IE version 6 users scoring 20 points lower than Firefox users, and more than 40 points lower than users of Opera. The results were supposed to be based on a survey of 100,000 people recruited through ads on websites.  This turns out not to be the case.

What makes the story interesting is how many reasons there were not to believe it. (more…)

August 2, 2011

Casual inference

From the NZ Herald:

“The survey found almost 65 per cent of women believed they were paid less because of their gender. Just under 43 per cent of men agreed but 47 per cent didn’t.”

Unfortunately the Herald doesn’t tell us what the actual question was. Were people asked whether they, personally, were paid differently because of their gender, or whether women, on average, were paid less because of their gender?   In either case, my sympathies are with Women’s Affairs Minister Hekia Parata, who refused to offer her own answer to the “simplistic” poll question.

There are two statistical problems here. The first is what we mean by “because of their gender”.  After that’s settled, we have the problem of finding data to answer the question.  Inference about cause and effect from non-experimental data will always be hard because of both problems, but that’s what we’re here for.

Usually, when we say that income or health is worse because of some factor, we mean that if you could experimentally change the factor you would change health or income. We say high blood pressure causes strokes, and we mean that if you lower the blood pressure of a bunch of people, fewer of them will get strokes.  This isn’t possible for gender — not only can we not assign gender at random, we can’t even say what it would mean to do that.  Would a female Dan Carter be his sister Sarah, or Irene van Dyk? (more…)

July 26, 2011

That trick never works.

Q: So, have you seen the article about Vitamin D and diabetes?

A: Of course. The tireless staff of StatsChat read even West Island newspapers. It’s a good report, too.

Q: What did the researchers do?

A: They studied 5200 people without diabetes, following them up for five years. 199 of them developed diabetes. The people who ended up with diabetes started off with lower vitamin D levels in their blood.

Q: Where did you get those details?

A: The abstract for the study publication (you can also get the full text there free if you’re at a university or if you wait until next year).

Q: Isn’t it annoying that newspaper websites don’t provide any links to that sort of information?

A: It’s like you’re reading my mind.

Q: One of the study authors is quoted as saying “”It’s hard to underestimate how important this might be.” What do you think?

A: I think he meant “overestimate”.

Q: So, how important is this finding?

A: If it really is an effect of vitamin D, it would be really important.  A simple supplement would be able to dramatically reduce the risk of diabetes.

Q: How can we tell?

A: Someone needs to do a randomized trial, where half the participants get vitamin D and half get a dummy pill. If the effect is real, fewer people getting vitamin D will end up with diabetes.

Q: That sounds like a good idea. Is someone doing a trial?

A: Yes, Professor Peter Ebeling, of the the University of Melbourne.

Q: Is there some useful website where I can find more information about the trial?

A: Indeed.

Q: Will it work?

A: No.

Q: Are you sure?

A: No, that’s why we need the trial.  But it’s a trial of vitamin supplementation, which almost always has disappointing results, and it’s a trial  in adult-onset diabetes, which almost always has disappointing results.

 

July 24, 2011

Every 48 seconds.

Last week, in Wellington, I saw an ad warning that someone is hurt in a household accident every 48 seconds. Is that a lot?

We need some numbers. The population of NZ is about 4.5 million, and there are about 30 million seconds in a year(*).

Since I’m doing this while trying to avoid traffic in the jaywalking capital of New Zealand, let’s round off the accident rate to one every 50 seconds. That gives 600,000 accidents per year. If we had a population of six million, this would be one accident every ten years, but we only have 4.5 million, so it’s more like one accident per person per seven years.

Somehow, it sounds less impressive that way.

 

* For real nerds, an even more accurate approximation is π times 10 million seconds.