Reading a study without a statistics degree
You do not need to follow the maths to read a health study well. Almost everything that matters sits in three things: how big the effect was, how uncertain that number is, and what the study was actually able to compare.
A headline says a habit “raises your risk by 40%”. The paper underneath it usually says something narrower, more careful and more useful. The gap between the two is not really a statistics problem — it is a reading problem, and a handful of habits close most of it.
Start with the effect size, not the p value
The effect size is the answer to the question you actually asked: how much difference was there between the groups? Ten percentage points? Two kilograms? Half a point on a scale nobody outside the field has heard of? Find that number first, before anything about significance, because it is the only one that tells you whether the finding would matter if it were true.
Be careful about which kind of number you are being shown. “A 40% increase in risk” is a relative figure, and relative figures are unmoored from how common something is. If a condition affects 2 people in 1,000, a 40% increase takes it to roughly 3 in 1,000. That is a real increase and it may be worth acting on at population scale, but it is a very different sentence from the one most readers construct in their heads. Good papers report the absolute numbers as well; if you cannot find them, that is worth noticing.
The confidence interval is the honest part
Every estimate from a sample is a guess about a population, and the confidence interval is the range of values the data are reasonably compatible with. A wide interval means the study could not pin the answer down. A narrow one means it could. Reading the interval instead of the single number is the fastest upgrade available to a non-statistical reader.
View the data as a table
| Row | Estimate | 95% interval | What you can say |
|---|---|---|---|
| Small study | +4.0 | −2.0 to +10.0 | Anything from a small harm to a large benefit |
| Large study, same estimate | +4.0 | +2.5 to +5.5 | A real difference, and big enough to act on |
| Very large study | +0.4 | +0.1 to +0.7 | Real, and almost certainly too small to matter |
| Small study, big estimate | +7.0 | −1.0 to +9.5 | Suggestive at best; the study was too small to tell |
“Significant” is a threshold, not a verdict
Statistical significance answers one narrow question: if there were genuinely no effect, how surprising would data like these be? The conventional cut-off of p < .05 is a convention, not a law of nature, and it says nothing about whether an effect is large, useful, or real in the everyday sense. Rows two and three of the figure are both significant. Only one of them is worth a headline.
It runs the other way too. A result that misses the cut-off has not been shown to be absent — often the study was simply too small to tell, which is exactly what row four looks like. Treating “not significant” as “no effect” is one of the most common misinterpretations in the literature, and statisticians have been complaining about it, loudly and in public, for years.
A p value tells you how surprised to be. It does not tell you how much anything changed, or whether the change is worth doing something about.
Then ask what the study could actually see
Statistics summarise the data. They cannot repair the design. Three questions cover most of it:
- Who was compared with whom? If the groups differ in ways other than the exposure — age, income, how sick they were to begin with — some of the difference belongs to those things. Adjustment helps, but it only handles what was measured.
- Which came first? A study that measures everything at one moment can show that two things travel together. It cannot show which one moved first, and that limits what any number in it can mean.
- Who was left out? People who volunteer for studies are healthier, better educated and more securely employed than the population. That does not make a study wrong; it makes it about a narrower group than the headline implies.
None of this requires arithmetic. It requires reading the methods section for two minutes, which is roughly the whole trick.
A short checklist
- What is the effect size, in units I understand?
- Is it absolute or relative — and what is the underlying rate?
- How wide is the confidence interval, and what is at each end of it?
- What design was this, and what could it not rule out?
- Who was studied, and are they anything like the people the headline is about?
Answer those five and you will read most health research more carefully than the article reporting it did.
Where this comes from
- Greenland S, Senn SJ, Rothman KJ, et al. Statistical tests, P values, confidence intervals, and power: a guide to misinterpretations. European Journal of Epidemiology. 2016;31(4):337–350. doi:10.1007/s10654-016-0149-3
- Wasserstein RL, Lazar NA. The ASA statement on p-values: context, process, and purpose. The American Statistician. 2016;70(2):129–133. doi:10.1080/00031305.2016.1154108
- Amrhein V, Greenland S, McShane B. Scientists rise up against statistical significance. Nature. 2019;567(7748):305–307. doi:10.1038/d41586-019-00857-9