In the News
The Bradford Hill Criteria: 9 Questions That Separate Correlation From Cause
The Bradford Hill Criteria: 9 Questions That Separate Correlation From Cause

We’ve all come across a number of scientific headlines that state many different things. Sunlight increases the risk of cancer. Childhood burns increase the risk of melanoma development. Total sunlight exposure is linked to lower all-cause mortality.
They cannot both be settled, and the reason has nothing to do with which journal published which paper. An association is not a cause. Almost no health journalism makes that distinction, and a surprising amount of published research does not make it either.
In 1965, a statistician named Austin Bradford Hill gave the world a tool to stop failing this scientific literacy litmus test. He laid out nine considerations for deciding whether an observed pattern reflects something real and causal, or whether it’s just two things that happen to move together. You don’t need a PhD to use this framework. You need about ten minutes and a willingness to ask one more question before believing the headline.
Here are the nine, explained the way you’d actually use them.

Strength
How big is the effect?
A weak association is easy to manufacture. Bias in how a study was designed, a confounding variable nobody measured, a bit of statistical noise, any of these can produce a relative risk of 1.2 or 1.4. Almost nothing can produce a relative risk of 20 by accident.
Smoking raises lung cancer risk roughly 20-30 fold. That’s enormous. To explain it away, you would need a hidden confounder that is itself powerfully linked to both smoking and lung cancer, and so such thing exists.
When you see a study reporting a 15% increase in risk, that is a small effect sitting comfortably inside the range that study design flaws produce on their own. Ask how big before you ask anything else.
Consistency
Does the finding repeat?
One study is a lead, not a conclusion. The question is whether different research teams, in different countries, studying different populations, using different methods, keep landing in the same place.
Consistency matters because every study has its own specific weaknesses. A study in Sweden has different flaws than a study in Japan. When both find the same thing anyway, the flaws are unlikely to be the explanation.
The reverse is equally informative. If a finding shows up in one dataset and refuses to replicate elsewhere, that failure is data. Treat it as such.
Specificity
Does the exposure point to one outcome?
Hill’s original idea was that a genuinely causal exposure tends to produce a specific disease rather than a scattershot of unrelated effects.
This is the weakest of the nine, and Hill said so himself. Plenty of real causes do many things at once. Smoking causes lung cancer, heart disease, stroke, emphysema, and bladder cancer. Specificity strengthens a causal case when it is present. Its absence proves nothing.
Learn this one so you know when to ignore it.
Temporality
Did the cause come first?
This is the one criterion that is genuinely required. if the exposure did not preceded the outcome, there is no causal claim to evaluate. Full stop.
It sounds too obvious to state, and it fails constantly in practice. Studies find that depressed people exercise less and conclude that inactivity causes depression. The arrow may well run the other way. Studies find that people with a disease report different past habits, but the reporting happens after diagnosis, and a diagnosis reshapes memory.
Whenever you read a study, find the moment the exposure was measured and the moment the outcome appeared. If they were measured at the same time, you are looking at a snapshot. Snapshots cannot establish direction.
Biological Gradient
More exposure, more effect?
A dose-response relationship is powerful evidence. Two cigarettes a day carries less risk than twenty, which carries less risk than forty. The graded relationship is hard to explain by chance or bias.
Two cautions. Some real causal relationships have thresholds, where nothing happens until a certain dose and then everything happens. Others are U-shaped, where too little and too much are both harmful. And a dose-response curve can appear when the underlying driver is a confounder that happens to track with the exposure.
A gradient makes a causal claim stronger. It doesn’t close the case.
Plausibility
Is there a believable mechanism?
Can you tell a coherent biological story connecting the exposure to the outcome? Something the body actually does to produce the result?
Hill attached a warning to this one. Plausibility is limited to what science currently knows. John Snow traced cholera to a contaminated water pump decades before anyone understood germ theory. The mechanism was missing, but the causal claim was correct.
The opposite failure is more common today. Plausibility is the easiest criterion to satisfy, because a competent biologist can construct a plausible mechanisms for almost any association. A mechanism being plausible does not make it the dominant one, or even a real one.
Coherence
Does it fit everything else we know?
A causal claim has to live alongside the rest of the evidence without contradiction. If the mechanism is real, the population data should reflect it. If exposure has risen sharply, disease should follow. If exposure has fallen, disease should decline.
This is where a lot of confident claims fall apart. The laboratory story is elegant, the population trends point somewhere else, and nobody reconciles the two. When incidence and mortality move in opposite directions, or when a supposed cause has been declining for 30 years while the disease climbs, the story has a hole in it.
Coherence is the criterion that catches contradictions the individual studies were never designed the notice.
Experiment
Remove the exposure. Does the outcome change?
This carries the most weight, because it’s the closest thing to an actual test rather than an observation.
Sometimes this means a randomized controlled trial. Often it means a natural experiment like a policy change, a ban, a supply disruption, a population that moved. Smoking bans were followed by measurable drops in hospital admissions for heart attacks. That’s the causal model making a prediction and the prediction coming true.
When an intervention has been widely adopted for decades and the disease hasn’t moved, the causal model made a prediction and the prediction failed. Take that seriously.
Analogy
Have we seen something like this before?
If one drug taken in pregnancy caused birth defects, we are willing to entertain the idea that a chemically similar drug might do the same. Prior example lower the bar for taking a new claim seriously.
This is the softest of the nine. Analogies are chosen by the person making the argument, and a well-chosen analogy can make almost anything sound reasonable. Use it as a sanity check and nothing more.
How To Use This
Hill’s closing point is the one most people skip. He wrote that none of the nine can be demanded as absolute proof, and that they are not a checklist to be scored and totaled. They are a structure for judgment.
Weight them accordingly. Temporality is required. Strength, consistency, and experiment carry real force. Gradient and coherence are strong supporting evidence. Plausibility, specificity, and analogy are the light end of the scale.
A solid causal claim holds up across most of these. A claim that satisfies one or two while straining against the rest is an association wearing a costume.
The practical version is a handful of questions you can ask about any health headline in under a minute. How big is the effect. Has anyone else found it. Which came first. Is there a dose curve. Does it fit the population data. Has anyone tested it by intervening.
Most of what gets reported as settled science does not survive those six questions. That is worth knowing before you rearrange your life around a press release.



