Confounding
A confounder is a hidden third factor that drives both the treatment and the outcome, manufacturing a link between them — or hiding a real one.
- 7 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
A confounder is a third factor that pushes on both the thing you changed and the thing you measured. It creates a connection that is not really there.
People who carry matchboxes get lung cancer far more often than people who do not. Matchboxes do not cause cancer. Smokers carry matchboxes, and smoking causes cancer. Smoking is the confounder — the common cause standing behind both sides of the pattern, pulling the strings.
Once you see this shape, it appears everywhere. Kids with bigger shoe sizes read better (age drives both). Towns with more temples have more crime (population drives both).
Why it exists
Any comparison between groups that formed themselves is exposed to confounding. Supplement takers versus non-takers. Users of a feature versus non-users. Cities that adopted a policy versus cities that did not. In each case, something decided who ended up in which group. Whatever did the deciding is a suspect.
This is the precise reason the previous lesson's naive estimate flipped sign: sickness decided who took the drug, and sickness drove the outcome.
How it works
smoking (confounder)
↙ ↘
matchbox lung cancer
✗ ─ ─ ─ ✗ no arrow here — yet the data shows a strong linkThe main defence is adjustment: compare like with like. Split people into groups where the confounder is the same — smokers with smokers, non-smokers with non-smokers — and compare matchbox-carriers within each group. Inside a group, the confounder cannot vary, so it cannot pull strings. Then average the within-group answers.
The defence has a hard limit: you can only adjust for confounders you measured. The one you never thought of stays in the data, pulling.
A real example you have seen
For years, studies showed coffee drinkers had more heart disease. Later work found the confounder: smoking travelled with coffee (the chai-and-cigarette pattern exists worldwide). After adjusting for smoking, coffee's "harm" largely evaporated. Headlines about food causing or curing disease are, more often than not, confounding stories in disguise.
Remember this
- A confounder is a common cause of treatment and outcome.
- It can invent a link, hide one, or reverse its direction.
- Adjustment means comparing like with like — and only works for confounders you measured.
What to learn next
- Causal DAGs and what to control for — the map that says what to adjust and what to leave alone.
- Propensity score matching — like-with-like at scale, when strata multiply.
- Simpson's paradox — confounding strong enough to reverse every table it touches.
Developer — Code and libraries.
Setup
pip install numpyVerified with numpy 1.26.4.
Watching age poison a comparison — and fixing it
A supplement genuinely speeds recovery by 3 days. Older people take it more often and recover slower. Watch what the raw comparison says.
import numpy as np
rng = np.random.default_rng(2)
n = 20_000
age = rng.integers(20, 70, n).astype(float)
takes = rng.random(n) < (age - 15) / 70 # older people take the supplement more
recovery = 30 + 0.4 * age - 3.0 * takes + rng.normal(0, 3, n) # true effect: -3 days
naive = recovery[takes].mean() - recovery[~takes].mean()
print(f"naive comparison: {naive:+.2f} days <- looks harmful")
# adjust: compare like with like inside 5-year age bands, then average
diffs, sizes = [], []
for lo in range(20, 70, 5):
band = (age >= lo) & (age < lo + 5)
if takes[band].any() and (~takes[band]).any():
diffs.append(recovery[band & takes].mean() - recovery[band & ~takes].mean())
sizes.append(band.sum())
adjusted = np.average(diffs, weights=sizes)
print(f"within-age comparison: {adjusted:+.2f} days (truth: -3.00)")naive comparison: +1.86 days <- looks harmful within-age comparison: -2.99 days (truth: -3.00)
The walkthrough
The naive comparison says +1.86 — a helpful supplement looks harmful. Takers skew old (the assignment line makes probability rise with age), and age adds 0.4 days per year. The age gap between the groups injects more than 4 days of bias, swamping the -3 benefit.
Stratification recovers -2.99. Inside one 5-year band, takers and non-takers have nearly the same age, so age cannot tilt the scales. Averaging the ten within-band differences, weighted by band size, rebuilds the population answer. This is the oldest adjustment method and still the most transparent.
The same logic wears many coats. A regression of recovery on takes and age performs a smooth version of this stratification — see linear regression. Propensity matching and inverse weighting are two more coats for the same body: compare like with like.
What no method here can do: rescue you from an unmeasured confounder. Delete age from the dataset and the bias is invisible and unfixable from this data alone. That is why the strongest designs — instruments, discontinuities, experiments — earn their complexity.
Common mistakes
Adjusting for everything available. Some variables create bias when adjusted for — mediators and colliders, coming in the next two lessons. Adjustment needs a causal map, not a kitchen sink.
Declaring victory after adjusting for whatever was measured. Age, gender, city — measured and adjusted. Motivation, severity, wealth — unmeasured, still pulling. Observational conclusions deserve the sentence "assuming no unmeasured confounding", said aloud.
Empty comparison groups. If no 65-year-old refuses the supplement, that band contributes nothing — and the method quietly extrapolates. The code's if guard skips such bands; real analyses must report them. This is the positivity problem, and it returns with force in inverse probability weighting.
Confusing confounders with predictors. A variable can predict the outcome strongly yet confound nothing (it never influenced treatment). Adjusting for pure outcome predictors is harmless — helpful, even, for precision — but it is not what removes bias.
Try it yourself
Delete age from the adjustment and instead adjust for a noisy proxy: proxy = age + rng.normal(0, 15, n). See how much bias survives — partial measurement gives partial protection. Then make the confounding negative (young people take the supplement more) and predict the naive estimate's direction before running.
What to learn next
- Causal DAGs and what to control for — the map that says what to adjust and what to leave alone.
- Propensity score matching — like-with-like at scale, when strata multiply.
- Simpson's paradox — confounding strong enough to reverse every table it touches.
Researcher — Mathematics and papers.
Definition and the adjustment formula
$X$ confounds the $T \to Y$ relation when $X$ causally affects both $T$ and $Y$ through paths not through $T$. Under ignorability given $X$ and positivity (defined in potential outcomes), the interventional mean is identified by standardisation (g-formula):
$$ \mathbb{E}[Y \mid do(T = t)] = \sum_x \mathbb{E}[Y \mid T = t, X = x]\, P(X = x) $$
- $do(T=t)$ — setting treatment by intervention.
- The inner term — the within-stratum mean; the outer sum — reweighting by the population distribution of $X$, not the conditional one.
The stratified estimator in the developer block is the empirical g-formula with discrete strata. Its regression, matching and weighting siblings estimate the same functional under the same assumptions — differing in variance, extrapolation behaviour and failure modes.
Omitted variable bias, quantified
In the linear model $Y = \beta T + \gamma X + \varepsilon$ with $\operatorname{Cov}(T, X) \ne 0$, omitting $X$ gives
$$ \hat{\beta}_{\text{short}} \xrightarrow{p} \beta + \gamma \, \frac{\operatorname{Cov}(T, X)}{\operatorname{Var}(T)} $$
The bias term is the product of two slopes: confounder→outcome ($\gamma$) and the regression of confounder on treatment. In the simulation, $\gamma = 0.4$ and treated-versus-untreated age gaps of ~12 years produce ~+4.9 bias against a true $\beta = -3$ — matching the observed +1.86.
Sensitivity analysis
Since no-unmeasured-confounding is untestable, quantify how strong a hidden confounder must be to overturn the result:
- Cornfield conditions (Cornfield et al., 1959) — the historical smoking-cancer bound: a confounder explaining the association must be at least as strongly associated with exposure as the observed risk ratio.
- Rosenbaum bounds (Rosenbaum, 2002, Observational Studies) — sensitivity parameter $\Gamma$ on odds of treatment for matched pairs.
- E-values (VanderWeele and Ding, 2017, Annals of Internal Medicine) — the minimum confounder-outcome and confounder-treatment risk ratio, on a shared scale, needed to explain away the estimate; computable from any risk ratio in one line.
- Partial $R^2$ framing (Cinelli and Hazlett, 2020, JRSS-B) — bias as a function of the confounder's partial variance explained in $T$ and $Y$; their
sensemakrpackage operationalises it.
Reading
- Hernán and Robins (2020), Causal Inference: What If, chapters 7 and 12 — confounding and its adjustment, free online.
- Greenland, Robins and Pearl (1999), Confounding and collapsibility in causal inference, Statistical Science — why "the association changed when I adjusted" does not itself prove confounding.
- VanderWeele (2019), Principles of confounder selection, European J. Epidemiology — practical selection rules ahead of the full DAG machinery.
What to learn next
- Causal DAGs and what to control for — the map that says what to adjust and what to leave alone.
- Propensity score matching — like-with-like at scale, when strata multiply.
- Simpson's paradox — confounding strong enough to reverse every table it touches.