Causal Inference Basics

Potential outcomes and counterfactuals

Every person has two possible futures — with and without the treatment — but reality shows you only one, and all of causal inference is about coping with the missing half.

On this page 5
  1. Why it exists
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

A potential outcome is what would happen to one specific person under one specific choice. Every person carries one such outcome for each choice, and you only ever get to see one.

Two roads lead from your house to the station. Each morning you take one, and you experience that road's traffic. The other road's traffic that same morning existed — you were not on it to see it. The road not taken is the counterfactual: the outcome of the choice you did not make.

The effect of choosing road A over road B, for you, this morning, is the difference between those two travel times. One is real, one is forever invisible. That invisibility is the deepest problem in this whole section.

Why it exists

Arguments about "did it work?" go in circles without this idea. A patient takes a tonic and recovers in six days. Did the tonic work? The honest answer needs the counterfactual: how long would this same patient have taken without it? Five days — the tonic hurt. Ten — it helped.

The potential-outcomes framing turns fuzzy debates into a precise bookkeeping problem: two columns per person, one always missing. Statisticians call the missing one the fundamental problem of causal inference.

How it works

              recovery days       recovery days
patient       WITH tonic          WITHOUT tonic       effect
Asha              6                    9                -3
Vikram            8                    7                +1
Meena             5                    9                -4

reality shows one column entry per person, never both

You can never fill one person's whole row. But you can aim at the average effect across many people. If a coin decides who takes the tonic, the takers are — on average — interchangeable with the refusers. The refusers' average stands in for the takers' missing column. This trick is the entire justification for randomised experiments.

A real example you have seen

"Would I have got the job if I had worn the other shirt?" is a counterfactual your brain runs daily. So is every "player X would have scored if the captain had not declared" argument in cricket. Humans reason counterfactually by instinct. Causal inference is that instinct with bookkeeping strict enough to trust.

Remember this

  • Each person has a potential outcome per choice; reality reveals exactly one.
  • An individual's effect is the difference between their two columns — never observable directly.
  • Randomisation makes the groups stand in for each other's missing column, on average.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install numpy

Verified with numpy 1.26.4.

Playing god with the full table

A simulation can write down both columns — then we can watch which real-world procedures recover the truth and which do not.

potential_outcomes.py
import numpy as np

rng = np.random.default_rng(4)
n = 10_000
# the god view: BOTH outcomes for every patient. Real data never shows this.
y0 = rng.normal(60, 8, n)                   # recovery days without the drug
y1 = y0 - rng.normal(5, 2, n)               # with the drug: about 5 days faster
print(f"true average effect: {(y1 - y0).mean():+.2f} days")

# a randomised trial: a coin decides who we observe under the drug
coin = rng.random(n) < 0.5
rct = y1[coin].mean() - y0[~coin].mean()
print(f"randomised estimate: {rct:+.2f}")

# self-selection: the sickest patients seek out the drug themselves
seek = y0 + rng.normal(0, 4, n) > 62
naive = y1[seek].mean() - y0[~seek].mean()
print(f"self-selection estimate: {naive:+.2f}  <- the drug looks harmful")
Output
true average effect: -5.01 days
randomised estimate: -5.16
self-selection estimate: +6.35  <- the drug looks harmful

The walkthrough

y0 and y1 exist side by side for everyone — the luxury reality never grants. The true average effect, -5.01 days, is computed from the full table. Everything after is about approximating it while seeing half.

The randomised estimate lands at -5.16. The coin knows nothing about the patients, so takers and refusers match on average. y0[~coin] is a fair stand-in for the takers' hidden y0. The small gap from -5.01 is sampling noise, shrinking with n.

Self-selection flips the sign: +6.35. Sicker patients (large y0) sought the drug, so the treated group started worse. Their observed outcomes are higher despite the drug helping each of them. A helpful drug shows up as harmful — not a small distortion, a reversed conclusion. The anatomy of this failure gets its own lesson: confounding.

Look at what the estimator compares: y1[seek] against y0[~seek] — treated people's treated column against different people's untreated column. The comparison is only fair when the two groups of people are interchangeable.

Common mistakes

Speaking of "the" effect of a treatment as one number for everyone. The table has a row per person; effects vary (Vikram's +1 against Meena's -4). Averages hide this — and uplift modelling exists precisely to find whose row differs.

Treating before/after as with/without. A patient's recovery after the tonic versus their state before mixes the tonic with time, weather and regression to the mean. The counterfactual is "the same moment, other choice" — not "an earlier moment".

Forgetting SUTVA. The two-column bookkeeping silently assumes my treatment does not touch your outcome. Vaccines, marketplaces and social feeds break this — see interference.

Trusting the naive comparison because n is huge. Selection bias is structural. With n = 10_000_000, the naive estimate above is still near +6, only with tighter error bars around the wrong answer.

Try it yourself

Make the drug's benefit depend on the patient: y1 = y0 - rng.normal(5, 2, n) * (y0 > 60) — it only helps the badly ill. Recompute the true average effect and the effect among those who chose the drug. The two now answer different questions; say which question a hospital deciding a purchase should ask.

What to learn next

Researcher — Mathematics and papers.

The Rubin causal model

For unit $i$ and binary treatment $T_i$, posit potential outcomes $Y_i(1), Y_i(0)$; the observed outcome is $Y_i = T_i Y_i(1) + (1 - T_i) Y_i(0)$ (the consistency assumption). Common estimands:

$$ \text{ATE} = \mathbb{E}[Y(1) - Y(0)], \qquad \text{ATT} = \mathbb{E}[Y(1) - Y(0) \mid T = 1] $$

  • ATE — the average treatment effect over the whole population.
  • ATT — the effect among the treated: what treatment did for those who received it.

The naive contrast decomposes as

$$ \mathbb{E}[Y \mid T{=}1] - \mathbb{E}[Y \mid T{=}0] = \underbrace{\text{ATT}}_{\text{wanted}}

  • \underbrace{\mathbb{E}[Y(0) \mid T{=}1] - \mathbb{E}[Y(0) \mid T{=}0]}_{\text{selection bias}} $$

The self-selection simulation makes the second term large and positive (sicker patients treated), overwhelming a negative ATT.

Identification assumptions

Observational identification of the ATE from $(Y, T, X)$ requires:

  1. Ignorability / exchangeability: $(Y(0), Y(1)) \perp T \mid X$ — treatment is as-good-as-random within covariate strata.
  2. Positivity / overlap: $0 < \Pr(T = 1 \mid X = x) < 1$ for all relevant $x$.
  3. Consistency and SUTVA: well-defined treatment versions; no interference (Rubin, 1980).

Randomisation delivers (1) unconditionally and (2) by design — the precise sense in which experiments are privileged. Every later lesson in this section is a strategy for earning (1) without a coin: adjustment, propensity scores, instruments, discontinuities, parallel trends.

Causal inference as missing data

Half the potential-outcome matrix is missing by construction, with missingness mechanism $T$. Randomisation makes it missing-completely-at-random; ignorability makes it missing-at-random given $X$ — importing Rubin's missing-data taxonomy (Rubin, 1976) wholesale. Imputation-flavoured estimators (outcome regression / g-computation) and weighting-flavoured ones (IPW) are the two classical responses, unified by doubly robust estimators.

History and reading

  • Neyman (1923, translated 1990) — potential outcomes for randomised agricultural trials.
  • Rubin (1974), Estimating causal effects of treatments in randomized and nonrandomized studies, J. Educational Psychology — the general framework.
  • Holland (1986), Statistics and causal inference, JASA — names the fundamental problem; the standard survey.
  • Imbens and Rubin (2015), Causal Inference for Statistics, Social, and Biomedical Sciences — the graduate reference.

What to learn next