Causal Inference Basics

Propensity score matching

The propensity score compresses every confounder into one number — the probability of receiving treatment — so each treated unit can be paired with an untreated twin.

Read these first

On this page 5
  1. Why it exists
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Propensity score matching compares each treated person against an untreated "twin" — someone who was equally likely to get the treatment but happened not to.

Two rice merchants want to compare their crops fairly. They cannot swap fields, but they can pair up plots. This irrigated lowland plot of mine against that irrigated lowland plot of yours. My dry hillside plot against your dry hillside plot. Pair by pair, the growing conditions match — so the remaining difference speaks about the seed, not the soil.

Matching does this with people. For each customer who took the offer, find a non-taker who looks the same on everything that mattered. Compare within pairs, average, done.

Why it exists

The like-with-like idea from confounding hits a wall as confounders multiply. Stratify by age — ten bands, fine. By age, income, city, device and tenure — thousands of cells, most containing one person or none. Exact twins do not exist in high dimensions.

The rescue is a 1983 discovery: you do not need to match on everything. Matching on one number — the propensity score, each person's probability of receiving treatment given their characteristics — balances all the measured characteristics at once. The impossible thousand-cell matching collapses into pairing on a single scale.

How it works

each person's traits  →  model  →  chance they get treated (0 to 1)

treated, score 0.82   ↔   untreated, score 0.81     paired
treated, score 0.45   ↔   untreated, score 0.46     paired
treated, score 0.97   ↔   nobody close             unmatched — set aside

effect ≈ average gap inside the pairs

The unmatched case matters. A treated person whose score is 0.97 has no untreated counterpart — everyone like them got treated. No honest comparison exists for them, and matching's quiet virtue is saying so instead of extrapolating.

A real example you have seen

Did attending a coaching institute raise a student's rank, or did coaching students differ from the start? Nobody will randomise children's futures to find out. Studies instead match each coaching student with a non-coaching student of similar marks, school type, city and family background, and compare ranks within pairs. Medicine does the same daily — smoking studies, vaccine safety in groups no trial covered.

Remember this

  • The propensity score is one number: the probability of getting treated, given your traits.
  • Matching on that one number balances all measured traits at once.
  • People with no counterpart get set aside — visible honesty, not a flaw.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install numpy scikit-learn

Verified with numpy 1.26.4 and scikit-learn 1.7.2.

Naive versus matched, on data where we know the truth

A premium card offer truly adds 40 rupees of monthly spend — but it was offered mostly to rich, urban customers, who spend more anyway.

matching.py
import numpy as np
from sklearn.linear_model import LogisticRegression

rng = np.random.default_rng(12)
n = 6000
income = rng.normal(0, 1, n)               # standardised income score
city = (rng.random(n) < 0.5).astype(float)
# richer, urban customers get offered the premium card more often
p_offer = 1 / (1 + np.exp(-(1.2 * income + 0.8 * city - 0.5)))
offer = rng.random(n) < p_offer
spend = 500 + 300 * income + 100 * city + 40 * offer + rng.normal(0, 50, n)

naive = spend[offer].mean() - spend[~offer].mean()
print(f"naive comparison: {naive:+.1f} rupees  (truth: +40)")

X = np.column_stack([income, city])
ps = LogisticRegression().fit(X, offer).predict_proba(X)[:, 1]

# match each offered customer to the nearest un-offered propensity score
c_ps, c_spend = ps[~offer], spend[~offer]
nearest = np.abs(c_ps[None, :] - ps[offer][:, None]).argmin(axis=1)
matched = (spend[offer] - c_spend[nearest]).mean()
print(f"matched estimate: {matched:+.1f} rupees")
Output
naive comparison: +334.3 rupees  (truth: +40)
matched estimate: +42.4 rupees

The walkthrough

The naive estimate is +334 — eight times the truth. Income and city stack the offered group with big spenders. Any executive shown that raw number would fund the card programme for the wrong reason.

Logistic regression estimates the score. predict_proba gives each customer's offer probability from their traits — the compression of two confounders (or two hundred) into one number. Note the model's target is treatment, not spend: propensity models predict who gets treated.

The matching line is a nearest-neighbour search. For each offered customer, argmin over absolute score distance finds the closest non-offered customer. Matching with replacement (one control can serve several treated) keeps bias low at some variance cost. The pair differences average to +42.4 — the truth, within noise.

This estimates the effect on the treated (ATT). Every offered customer got a twin; non-offered customers far from any treated profile were never asked to stand in. Compare with inverse probability weighting, which reweights everyone and targets the population effect.

The unwaivable caveat: matching balances measured traits only. A hidden confounder — ambition, say — passes straight through. The no-unmeasured-confounding assumption from confounding is doing all the load-bearing here.

Common mistakes

Skipping the balance check. After matching, verify treated and matched-control groups actually align on each covariate (standardised mean differences below 0.1 is the convention). Matching that fails to balance is decoration.

Matching without a caliper. A treated unit's "nearest" control can still be far away. Impose a maximum score distance (a caliper, commonly 0.2 standard deviations of the logit score); discard treated units with no control inside it — and report how many.

Judging the propensity model by classification accuracy. A propensity model that predicts treatment too well is a red flag, not an achievement — near-separable treatment means near-zero overlap, and matching degenerates. The model's job is balance, not model evaluation glory.

Treating matched estimates like experimental ones. The output deserves the sentence "assuming all confounders were measured". Sensitivity analysis (how strong a hidden confounder would overturn this?) belongs in the report.

Try it yourself

Add an unmeasured confounder: generate ambition, let it raise both p_offer and spend, but exclude it from X. Watch the matched estimate drift from +40 — the gap is the hidden confounding. Then add it to X and watch the estimate return.

What to learn next

Researcher — Mathematics and papers.

The balancing property

Define $e(x) = \Pr(T = 1 \mid X = x)$. Rosenbaum and Rubin (1983), The central role of the propensity score in observational studies for causal effects (Biometrika), prove:

  1. Balance: $X \perp T \mid e(X)$ — within a score stratum, covariates are distributed identically across arms.
  2. Sufficiency: if $(Y(0), Y(1)) \perp T \mid X$ (ignorability) and $0 < e(X) < 1$ (overlap), then $(Y(0), Y(1)) \perp T \mid e(X)$.
  • $e(x)$ — the propensity score; $X$ — measured covariates; $T$ — treatment.

Conditioning on the scalar $e(X)$ therefore suffices where conditioning on all of $X$ would — the dimension-reduction theorem that makes high-dimensional adjustment feasible.

Estimators and their properties

Nearest-neighbour matching on $\hat{e}(X)$ (with replacement, caliper) estimates the ATT. Known results:

  • Matching estimators are generally $\sqrt{n}$-consistent but not efficient; with a fixed number of matches, a bias term of order $O(n^{-1/k})$ from imperfect matches appears in $k$ continuous covariates (Abadie and Imbens, 2006, Econometrica) — matching on the scalar score sidesteps the dimension but inherits first-stage estimation error.
  • Standard errors must account for the estimated propensity and the matching step; Abadie and Imbens (2016, Econometrica) give the correction. Bootstrap is invalid for fixed-k nearest-neighbour matching (Abadie and Imbens, 2008).
  • Alternatives on the same score: stratification into quintiles (Rosenbaum and Rubin, 1984), kernel matching, full matching (Hansen, 2004) — trading bias against variance along familiar lines.

The King–Nielsen critique and modern practice

King and Nielsen (2019), Why propensity scores should not be used for matching (Political Analysis), show PSM can increase imbalance as pairs are pruned (the "PSM paradox"), because the score discards within-stratum covariate information. Responses in practice:

  • Prefer covariate-based matching — Mahalanobis distance, coarsened exact matching (Iacus, King and Porro, 2012) — or genetic matching (Diamond and Sekhon, 2013) when feasible.
  • Prefer weighting (IPW, overlap weights — Li, Morgan and Zaslavsky, 2018, JASA) or doubly robust estimators (AIPW, TMLE) for efficiency and cleaner asymptotics.
  • Whatever the method, the balance table is the deliverable; the estimator is secondary to demonstrated overlap and balance (Imbens and Rubin, 2015, chapters 13–15).

Machine-learned propensities (gradient boosting, see XGBoost) improve fit but can sharpen scores toward 0/1, shrinking overlap; calibration and trimming matter more than accuracy — the theme continued in double machine learning.

What to learn next