Inverse probability weighting
IPW rebalances an unfair comparison by counting surprising cases more — each person is weighted by one over the probability of the group they landed in.
- 8 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Inverse probability weighting makes each rare, surprising case speak for all the similar people who went the other way — by counting it extra.
A survey team wants a village panchayat's true opinion but could reach only two farmers out of the village's hundred, alongside fifty townspeople. The fix is old and sensible: count each farmer's answer fifty times, each townsperson's once. Every hard-to-reach voice stands in for the many like it who went unheard.
IPW (inverse probability weighting) applies the same fix to treatment groups. A poor, rural customer who did get the premium offer is a rarity. So in the treated group's average, her outcome is counted many times. She stands in for all the similar customers the offer skipped.
Why it exists
Matching fixes unfair comparisons by finding twins and discarding everyone unmatched. That discarding costs data, and the final answer describes only the treated kind of person.
Weighting keeps everyone. Instead of hunting twins, it reshapes each group until both look like the full population. The treated group is reweighted to resemble everyone, and so is the untreated group. Then the comparison is fair, and it answers the population question: what would happen if everyone got, or nobody got, the treatment.
How it works
person chance of landed how much
being treated in they count
rich, urban high (usual) treated about once (expected)
poor, rural low (rare) treated ten times (surprising)
rich, urban high (usual) control ten times (surprising the other way)
poor, rural low (rare) control about once (expected)
each group, reweighted, now mirrors the whole populationThe rule is short: the rarer it was for someone to land where they landed, the more times their answer counts. Expected cases count about once. Surprising cases count a lot — they carry the missing voices. The developer tab turns the rule into the exact arithmetic.
The soft spot is visible in the table: someone who almost never lands where they landed gets counted a runaway number of times. One extremely surprising person can dominate the whole answer, so real analyses cap or trim extreme weights and always look at their spread.
A real example you have seen
Election pollsters live on this method. Their phone samples over-reach some groups and under-reach others. So every published poll number is a weighted average, with each respondent counted by how under-represented their kind is. When a poll spectacularly misses, an unmodelled group with tiny reach — and thus wild weights — is a usual suspect.
Remember this
- The rarer your landing, the more times your outcome counts.
- Surprising cases speak for the many who went the other way.
- Near-zero probabilities make explosive weights — trim, cap, and inspect them.
What to learn next
- Difference-in-differences — leaving the ignorability world: identification from time.
- Double machine learning — IPW's doubly robust, ML-powered descendant.
- Propensity score matching — the pairing alternative, and when each wins.
Developer — Code and libraries.
Setup
pip install numpy scikit-learnVerified with numpy 1.26.4 and scikit-learn 1.7.2.
The same biased offer data, fixed by weights
Identical setup to the matching lesson: the offer truly adds 40 rupees, and rich urban customers get it far more often.
import numpy as np
from sklearn.linear_model import LogisticRegression
rng = np.random.default_rng(12)
n = 6000
income = rng.normal(0, 1, n)
city = (rng.random(n) < 0.5).astype(float)
p_offer = 1 / (1 + np.exp(-(1.2 * income + 0.8 * city - 0.5)))
offer = rng.random(n) < p_offer
spend = 500 + 300 * income + 100 * city + 40 * offer + rng.normal(0, 50, n)
X = np.column_stack([income, city])
ps = LogisticRegression().fit(X, offer).predict_proba(X)[:, 1]
# weight each person by 1 / probability of the arm they landed in
w = np.where(offer, 1 / ps, 1 / (1 - ps))
treated = np.average(spend[offer], weights=w[offer])
control = np.average(spend[~offer], weights=w[~offer])
print(f"naive: {spend[offer].mean() - spend[~offer].mean():+.1f} IPW: {treated - control:+.1f} (truth: +40)")
print(f"largest weight: {w.max():.1f} (a weight of 50 means one person stands in for 50)")
trim = np.clip(ps, 0.02, 0.98) # cap extreme propensities before weighting
wt = np.where(offer, 1 / trim, 1 / (1 - trim))
t2 = np.average(spend[offer], weights=wt[offer])
c2 = np.average(spend[~offer], weights=wt[~offer])
print(f"IPW with trimmed weights: {t2 - c2:+.1f}")naive: +334.3 IPW: +39.5 (truth: +40) largest weight: 70.3 (a weight of 50 means one person stands in for 50) IPW with trimmed weights: +43.2
The walkthrough
From +334 to +39.5 with one weighted average. The propensity model is the same one matching used; only its deployment changed. np.average(..., weights=...) computes the reweighted group means, and their difference lands next to the truth.
One customer carries weight 70. Somewhere in the data sits a person whose profile almost guaranteed the other arm. Their good or bad month moves the estimate noticeably — the estimator is unbiased but can be jumpy. Checking w.max() and the weight distribution is not optional hygiene; it is the diagnostic.
Trimming trades a little bias for a lot of calm. Clipping propensities into [0.02, 0.98] caps any weight at 50. The estimate moves to +43.2 — slightly biased now, far more stable across reruns. Production analyses routinely report both.
Weighting versus matching, in one line each: matching answers "what did treatment do for the treated kind of person" and discards the unmatched; weighting answers "what would treatment do for everyone" and keeps all rows at the price of variance. Same score, different questions.
Common mistakes
Using 1/ps for everyone. The control group's weight is 1/(1-ps) — the probability of the arm they landed in. This sign slip produces silently absurd estimates.
Ignoring the positivity requirement. If some profile is never treated, no weight can conjure their treated outcome — the data contains no such voice to amplify. Weights explode exactly where positivity frays; huge w.max() is that alarm ringing. The same requirement appeared as empty strata in confounding.
Reporting naive standard errors. Weighted estimates have variance driven by the weights; treating rows as equally-weighted i.i.d. understates uncertainty badly. Use robust (sandwich) errors or the bootstrap.
Building the propensity model to maximise accuracy. As with matching: a sharper classifier pushes probabilities toward 0 and 1, which worsens weights. Calibrated, modest models serve weighting better than leaderboard champions.
Try it yourself
Break positivity on purpose: regenerate offers with 1.2 replaced by 4.0, making treatment nearly deterministic in income. Watch w.max() and the untrimmed estimate across a few seeds. Then compute the stabilised weights — multiply each weight by its arm's share (offer.mean() or 1 - offer.mean()) — and confirm the estimate matches while the weights shrink.
What to learn next
- Difference-in-differences — leaving the ignorability world: identification from time.
- Double machine learning — IPW's doubly robust, ML-powered descendant.
- Propensity score matching — the pairing alternative, and when each wins.
Researcher — Mathematics and papers.
The estimator family
With propensity $e(X)$ and the ignorability-plus-positivity assumptions of the potential outcomes lesson, the Horvitz–Thompson (1952, JASA — from survey sampling) form identifies mean potential outcomes:
$$ \mathbb{E}[Y(1)] = \mathbb{E}!\left[\frac{T\,Y}{e(X)}\right], \qquad \mathbb{E}[Y(0)] = \mathbb{E}!\left[\frac{(1-T)\,Y}{1 - e(X)}\right] $$
- $T$ — treatment indicator; $Y$ — observed outcome; $e(X)$ — propensity score.
The proof is one line of iterated expectations: $\mathbb{E}[TY/e(X)] = \mathbb{E}[\mathbb{E}[T \mid X]\, Y(1)/e(X)] = \mathbb{E}[Y(1)]$. The Hájek variant (used in the code) normalises by the sum of weights instead of $n$ — biased in finite samples, usually lower variance, invariant to outcome location shifts; prefer it in practice.
Stabilised weights $sw = \Pr(T{=}t)/\Pr(T{=}t \mid X)$ (Robins, Hernán and Brumback, 2000, Epidemiology) keep expectation 1 per arm, taming variance without changing the estimand — essential in the longitudinal case, where weights multiply across time and marginal structural models are built on them.
Variance and the positivity frontier
The semiparametric efficiency bound for the ATE (Hahn, 1998, Econometrica) contains $\mathbb{E}[\sigma^2_1(X)/e(X)]$-type terms: variance diverges as $e(X) \to 0$ or $1$. Under weak overlap:
- Trimming samples with $e(X)$ outside $[\alpha, 1-\alpha]$ (Crump, Hotz, Imbens and Mitnik, 2009, Biometrika — the 0.1 rule) changes the estimand to a trimmed population, honestly.
- Overlap weights $e(X)(1-e(X))$ (Li, Morgan and Zaslavsky, 2018, JASA) target the population of clinical equipoise and are bounded by construction.
- Interestingly, weighting by the estimated rather than true propensity lowers asymptotic variance (Hirano, Imbens and Ridder, 2003, Econometrica) — the estimation step performs implicit covariate balancing.
Doubly robust combination
The augmented IPW (AIPW) estimator
$$ \hat{\mu}_1 = \frac{1}{n}\sum_i \left[ \frac{T_i (Y_i - \hat{m}_1(X_i))}{\hat{e}(X_i)} + \hat{m}_1(X_i) \right] $$
- $\hat{m}_1(x)$ — an outcome-regression estimate of $\mathbb{E}[Y \mid T{=}1, X{=}x]$.
is consistent if either $\hat{e}$ or $\hat{m}_1$ is correct (Robins, Rotnitzky and Zhao, 1994; Scharfstein, Rotnitzky and Robins, 1999), and efficient when both are. This double robustness plus Neyman orthogonality is the gateway to ML nuisance estimation — the exact construction in double machine learning. TMLE (van der Laan and Rubin, 2006) is the targeted-likelihood sibling with better finite-sample behaviour under weight instability.
Reading
- Hernán and Robins (2020), Causal Inference: What If, chapters 12–13 — IPW and marginal structural models.
- Austin and Stuart (2015), Moving towards best practice when using IPW, Statistics in Medicine — the applied checklist (balance, weight diagnostics, robust variance).
- Cole and Hernán (2008), Constructing inverse probability weights for marginal structural models, AJE.
What to learn next
- Difference-in-differences — leaving the ignorability world: identification from time.
- Double machine learning — IPW's doubly robust, ML-powered descendant.
- Propensity score matching — the pairing alternative, and when each wins.