Causal Inference Basics

Uplift modelling

Uplift modelling predicts who a treatment will actually change — separating the persuadables from people who would buy anyway and people the nudge will annoy.

On this page 5
  1. Why it exists
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Uplift modelling predicts, for each individual person, how much a treatment would change their behaviour — so you spend the treatment only where it changes something.

A sweet-shop owner considers handing out discount coupons. Four kinds of customer walk in. The regular aunty buys her weekly kaju katli, coupon or not — a coupon for her is charity from the till. The window-shopper leaves either way — the coupon is wasted paper. The hesitant student buys only if nudged — the coupon works precisely on him. And one proud customer finds the discount insulting and walks out — the coupon backfires.

Marketing names for the four: sure things, lost causes, persuadables, and do-not-disturbs. Uplift modelling is the craft of finding the persuadables in advance.

Why it exists

Ordinary prediction models target the wrong customers by design. Ask a model "who will buy after getting the coupon?" and it ranks the regular aunty at the top. She buys after a coupon, after a monsoon, after anything. Campaigns built on purchase-likelihood shower money on the sure things.

The question that pays is not "who buys?" but "whose behaviour changes?" — the difference between buying-if-nudged and buying-if-left-alone. That difference is a per-person causal effect from potential outcomes, and no single customer ever reveals both halves of it. Uplift models estimate it anyway, from a randomised campaign.

How it works

randomised trial data:
   half got the coupon, half did not

train:  model A — buying, among coupon-getters
        model B — buying, among the rest

per customer:
   uplift = A's prediction - B's prediction

  aunty:    0.95 - 0.94 = +0.01   sure thing — skip
  student:  0.60 - 0.20 = +0.40   persuadable — target!
  proud:    0.30 - 0.45 = -0.15   do-not-disturb — avoid

Each customer gets two predicted futures — nudged and unnudged — and the gap between the futures is the score. Spend flows to the biggest gaps, not the biggest buyers.

A real example you have seen

Telecom churn teams learned this painfully. Calling customers who might leave with a retention offer sometimes reminds them to leave — do-not-disturbs are real, and dialling them costs customers. Modern retention, coupon and political-canvassing operations all run uplift targeting. The coupon that mysteriously arrives exactly when you were hesitating about a purchase is not coincidence.

Remember this

  • Uplift = predicted change caused by the treatment, per person.
  • Likely buyers and persuadable buyers are different people — targeting the first wastes money.
  • Training needs randomised data: a campaign where a coin chose who got the nudge.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install numpy scikit-learn

Verified with numpy 1.26.4 and scikit-learn 1.7.2.

The T-learner: two models, one subtraction

Simulated randomised campaign: loyalty drives buying regardless; the coupon works only on price-sensitive newcomers. Because we built the world, each customer's true effect is known — so the targeting strategies can be scored honestly.

uplift.py
import numpy as np
from sklearn.ensemble import GradientBoostingClassifier

rng = np.random.default_rng(22)
n = 20_000
loyal = rng.random(n)                       # 0 = brand new, 1 = die-hard regular
sensitive = rng.random(n)                   # how much price moves this person
coupon = rng.random(n) < 0.5                # a properly randomised campaign

# loyal customers buy anyway; the coupon only moves price-sensitive newcomers
lift = 0.25 * sensitive * (1 - loyal)
buy = rng.random(n) < (0.15 + 0.55 * loyal + coupon * lift)

X = np.column_stack([loyal, sensitive])
m1 = GradientBoostingClassifier(random_state=0).fit(X[coupon], buy[coupon])
m0 = GradientBoostingClassifier(random_state=0).fit(X[~coupon], buy[~coupon])
uplift = m1.predict_proba(X)[:, 1] - m0.predict_proba(X)[:, 1]
likely = m0.predict_proba(X)[:, 1]          # "who buys anyway" — the naive target

top_uplift = uplift >= np.quantile(uplift, 0.8)
top_likely = likely >= np.quantile(likely, 0.8)
print(f"true effect on the top-20% by uplift score:  {lift[top_uplift].mean():.3f}")
print(f"true effect on the top-20% likely buyers:    {lift[top_likely].mean():.3f}")
print(f"true effect on a random 20%:                 {lift.mean():.3f}")
Output
true effect on the top-20% by uplift score:  0.129
true effect on the top-20% likely buyers:    0.018
true effect on a random 20%:                 0.062

The walkthrough

Uplift targeting doubles random; likely-buyer targeting falls below random. Per coupon: 0.129 extra purchases for uplift targeting, 0.062 for random spray — and 0.018 for the "target our best customers" strategy. Same budget, same models, different question asked.

Why do likely buyers underperform a coin toss? The likeliest buyers are the loyal regulars — exactly the people whose behaviour the coupon cannot change (the (1 - loyal) factor in lift zeroes them out). Random spray at least hits some persuadables by accident; likely-buyer targeting systematically selects the unmovable. This is the sweet-shop aunty receiving charity from the till, at scale.

This two-model recipe is the T-learner ("T" for two). Simple, uses any classifier (gradient boosting here), and its weakness is visible in the code: each model optimises its own arm's accuracy, and their errors do not cancel when subtracted. Small individual errors can dominate the small differences being estimated. The S-, X- and R-learner variants (researcher block) attack exactly this.

Real campaigns cannot print lift. True effects stay hidden; evaluation uses held-out randomised data ranked by score — checking that top-decile customers show the largest treated-versus-control gap. The standard picture is the Qini curve, cumulative incremental buys against people targeted.

Common mistakes

Training on observational campaign data. If last year's coupons went to customers marketing chose, the two models learn the choosing policy's fingerprints — confounding inside a subtraction. The randomised requirement is not negotiable without the heavy machinery of earlier lessons.

Evaluating with accuracy or AUC. Uplift has no observable label — no customer reveals both futures. Model-quality claims need uplift-specific tools (Qini/AUUC on randomised holdout), not classification metrics.

Reading per-person scores as precise. Individual uplift estimates are noisy differences of noisy predictions. Ranking and decile statements are defensible; "Ramesh's uplift is exactly 0.40" is not.

Ignoring negative uplift. The do-not-disturb segment justifies the whole method in retention settings — skipping the bottom of the ranking is often worth more than refining the top.

Try it yourself

Add a do-not-disturb segment: subtract 0.3 * (loyal > 0.9) * coupon-style annoyance from the buying probability and retrain. Check the bottom decile's true effect turns negative — and compute how much the campaign gains by excluding them versus targeting the top only.

What to learn next

Researcher — Mathematics and papers.

Estimand and meta-learners

The target is the conditional average treatment effect (CATE):

$$ \tau(x) = \mathbb{E}[Y(1) - Y(0) \mid X = x] $$

  • $\tau(x)$ — the per-covariate-profile effect; identified under the randomisation (or ignorability) of $T$.

Meta-learner taxonomy (Künzel, Sekhon, Bickel and Yu, 2019, PNAS):

  • S-learner: one model $\hat{\mu}(x, t)$; $\hat{\tau}(x) = \hat{\mu}(x,1) - \hat{\mu}(x,0)$. Regularisation shrinks the treatment term toward zero — biased when effects are small relative to signal.
  • T-learner: the code above; separate $\hat{\mu}_1, \hat{\mu}_0$. No sharing across arms; error non-cancellation.
  • X-learner: imputed individual effects ($D_i^1 = Y_i - \hat{\mu}_0(X_i)$ for treated, mirrored for control) fed to second-stage regressions, blended by propensity — dominant under unequal arm sizes.
  • R-learner (Nie and Wager, 2021, Biometrika): minimises the residualised objective from Robinson's decomposition — a CATE-targeted cousin of DML's partialling out.
  • Causal forests (Wager and Athey, 2018, JASA; Athey, Tibshirani and Wager, 2019) — honest trees splitting on effect heterogeneity, with asymptotic normality per leaf-neighbourhood; econml and grf implement all of the above.

Evaluation without individual labels

  • Qini curve / AUUC: rank by $\hat{\tau}$, plot cumulative incremental outcome (treated minus reweighted control) against fraction targeted; Qini coefficient = area versus random (Radcliffe, 2007). The uplift analogue of ranking metrics.
  • Class-variable transformation (Jaskowski and Jaroszewicz, 2012): with 50/50 randomisation, $Z = Y \cdot T + (1-Y)(1-T)$ satisfies $\mathbb{E}[2Z - 1 \mid X] = \tau(X)$ — turning CATE into a single-model target.
  • Pseudo-outcome validation: doubly robust scores $\hat{\Gamma}_i$ (AIPW residuals) provide unbiased per-unit effect proxies for model selection (Schuler et al., 2018; Curth and van der Schaar, 2023, on the pitfalls).

Decision layer

Targeting is a policy problem: treat when $\hat{\tau}(x) \cdot v > c$ (value $v$, cost $c$), connecting to policy learning with regret guarantees (Athey and Wager, 2021, Econometrica). Budget constraints turn it into a knapsack on $\hat{\tau}$-ranked units. Persistent subtlety: optimising the policy tolerates far more CATE estimation error than the estimation literature suggests — ranking quality, not calibration, binds.

Reading

  • Radcliffe and Surry (2011), Real-world uplift modelling with significance-based uplift trees — the practitioner lineage.
  • Künzel et al. (2019) — meta-learners; Nie and Wager (2021) — the R-learner.
  • Gutierrez and Gérardy (2017), Causal inference and uplift modelling: a review (PMLR) — the bridge survey.

What to learn next