Causal Inference Basics

Double machine learning

Double ML lets flexible models soak up messy confounding — predict the treatment from covariates, predict the outcome from covariates, and read the causal effect from the two leftovers.

Read these first

On this page 5
  1. Why it exists
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Double machine learning measures a cause's effect by first removing everything else's influence from both the cause and the outcome — then comparing the leftovers.

Two friends work in the same office, and you wonder whether one's mood genuinely lifts the other's. The office itself moves both of them — deadlines darken everyone, Fridays brighten everyone. To see the private connection, first account for the office: work out what the office alone would predict for each friend's mood today. Then look at the leftovers — how far each friend sits above or below their office-predicted mood. If one friend's leftover tracks the other's, that link is theirs alone; the office's fingerprints were wiped first.

Double ML does this wiping with machine learning, twice — once for the treatment, once for the outcome. Hence "double".

Why it exists

Adjusting for confounders with straight lines fails when reality curves. Age's effect on spending is not a line; neither is income's on risk. Powerful learners — random forests, boosting — capture curves happily. But plugging them naively into causal work backfires. They absorb some of the effect being measured, and their slow, regularised learning contaminates the estimate.

The 2018 breakthrough was a recipe that makes flexible models safe for one precious number. Use them only to predict away nuisance — the confounders' influence. Keep the final step deliberately dumb. And never let a model grade its predictions on the data it trained on.

How it works

step 1:  ML predicts treatment from covariates   → treatment leftovers
step 2:  ML predicts outcome  from covariates    → outcome leftovers
step 3:  a plain line through (treatment leftovers, outcome leftovers)
         → its slope is the causal effect

rule throughout: every prediction is made on data
                 the model never trained on  ("cross-fitting")

The leftovers are the parts of treatment and outcome that the confounders cannot explain. Whatever connection survives in both leftovers is the treatment's own doing.

A real example you have seen

Pricing teams at large tech companies ask "what does a 10-rupee price change do to demand?" They ask it of history where prices were set by algorithms reacting to demand, seasonality and stock. That is confounding of the deepest kind. Double ML is the standard modern answer, powering open-source libraries built at Microsoft (EconML) and used across pricing, advertising and policy evaluation.

Remember this

  • Two ML models wipe the confounders' fingerprints off treatment and outcome.
  • The effect is read from a plain comparison of the leftovers.
  • Never predict on training data — cross-fitting is what keeps the ML honest.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install numpy scikit-learn

Verified with numpy 1.26.4 and scikit-learn 1.7.2.

Curvy confounding: lines fail, DML survives

The confounder acts through a sine and a square — shapes a linear adjustment cannot absorb. True effect of the treatment: +1.5.

dml.py
import numpy as np
from sklearn.ensemble import RandomForestRegressor
from sklearn.linear_model import LinearRegression
from sklearn.model_selection import cross_val_predict

rng = np.random.default_rng(30)
n = 4000
X = rng.uniform(-2, 2, (n, 3))
confound = np.sin(2 * X[:, 0]) + X[:, 1] ** 2          # curvy: no line captures this
t = confound + rng.normal(0, 1, n)                     # treatment driven by X
y = 1.5 * t + 3 * confound + rng.normal(0, 1, n)       # true effect of t: +1.5

# attempt 1: linear regression with X tacked on as extra columns
linear = LinearRegression().fit(np.column_stack([t, X]), y).coef_[0]

# attempt 2: DML — predict away X's influence on BOTH t and y, relate the leftovers
def forest():
    return RandomForestRegressor(n_estimators=200, min_samples_leaf=20, random_state=0)

t_res = t - cross_val_predict(forest(), X, t, cv=5)    # cross-fitting: held-out folds
y_res = y - cross_val_predict(forest(), X, y, cv=5)
theta = (t_res * y_res).sum() / (t_res ** 2).sum()

print(f"linear adjustment: {linear:+.3f}")
print(f"double ML:         {theta:+.3f}  (truth: +1.500)")
Output
linear adjustment: +3.426
double ML:         +1.593  (truth: +1.500)

The walkthrough

Linear adjustment reports +3.43 — over double the truth. It "controlled for X", and the columns of X are even the right variables. But the confounding flows through sin and a square, and a line through raw columns leaves most of it standing. Controlling for the right variables in the wrong shape still fails.

cross_val_predict is the honesty mechanism. Each row's prediction comes from a model trained on the other four folds. Without this, the forest partially memorises its training rows, the residuals shrink dishonestly, and the final slope biases. Cross-fitting is not a nicety — the method's guarantees name it.

The final line is deliberately dumb. (t_res * y_res).sum() / (t_res ** 2).sum() is one-variable least squares — the slope of outcome-leftovers on treatment-leftovers. All flexibility lives in the nuisance predictions; the causal parameter itself is estimated by arithmetic a referee can check.

Why two models beat one: predicting only the outcome leaves treatment-selection bias; predicting only the treatment wastes outcome signal. The product structure of the leftover regression makes small errors in each model cancel to second order — one model can be somewhat wrong if the other is decent. That forgiveness is the "double" earning its keep.

Common mistakes

Skipping cross-fitting because "the forest is regularised". Regularisation does not remove own-fit contamination. The bias is largest exactly when the learner is most flexible — the case DML exists for.

Reading the forests' feature importances causally. The nuisance models are prediction tools; their internals inherit every trap from why prediction is not causation. Only theta carries causal meaning here.

Using DML to dodge the thinking. DML relaxes the functional form assumption, not the causal ones: all confounders must be in X (no unmeasured confounding), and X must not contain mediators or colliders. Garbage causal structure in, confident garbage out.

No uncertainty on theta. The residual regression yields standard errors almost for free (heteroskedasticity-robust ones on the final OLS are asymptotically valid). A point estimate alone invites overconfidence.

Try it yourself

Delete the cross_val_predict calls and fit-predict on the full data instead — watch theta drift and the residual variance shrink suspiciously. Then swap the forests for LinearRegression nuisances and confirm you recover the broken +3.4: DML with linear nuisances is linear adjustment.

What to learn next

Researcher — Mathematics and papers.

Neyman orthogonality and the partialling-out estimator

In the partially linear model $Y = \theta T + g(X) + \varepsilon$, $\;T = m(X) + v$, with $\mathbb{E}[\varepsilon \mid X, T] = 0$, $\mathbb{E}[v \mid X] = 0$:

  • $\theta$ — the target causal parameter; $g, m$ — unknown nuisance functions.

Robinson (1988) partialling-out estimates $\theta$ by regressing $Y - \mathbb{E}[Y \mid X]$ on $T - m(X)$ — the code's residual slope. Chernozhukov, Chetverikov, Demirer, Duflo, Hansen, Newey and Robins (2018), Double/debiased machine learning for treatment and structural parameters (Econometrics Journal), show the corresponding moment function

$$ \psi(W; \theta, \eta) = \big(Y - \ell(X) - \theta\,(T - m(X))\big)\,(T - m(X)) $$

is Neyman-orthogonal: its Gateaux derivative with respect to the nuisances $\eta = (\ell, m)$ vanishes at the truth. First-order nuisance errors therefore do not propagate; the estimator error involves products of nuisance errors, so $\sqrt{n}$-normality of $\hat{\theta}$ holds whenever both nuisances converge faster than $n^{-1/4}$ — attainable by forests, boosting, lasso and neural nets under structure. Naive plug-in (non-orthogonal) moments transmit the full first-order regularisation bias — the formal reason "predict Y, read off T's coefficient" fails.

Cross-fitting

Sample-splitting removes the own-observation bias term $\mathbb{E}[\hat{g}(X_i)\,\varepsilon_i] \ne 0$ that arises when $\hat{g}$ is trained on observation $i$; K-fold cross-fitting (estimate nuisances on $K{-}1$ folds, evaluate moments on the held-out fold, rotate, average) restores full-sample efficiency. Both ingredients — orthogonality and cross-fitting — are needed for the main theorem; each alone fails.

Beyond the partially linear model

  • AIPW/interactive model: fully nonparametric ATE via the doubly robust score from IPW, with ML nuisances and cross-fitting — DML and TMLE converge on the same efficient influence function.
  • CATE: the R-learner objective (Nie and Wager, 2021) is orthogonalised CATE regression — the bridge to uplift modelling.
  • IV-DML: orthogonal scores for instrumented parameters; also panel, mediation and dynamic extensions.
  • Software: DoubleML (Bach et al., 2022, JMLR) and Microsoft's EconML implement the full zoo with valid inference.

Cautions from the frontier

Nuisance quality still binds: slower-than-$n^{-1/4}$ learning (very high dimensions, weak signals) breaks the guarantee, and overlap failures inflate the residual denominator's variance exactly as in IPW. Chernozhukov et al. recommend reporting nuisance cross-validated fits alongside $\hat{\theta}$; sensitivity analysis for unmeasured confounding (Chernozhukov, Cinelli, Newey, Sharma and Syrgkanis, 2022 — "Long Story Short") extends the omitted-variable-bias toolkit to the DML setting.

What to learn next