Why a great model answers the wrong question
A model can predict brilliantly and still be useless for decisions, because prediction describes the world as it is — while decisions change the world.
- 7 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
A predictive model tells you what usually happens; a decision needs to know what would happen if you stepped in and changed something. These are different questions.
Count the umbrellas on your street and you can predict rain impressively well. Many umbrellas — rain is coming or already here. Now the wrong conclusion: ban umbrellas to stop the rain. The prediction was real. The lever is fake. Umbrellas and rain travel together because a third thing — the weather — drives both.
Every model that learns patterns from the world inherits this trap. Patterns tell you what goes together. They stay silent about what causes what.
Why it matters
Machine learning is superb at "what goes together". A model sees that customers who use your app's search feature rarely cancel. True pattern, strong signal.
The product team now spends a quarter pushing everyone to use search — and cancellations do not move. Search users stayed because they were engaged; search was the umbrella, engagement was the weather. The model answered "who will cancel?" perfectly. The team asked "what should we change?" — a question the model never heard.
How it works
what the model learned:
weather ──→ rain
│
└─────→ umbrellas umbrellas and rain: correlated!
what the decision assumed:
umbrellas ──→ rain ← this arrow does not existThe tell is the direction of use. Reading the world through a model is safe: umbrellas out, carry a raincoat. Changing the world through a model needs arrows — knowledge of which lever moves which outcome. Getting arrows requires experiments, or the careful methods in the rest of this section.
A real example you have seen
Hospitals found that asthma patients with pneumonia had better survival than average. The pattern was strong enough that a naive model would send asthma patients home as low-risk. The truth: doctors rushed asthma patients into intensive care, and the extra care caused the survival. Acting on the prediction would have removed exactly the thing producing it.
Remember this
- Prediction answers "what will I see?"; decisions need "what happens if I act?".
- Patterns travel together because of hidden common causes — the weather behind the umbrellas.
- Accuracy is no defence: the asthma model was accurate and would have killed people.
What to learn next
- Potential outcomes and counterfactuals — the formal language for "what would have happened".
- Confounding — the anatomy of the umbrella trap.
- Randomisation units and assignment — the cleanest escape: run the intervention yourself.
Developer — Code and libraries.
Setup
pip install numpy scikit-learnVerified with numpy 1.26.4 and scikit-learn 1.7.2.
A 97%-accurate model that fails the intervention test
Because this is a simulation, we can do what reality forbids: run the world twice — once as observed, once with an intervention forced on it.
import numpy as np
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import accuracy_score
rng = np.random.default_rng(0)
n = 10_000
def world(umbrella_ban=False):
storm = rng.random(n) < 0.3 # the real cause of both things
umbrellas = np.where(storm, rng.poisson(40, n), rng.poisson(5, n))
if umbrella_ban:
umbrellas = np.zeros(n, dtype=int) # the intervention: force it to 0
wet = storm & (rng.random(n) < 0.9) # rain soaks people, not umbrellas
return umbrellas.reshape(-1, 1), wet
X, y = world()
model = LogisticRegression().fit(X, y)
print(f"prediction accuracy from umbrella counts: {accuracy_score(y, model.predict(X)):.3f}")
Xb, yb = world(umbrella_ban=True)
print(f"wet people in the normal world: {y.sum():,}")
print(f"wet people after banning umbrellas: {yb.sum():,}")prediction accuracy from umbrella counts: 0.967 wet people in the normal world: 2,721 wet people after banning umbrellas: 2,722
The walkthrough
96.7% accuracy — a genuinely good classifier. Umbrella counts separate wet days from dry days almost perfectly, because both are downstream of storms. Nothing about the model is broken.
The ban changes nothing: 2,721 versus 2,722 wet people. The intervention severs umbrellas from their cause, and the outcome does not care. The model's accuracy measured the strength of the storm → umbrella arrow, which the ban cannot use.
The umbrella_ban flag is the star of this code. It implements what causal inference calls an intervention — reaching into the system and setting a variable, rather than observing it. Interventions in simulations are free; in reality they cost an A/B test or the observational methods later in this section.
Train/test discipline would not have caught this. Held-out accuracy from the same world is still 97% — see train test split for what that machinery does and does not check. The failure only appears under a changed world, which no split of observational data contains.
Common mistakes
Reading feature importance as leverage. "Feature X tops the importance chart, so changing X changes the outcome" — importance measures predictive contribution, which common causes inflate freely. The umbrella feature has maximal importance and zero leverage.
Fixing it with more data. Ten times the rows sharpens the estimate of the same non-causal pattern. Bias does not average out; it compounds confidence.
Believing complex models escape the trap. Gradient boosting and deep networks fit the correlation structure better, which makes the wrong answer more precise. The trap is in the question, not the model class.
Concluding prediction is useless. Prediction is the right tool when you will not intervene on the inputs: forecasting demand, ranking risk, routing tickets. The failure begins where the model's inputs become your levers.
Try it yourself
Add a cloud_seeding variable that actually causes rain, and give the model both features. Verify the model weights umbrellas heavily anyway — then check which intervention (banning umbrellas versus stopping seeding) changes wet counts. Feature importance and leverage part ways before your eyes.
What to learn next
- Potential outcomes and counterfactuals — the formal language for "what would have happened".
- Confounding — the anatomy of the umbrella trap.
- Randomisation units and assignment — the cleanest escape: run the intervention yourself.
Researcher — Mathematics and papers.
Two different conditional objects
Prediction estimates the observational conditional $P(Y \mid X = x)$: the distribution of $Y$ among units where $X$ happens to equal $x$. Decisions need the interventional conditional $P(Y \mid do(X = x))$: the distribution of $Y$ when $X$ is set to $x$ regardless of its usual causes.
- $do(X = x)$ — Pearl's do-operator: an intervention replacing $X$'s structural mechanism with the constant $x$.
- The two coincide iff $X$ has no unblocked backdoor path to $Y$ — no confounding, in the sense made precise in confounding and causal DAGs.
In the umbrella system: $P(\text{wet} \mid \text{umbrellas} = 0) \approx 0$ (no umbrellas implies no storm), while $P(\text{wet} \mid do(\text{umbrellas} = 0)) = P(\text{storm}) \cdot 0.9 \approx 0.27$. The simulation prints exactly this divergence.
Pearl's ladder
Pearl (2009, Causality, 2nd ed.; Pearl and Mackenzie 2018, The Book of Why) organises queries into three rungs: association ($P(y \mid x)$), intervention ($P(y \mid do(x))$), and counterfactuals ($P(y_x \mid x', y')$). Supervised learning, however large, operates on rung one; no amount of rung-one data alone identifies rung-two quantities without structural assumptions — the formal content of the "correlation is not causation" slogan.
When predictive deployment is itself an intervention
Deploying a model changes behaviour, shifting the distribution it learned from:
- Performative prediction (Perdomo, Zrnic, Mendler-Dünner and Hardt, 2020, ICML) — predictions that influence the outcomes they predict, with fixed-point notions of stability.
- The asthma-pneumonia case — Caruana et al. (2015), Intelligible Models for HealthCare (KDD): the pattern was real, treatment-mediated, and acting on it inverts it. Their remedy was interpretable models so clinicians could spot treatment-mediated features.
- Strategic gaming / Goodhart's law — once a proxy becomes a target, its correlation with the goal degrades.
The estimand-first discipline
The methodological repair is asking, before modelling: what quantity would answer the decision? If it is $\mathbb{E}[Y \mid do(T=1)] - \mathbb{E}[Y \mid do(T=0)]$, then identification (experiment, adjustment, instrument, discontinuity — each of the coming lessons) precedes estimation. Hernán, Hsu and Healy (2019), A second chance to get causal inference right (CHANCE), and Hernán and Robins, Causal Inference: What If (2020) — the free standard textbook — both open with exactly this framing.
What to learn next
- Potential outcomes and counterfactuals — the formal language for "what would have happened".
- Confounding — the anatomy of the umbrella trap.
- Randomisation units and assignment — the cleanest escape: run the intervention yourself.