Data Leakage and Results That Are Too Good

Target leakage: features that contain the answer

Target leakage is a feature that only exists because the outcome already happened — the single most common way models learn to copy answers instead of predicting them.

On this page 5
  1. Why it exists
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Target leakage is a feature that was created by the answer, instead of being a clue that predicts it.

Look out of the window and see the ground is wet. You conclude it rained. That works — but only because the rain already happened. The wet ground did not predict the rain. The rain caused the wet ground.

A model trained on "wet ground" will look brilliant at rain detection. Ask it to forecast tomorrow's rain, and it has nothing. That is target leakage: a feature that is a footprint of the outcome, not a sign pointing towards it.

Why it exists

Databases record everything that ever happened to a customer, a patient, a parcel. When you export a table for training, columns written before the outcome and columns written after it sit side by side, looking identical.

A hospital table might hold a patient's age next to the medicine they were prescribed. But prescriptions are written after diagnosis. A model predicting the disease from the medicine is reading the doctor's conclusion, not making its own.

Nothing in the file marks which columns are footprints. You have to ask, for every single feature: was this known before the moment of prediction?

How it works

time →──────────────────────────────────────→

  age recorded      diagnosis      medicine prescribed
       │                │                  │
   safe to use      the answer      FOOTPRINT of the answer
                                    (arrived too late)

The prediction moment sits at the diagnosis. Everything to the left is fair. Everything to the right is leakage, no matter how innocent the column name looks.

A real example you have seen

Insurance apps ask for details before showing a premium. Imagine training a claims model on last year's table, including a column "number of claim phone calls". Customers who claimed made calls; customers who did not made none. The model looks perfect and is completely useless — nobody has made claim calls at the moment the policy is priced.

Remember this

  • Target leakage means a feature caused by the outcome, not one predicting it.
  • The test is about time: was this value known before the prediction moment?
  • Column names lie. Ask how and when each value gets written.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install scikit-learn pandas

Verified with scikit-learn 1.7.2, pandas 2.2.3, numpy 1.26.4, CPU. Seeded, so your numbers should match.

The one-feature scan

The fastest leak detector is embarrassingly basic: score every feature alone against the target. A single feature should not be a brilliant predictor of a genuinely hard outcome. We use AUC — a score where 0.5 means coin-flip guessing and 1.0 means perfect ranking (covered in model evaluation).

single_feature_scan.py
import numpy as np
import pandas as pd
from sklearn.metrics import roc_auc_score

rng = np.random.default_rng(1)
n = 600
age = rng.normal(45, 12, n).round(0)
blood_pressure = rng.normal(120, 15, n).round(0)
has_disease = ((age + rng.normal(0, 20, n)) > 55).astype(int)

# prescribed AFTER the doctor made the diagnosis
on_medicine = (has_disease & (rng.random(n) < 0.95)).astype(int)

df = pd.DataFrame({"age": age, "blood_pressure": blood_pressure,
                   "on_medicine": on_medicine})
for col in df.columns:
    auc = roc_auc_score(has_disease, df[col])
    print(f"{col:15s} alone predicts the label with AUC {auc:.3f}")
Output
age             alone predicts the label with AUC 0.750
blood_pressure  alone predicts the label with AUC 0.491
on_medicine     alone predicts the label with AUC 0.971

The walkthrough

Read the scan like a doctor reads a thermometer. age at 0.75 is a strong, believable risk factor. blood_pressure at 0.49 carries nothing here. on_medicine at 0.97 is the alarm: one binary column nearly reproduces the diagnosis.

A high single-feature AUC is a symptom, not a conviction. Some features are legitimately dominant — a spam filter's "sender is in your contacts" can be. The scan tells you where to interrogate. The interrogation question is always the same: when is this value written, and by whom?

The 0.95 in the code matters. Five percent of sick patients are not yet on medicine. Real leaks are rarely perfect copies of the target — they are noisy footprints. AUC 0.97 rather than 1.0 makes a leak harder to spot by eye, which is exactly why you run the scan.

Where target leaks come from in real tables

  • Post-outcome actions: refund issued, medicine prescribed, account closed, collections called.
  • Aggregates computed over all time: "total purchases" includes purchases after the churn you are predicting.
  • Human labels informed by the answer: a fraud analyst's "risk note" written during the investigation.
  • Proxy targets: predicting hospital readmission with "days until next appointment".

Common mistakes

Trusting column names. account_status sounds safe. If it flips to "closed" because of the churn you are predicting, it is the answer wearing a costume.

Scanning only once. New features appear as pipelines evolve. Put the scan in your training script and print it every run — it costs a few lines and catches regressions. Asserting on your data before you train shows the habit in full.

Dropping the leaky column but keeping its children. A feature engineered from a leaky column — a ratio, a rolling mean — inherits the leak. Trace descendants before declaring victory.

Confusing target leakage with a strong feature. The difference is causal direction, not strength. Ask what caused what, and when.

Try it yourself

Add a feature hospital_visits_this_year where sick patients get 2 to 6 visits and healthy ones 0 to 2. Run the scan. Decide whether it is a leak — the honest answer depends on when in the year the prediction happens, which is the entire point.

What to learn next

Researcher — Mathematics and papers.

Legitimacy, formally

Following Kaufman, Rosset and Perlich (2012), ACM TKDD: model the data as observations of random variables with timestamps. A feature $X_j$ observed for prediction at time $t$ is legitimate iff $X_j$ is measurable with respect to $\mathcal{F}_t$, the sigma-algebra of information available at $t$.

Where:

  • $\mathcal{F}_t$ — the information set at prediction time (everything knowable at $t$).
  • $X_j$ — the candidate feature.
  • $Y$ — the target, realised at some time $\geq t$.

Target leakage is the special case where $X_j = f(Y, \varepsilon)$ for some noise $\varepsilon$ — the feature is a descendant of the target in the causal graph. Conditioning on a descendant of $Y$ can make $Y$ arbitrarily easy to infer while carrying zero forward-looking information.

An information-theoretic reading

The single-feature scan approximates a mutual-information screen. For a binary target, univariate AUC is a monotone proxy for how much of $H(Y)$ one feature resolves. A feature with $I(X_j; Y)$ close to $H(Y)$ on a task believed hard is either a scientific discovery or a leak, and the prior strongly favours the leak.

  • $I(X_j; Y)$ — mutual information between feature and target.
  • $H(Y)$ — entropy of the target: the total uncertainty there is to resolve.

The two-schema discipline

The robust prevention is architectural, not statistical: maintain a prediction-time schema — the exact set of fields the serving system will possess at the moment of scoring — and train only on a historical reconstruction of that schema. This is why feature stores implement point-in-time joins: each training row is assembled from feature values as they stood at that row's timestamp, never from the current table. Training-serving skew and its infrastructure are covered from the production side in monitoring and drift.

Kapoor and Narayanan (2023), Patterns, propose "model info sheets" forcing authors to justify each feature's availability at prediction time — a paper-review version of the same schema discipline.

Notable published cases

  • Roberts et al. (2021) and DeGrave et al. (2021), Nature Machine Intelligence: COVID-19 imaging models exploiting acquisition artefacts correlated with outcome.
  • Kaufman et al. (2012) document the KDD Cup 2008 breast-cancer task, where patient IDs encoded the source institution and with it the label prevalence.
  • Filho et al. (2021), and the broader clinical-prediction literature, repeatedly flag "treatment given" features inflating diagnostic models.

The general lesson: leakage is a property of the data-generating process, so its detection ultimately requires domain reasoning. Statistical scans locate suspects; only provenance answers the question.

What to learn next