Data Leakage and Results That Are Too Good

Your model scores 99%: what to suspect first

An amazing score is usually a broken experiment, not a breakthrough — here is the checklist to run before you believe any number that looks too good.

On this page 6
  1. Why this lesson exists
  2. How it works
  3. The suspect list
  4. A real example you have seen
  5. Remember this
  6. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

A score that looks amazing is a reason to investigate, not a reason to celebrate.

Imagine a student who scores full marks on a mock exam. Later you find the answer key was printed on the back of the question paper. The student is not a genius. The exam was broken.

Machine learning has the same failure, and it happens constantly. When a model scores 99%, the most likely explanation is that the answer reached the model through a side door. This side door has a name: data leakage — information about the answer sneaking into the data the model learns from.

Why this lesson exists

New practitioners celebrate high scores. Experienced practitioners fear them.

That flip in instinct is one of the biggest differences between the two groups. Real problems are messy. Predicting who repays a loan, or which customer leaves, is genuinely hard. Honest models on hard problems score modestly.

So when a hard problem suddenly looks easy, something in the experiment is usually feeding the model the answer. If you present that 99% to your team, someone else will find the leak. It is much better to find it yourself.

How it works

what you think happened:
  features  →  model  →  brilliant predictions

what actually happened:
  answer ──┐
           ├─ hidden inside a feature  →  model  →  copied answers
  features ┘

The model is not cheating on purpose. It is doing its job: finding the strongest pattern. If the strongest pattern is a smuggled copy of the answer, the model will find that.

The suspect list

When a score looks too good, check these, in this order:

  1. One feature contains the answer. Something recorded after the outcome happened.
  2. The same rows appear in training and test. Duplicates straddle the split.
  3. Preparation touched the test data. Scaling or selecting features before splitting.
  4. Time ran backwards. The model saw the future while predicting the past.
  5. The same person is on both sides. The model recognises individuals, not patterns.

Each suspect gets its own lesson in this section.

A real example you have seen

In 2020, many models claimed to detect COVID-19 from chest X-rays with stunning accuracy. Reviews later found that many had learned shortcuts instead. Some learned which hospital a scan came from, because sick and healthy scans came from different machines. The scores were real. The skill was not.

Remember this

  • A surprising score means a broken experiment until proven otherwise.
  • Data leakage is answer information sneaking into your inputs.
  • Work through the suspect list yourself, before someone else does it for you.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install scikit-learn pandas

Verified with scikit-learn 1.7.2, pandas 2.2.3, numpy 1.26.4, on CPU. Seeded, so your numbers should match.

A 100% score, manufactured on purpose

We predict loan defaults. One feature, collections_calls, counts calls made to recover money. Those calls happen after a customer defaults. At prediction time, that column would not exist yet.

suspicious_score.py
import numpy as np
import pandas as pd
from sklearn.ensemble import RandomForestClassifier
from sklearn.model_selection import train_test_split

rng = np.random.default_rng(0)
n = 400
income = rng.normal(50, 15, n).round(1)
loan_size = rng.normal(5, 2, n).round(1)
defaulted = (rng.random(n) < 0.3).astype(int)

# written into the database AFTER the customer defaulted
collections_calls = defaulted * rng.integers(3, 9, n) + rng.integers(0, 2, n)

X = pd.DataFrame({"income": income, "loan_size": loan_size,
                  "collections_calls": collections_calls})
Xtr, Xte, ytr, yte = train_test_split(X, defaulted, test_size=0.5,
                                      random_state=0, stratify=defaulted)
model = RandomForestClassifier(random_state=0).fit(Xtr, ytr)
print("test accuracy:", round(model.score(Xte, yte), 3))
for name, imp in zip(X.columns, model.feature_importances_):
    print(f"  {name:18s} importance {imp:.3f}")
Output
test accuracy: 1.0
  income             importance 0.088
  loan_size          importance 0.090
  collections_calls  importance 0.821

A perfect score, on a held-out test set, with a proper split. Every standard check passes. The experiment is still worthless.

The walkthrough

The test set did its job — and it was not enough. The leak lives inside a feature, so it travels into the test set along with everything else. Holding data out cannot protect you from a poisoned column.

Feature importance is your first flashlight. One feature carries 82% of the importance. A single dominant feature on a supposedly hard problem is the classic leak signature. It is not proof — sometimes one feature is legitimately strong — but it tells you exactly where to look.

The question that catches it is not statistical. It is: "Would this column exist at the moment of prediction?" Collections calls happen after default. Case closed.

Now remove the leak

Append this to the same file, so Xtr, Xte, ytr, yte and defaulted are already in scope.

suspicious_score.py (continued)
Xtr2 = Xtr.drop(columns="collections_calls")
Xte2 = Xte.drop(columns="collections_calls")
model = RandomForestClassifier(random_state=0).fit(Xtr2, ytr)
print("accuracy without the leak:", round(model.score(Xte2, yte), 3))
print("share of customers who did not default:", round(1 - defaulted.mean(), 3))
Output
accuracy without the leak: 0.65
share of customers who did not default: 0.705

Two things worth sitting with. First, the honest score collapsed from 1.0 to 0.65. Second, 0.65 is below the 0.705 you would get by predicting "no default" for everyone. The remaining features carry almost no signal — which is true, because this synthetic data gave them none. Compare every score against that lazy baseline before reporting it.

Common mistakes

Trusting a score because the split was correct. A clean split does not clean the columns. Leakage rides inside features, duplicates and preprocessing, all of which cross the split untouched.

Skipping the baseline. Always compute the majority-class score first. A model at 94% on a dataset that is 93% one class has learned almost nothing.

Explaining a great score with a flattering story. "The model found deep patterns" is a hypothesis. So is "a column contains the answer". Test the boring one first — it wins most of the time.

Deleting the suspicious feature and moving on silently. Find out why that column existed and who else is using it. Leaky columns tend to appear in many projects at the same company.

Try it yourself

Add a feature account_flagged = defaulted with 10% of values flipped. Retrain and check the accuracy and importances. Then work out the highest accuracy this corrupted copy of the answer could possibly give.

What to learn next

Researcher — Mathematics and papers.

A formal definition of the smell

Let $X$ be the feature vector, $Y$ the target, and $t$ the moment a prediction must be made. A feature $X_j$ is legitimate only if it is measurable with respect to the information available at time $t$. Leakage is the use of any $X_j$ that is a function of information unavailable at $t$ — most often a descendant of $Y$ itself in the causal graph.

Where:

  • $X \in \mathbb{R}^d$ — the features presented to the model.
  • $Y$ — the outcome being predicted.
  • $t$ — prediction time; everything usable must exist strictly before it.

The reference formulation is Kaufman, Rosset and Perlich (2012), Leakage in Data Mining: Formulation, Detection, and Avoidance, ACM TKDD — which grew out of leakage repeatedly deciding KDD Cup competitions.

Why held-out testing cannot detect it

The train/test split estimates generalisation error under the assumption that both sets are drawn i.i.d. from the deployment distribution. Leakage violates the premise, not the estimator: the joint distribution $P(X, Y)$ at training time differs from the deployment distribution, where the leaky coordinate of $X$ is unavailable or uninformative. The held-out estimate is then an unbiased estimate of performance on a distribution you will never see again.

Scale of the problem

Kapoor and Narayanan (2023), Leakage and the Reproducibility Crisis in ML-based Science, Patterns, surveyed applied-ML literature and documented leakage-driven overoptimism in hundreds of papers spanning 17 scientific fields, from medicine to political science. Their proposed remedy — model info sheets declaring how each feature would be available at prediction time — is essentially Kaufman's legitimacy rule made into a checklist.

For the COVID-19 imaging example: Roberts et al. (2021), Common pitfalls and recommendations for using machine learning to detect and prognosticate for COVID-19 using chest radiographs and CT scans, Nature Machine Intelligence, reviewed 62 studies and found none of clinical use, with leakage and dataset shortcuts among the dominant failures. DeGrave, Janizek and Lee (2021), Nature Machine Intelligence, showed such models relied on source-specific artefacts — laterality markers and hospital-specific processing.

Detection heuristics that generalise

  • Dominance: a single feature with outsized importance or a near-perfect univariate AUC. See target leakage for the scan.
  • Performance cliff on temporal deployment: models that degrade sharply when evaluated strictly forward in time. See temporal leakage.
  • Ablation response: removing one feature collapses the score. See hunting a leak with ablations.
  • Too-flat learning curves: near-ceiling accuracy from tiny training fractions suggests the task is lookup, not learning.

The suspects are examined one by one across this section; the closing lesson assembles them into a hunting procedure.

What to learn next