Data Leakage and Results That Are Too Good
Leakage that survives cross-validation
Cross-validation validates whatever happened before it — including your mistakes — so anything computed on the full dataset poisons every fold at once.
- 8 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Cross-validation checks your split — it cannot check what you did to the data before splitting.
Think of a referee with a perfect stopwatch, timing a race with total honesty. What the referee cannot see: one runner walked the course the night before and moved the starting blocks. The timing is flawless. The race was rigged before the whistle.
Cross-validation — testing on several different held-out slices and averaging — is the referee. Many people believe it is leak-proof. It is only split-proof. Anything you computed from the full dataset before the folds were cut is baked into every fold, training and test alike.
Why this misunderstanding exists
Cross-validation genuinely fixes real problems: a lucky split, a small test set (the rows held back to score the finished model), a fluke result. Because it feels rigorous, it collects credit it has not earned. People run five folds, see five agreeing scores, and conclude the number is trustworthy.
But five folds agreeing can mean two things. Either the result is solid — or all five folds inherited the same poison. Agreement between folds proves consistency, not honesty.
How it works
the poison happens here:
full data ──→ [compute something using ALL rows] ──→ folds
│
└── every fold now contains
knowledge of every other fold
fold 1: train | test ✓ looks clean
fold 2: train | test ✓ looks clean ← all five agree,
fold 3: train | test ✓ looks clean all five are wrongAny step that learned from data — an average, a ranking, a category summary — must happen inside each fold, using that fold's training part only.
A real example you have seen
Group projects in college: five members "independently" verify a calculation, but all five copied the same formula sheet from one senior. Five confirmations, one shared error. Independent-looking checks that share an upstream source are one check, repeated.
Remember this
- Cross-validation guards the split, not the steps before the split.
- Anything computed on the full dataset contaminates every fold simultaneously.
- Five agreeing folds can be five copies of the same mistake.
What to learn next
- Choosing features on the full dataset — the most famous special case, with its Nobel-grade victims.
- Fitting the scaler before splitting — the same law at lower volume, if you skipped it.
- Overfitting your own test set — what repeated tuning does to any honest number.
Developer — Code and libraries.
Setup
pip install scikit-learn pandasVerified with scikit-learn 1.7.2, pandas 2.2.3, numpy 1.26.4, CPU. Seeded, so your numbers should match.
72% cross-validated accuracy on coin flips
Target encoding replaces a category (like a city) with the average label of rows in that category. It is powerful, popular — and the most reliable way to poison every fold at once when computed on the full dataset.
The labels below are coin flips. Honest accuracy is 50%. Watch cross-validation certify better.
import numpy as np
import pandas as pd
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import cross_val_score
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import TargetEncoder
rng = np.random.default_rng(6)
n = 300
city = rng.integers(0, 100, n) # 100 cities, ~3 customers each
y = rng.integers(0, 2, n) # labels are pure coin flips
# the leak: each city's average label, computed over ALL rows
city_mean = pd.Series(y).groupby(city).mean()
X_leaky = city_mean[city].to_numpy().reshape(-1, 1)
leaky = cross_val_score(LogisticRegression(), X_leaky, y, cv=5)
print("encode first, validate after :", round(leaky.mean(), 3))
# the fix: encoding happens inside each training fold
pipe = make_pipeline(TargetEncoder(random_state=6), LogisticRegression())
honest = cross_val_score(pipe, city.reshape(-1, 1).astype(str), y, cv=5)
print("encode inside the folds :", round(honest.mean(), 3))encode first, validate after : 0.723 encode inside the folds : 0.44
Same encoder, same model, same folds. The order of operations is the entire difference between 72% and honest chance.
The walkthrough
Why is the leak so strong here? Each city has about three rows. A city's average label is mostly made of each row's own label — with three rows, your own coin flip is a third of "your city's average". The feature quietly carries the answer into whichever fold the row lands in. Rare categories make target encoding radioactive.
Why does cross-validation certify it? By the time folds are cut, the poison is already inside the feature column. Every training fold contains encodings built partly from its test fold's labels. The referee timed a rigged race, five times, consistently.
The fix is structural. TargetEncoder inside a Pipeline gets re-fitted per fold, and — a lovely detail — scikit-learn's TargetEncoder (added in 1.3) additionally cross-fits internally during training, so a row is never encoded using its own label. The honest 0.44 is chance, with small-sample wobble.
This is the same law as the scaler lesson, at maximum volume. Scalers learn label-free averages: mild. Target encoders learn label averages: catastrophic.
The checklist for "does CV protect me here?"
Cross-validation does not protect against:
- Preparation fitted on the full dataset (this lesson).
- Duplicates — twins straddle fold boundaries.
- Groups — unless you pass
groups. - Time — unless folds respect order.
- Leaky features — poison rides inside columns.
- Tuning on your CV score hundreds of times — that is test-set overfitting against the folds.
CV protects against exactly one enemy: an unlucky single split.
Common mistakes
"I cross-validated, so the number is safe." The sentence that precedes most public leakage incidents. CV validates the pipeline you gave it, poison included.
Encoding with pandas before modelling with sklearn. The two-library workflow invites full-data groupbys. Keep every learned step inside the pipeline object so the fold machinery governs it.
Tuning hyperparameters on the same folds that produce the final score. The winning configuration was chosen because it flattered these folds. Use nested cross-validation — an outer loop scoring, an inner loop tuning — when the claim matters.
Trusting fold agreement as evidence of correctness. Low variance across folds measures stability. Shared contamination is perfectly stable.
Try it yourself
Change 100 cities to 5 cities and rerun. The leaky score falls close to honest. Work out why: with sixty rows per city, one row's label barely moves its city's average. Leak severity scales with category rarity — a fact worth keeping.
What to learn next
- Choosing features on the full dataset — the most famous special case, with its Nobel-grade victims.
- Fitting the scaler before splitting — the same law at lower volume, if you skipped it.
- Overfitting your own test set — what repeated tuning does to any honest number.
Researcher — Mathematics and papers.
The general theorem of the mistake
Let $T_{\hat\theta}$ be any data-dependent transform with $\hat\theta = g(D_{\text{full}}, Y_{\text{full}})$. Applying $T_{\hat\theta}$ before CV makes every fold's training set a function of every fold's test labels. The CV estimate then measures performance on a distribution where test-label information is embedded in the features — an estimate of nothing deployable.
The label-using case is the dangerous one. For target encoding with $m$ rows per category, a row's own label contributes weight $1/m$ to its encoding; the induced self-information scales as $O(1/m)$, which is why rare categories (small $m$) produce the strongest leak — 0.72 on coin flips at $m \approx 3$.
Where:
- $\hat\theta$ — fitted transform parameters (here, per-category label means).
- $m$ — category cardinality's inverse: rows per category.
Selection bias in model selection
The second CV-surviving leak is choosing using CV, then reporting the same CV. Cawley and Talbot (2010), On Over-fitting in Model Selection and Subsequent Selection Bias in Performance Evaluation, JMLR: the CV score of the selected configuration is an optimistically biased estimate, with bias growing with the number of configurations examined and shrinking with fold count and dataset size. Varma and Simon (2006), BMC Bioinformatics, measured the bias empirically in genomics and prescribed nested CV: an outer loop that never influences any choice.
The bias mechanism is the max-of-noisy-estimates argument developed quantitatively in overfitting your own test set.
Cross-fitting, properly
The correct discipline has a name across fields: cross-fitting. Out-of-fold target encoding (each row encoded by a model of other folds) is standard in gradient-boosting practice; scikit-learn's TargetEncoder implements internal cross-fitting in fit_transform — see its user-guide entry for the exact scheme. The identical construction underlies stacking (Wolpert, 1992, Stacked Generalization, Neural Networks) — out-of-fold predictions as meta-features — and debiased machine learning (Chernozhukov et al., 2018): nuisance functions estimated on complementary folds.
One estimator family, three literatures, one rule: nothing a row's own label touched may describe that row.
What fold agreement does measure
Fold-score variance estimates the variability of the training procedure under data resampling — useful for error bars, treacherous for validity. Contamination is a bias, orthogonal to this variance; no amount of fold agreement bounds it. Diagnostics that do detect it are behavioural: permutation tests (score should collapse to chance under label shuffling — Ojala and Garriga, 2010, JMLR) and the ablation hunt in the closing lesson.
What to learn next
- Choosing features on the full dataset — the most famous special case, with its Nobel-grade victims.
- Fitting the scaler before splitting — the same law at lower volume, if you skipped it.
- Overfitting your own test set — what repeated tuning does to any honest number.