Data Leakage and Results That Are Too Good
Overfitting your own test set
Every time you peek at your test score and adjust, a little of the test set's luck leaks into your choices — enough peeks, and the test set stops being a test.
- 7 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
A test set answers honestly only the first time — every reuse leaks a little of its answer key into your decisions.
Play the guessing game "hot or cold". Your friend hides a coin; each guess earns a hint — warmer, colder. After twenty hints, you find the coin. Nobody told you where it was. The hints, added up, were the location.
Your test score is a hint. Try a change, check the score, keep the change if the score rose. Each check is one "warmer or colder". Do it fifty times and your model has slowly been guided to this particular test set — without ever training on it.
Why it exists
Nobody sets out to cheat. The workflow feels responsible: make a change, measure it, keep improvements. That loop is the daily rhythm of model development.
The poison is where the measurement comes from. Improvements measured on the test set select for two things at once: genuine skill, and luck specific to those test rows. Genuine skill transfers to new data. The luck does not — and after enough iterations, luck is most of what you have collected.
How it works
try change → score on TEST → keep if better ─┐
▲ │
└──────────── repeat 50 times ────────────┘
each loop: a little real improvement
+ a little luck that fits THESE test rows only
new data arrives → the luck evaporates → score dropsThe defence is a hierarchy: a validation set — a slice reserved for these repeated checks — absorbs the peeking. The test set stays sealed for one final measurement.
A real example you have seen
Public AI leaderboards. Teams submit hundreds of entries, tuning against the leaderboard's hidden test data one score at a time. Winners sometimes fall down the rankings when a fresh, final test set is revealed — their climb was partly a fifty-hint search of the public one. Competition sites now split leaderboards into "public" and "private" for exactly this reason.
Remember this
- Each test-set evaluation you act on leaks information about it into your model.
- Tune against a validation slice; open the test set once, at the end.
- The more attempts, the luckier the best one looks — track your attempt count.
What to learn next
- Did the model already see your test set? — the LLM era's version, where the training corpus is the whole internet.
- Report a range, not a single number — the honest way to present results after all this discipline.
- Overfitting and underfitting — the classical overfitting this lesson generalises to your workflow.
Developer — Code and libraries.
Setup
pip install numpyVerified with numpy 1.26.4, CPU. Seeded, so your numbers should match.
Guessing our way to 71% on coin flips
No training, no features, no model. One thousand random prediction sets, and we keep the one the test set likes best. Watch expertise appear from nothing — and vanish on fresh data.
import numpy as np
rng = np.random.default_rng(8)
y_test = rng.integers(0, 2, 100) # your 100-example test set
best_score, best_model = 0.0, None
for trial in range(1000): # 1,000 "improvements", all random guessing
guesses = rng.integers(0, 2, 100)
score = (guesses == y_test).mean()
if score > best_score:
best_score, best_model = score, guesses
print("best accuracy found on the test set:", best_score)
y_fresh = rng.integers(0, 2, 100) # data the winner never influenced
print("the same winner on fresh data :", (best_model == y_fresh).mean())best accuracy found on the test set: 0.71 the same winner on fresh data : 0.48
Twenty-one points of accuracy, manufactured from coin flips by selection alone. On untouched data: chance.
The walkthrough
Nothing here trained on the test set. Every guess vector is independent random noise. The corruption enters through one operator: if score > best_score — keeping what the test set rewarded. Selection is the leak.
Your real workflow is this loop in slow motion. Each architecture tweak, threshold nudge, or prompt rewrite evaluated on the same data is one trial. Real changes carry real signal too — so the decay on fresh data is partial, not total. But the gap between "score that guided development" and "score on untouched data" grows with every acted-upon peek.
The arithmetic of the 0.71 is worth internalising: the best of $k$ random tries on $n$ examples lands near $0.5 + \sqrt{\ln k / 2n}$. With $k=1000$, $n=100$: about 0.69. Small test sets amplify the effect brutally — with $n=30$, the same thousand peeks manufacture roughly 84%.
The working discipline
- Three-way split: train / validation / test. Iterate freely against validation. The habit starts in train-test split.
- The test set is opened once, for the final report. If you acted on it, it has become a validation set — carve out a new test set if you can.
- Refresh validation when it wears out. Cross-validation re-cut with a new seed, or a re-sampled validation slice, resets accumulated luck.
- Log every evaluation. An experiment tracker gives you the attempt count $k$ — the number your final claim must be discounted by.
Common mistakes
"I never trained on it, so it's clean." Selection contaminates without training. The demo above trains nothing.
Ten reruns with different seeds, reporting the best. Seed shopping is this lesson with $k=10$. Report the spread — report a range, not a single number.
Early stopping on the test set. Choosing when to stop by test score is an acted-upon peek per epoch. Stop on validation loss.
A leaderboard-sized team sharing one test set. Five people, two hundred peeks each: the team's $k$ is a thousand. The test set degrades per organisation, not per person.
Try it yourself
Shrink the test set to n = 30 and rerun. Then try trials = 10 versus trials = 100000. Plot nothing — the two printed numbers against $0.5 + \sqrt{\ln k / 2n}$ tell the story. Watch the formula track your experiment.
What to learn next
- Did the model already see your test set? — the LLM era's version, where the training corpus is the whole internet.
- Report a range, not a single number — the honest way to present results after all this discipline.
- Overfitting and underfitting — the classical overfitting this lesson generalises to your workflow.
Researcher — Mathematics and papers.
The bound behind the demo
For $k$ independent classifiers of true accuracy $\tfrac12$ evaluated on $n$ i.i.d. binary outcomes, Hoeffding plus a union bound gives
$$ \mathbb{E}\left[\max_{j \le k} \hat{A}_j\right] \;\lesssim\; \frac{1}{2} + \sqrt{\frac{\ln k}{2n}} $$
Where:
- $\hat{A}_j$ — measured accuracy of attempt $j$ on the test set.
- $k$ — number of attempts (queries) evaluated.
- $n$ — test-set size.
$k = 1000$, $n = 100$ gives $\approx 0.686$; the run observed 0.71, within the approximation's slack. The $\sqrt{\ln k}$ dependence is the mercy — attempts hurt logarithmically — and $\sqrt{1/n}$ is the warning: small test sets are consumed fast.
Adaptive data analysis
The demo's attempts were independent; real development is adaptive — each query depends on previous answers, which breaks the union-bound analysis and can be exponentially worse. The foundational treatment is Dwork, Feldman, Hardt, Pitassi, Reingold and Roth (2015), The reusable holdout: Preserving validity in adaptive data analysis, Science: answering test-set queries through a differentially-private mechanism (adding calibrated noise, thresholding) supports quadratically more adaptive queries than naive reuse. Blum and Hardt (2015), The Ladder, ICML, give the leaderboard-specific mechanism: only report improvements exceeding a threshold, which bounds leaderboard error against arbitrarily many submissions.
Does the field actually overfit its benchmarks?
Empirically, less than the theory permits — an interesting open question. Recht, Roelofs, Schmidt and Shankar (2019), Do ImageNet Classifiers Generalize to ImageNet?, ICML: fresh test sets built to the original recipes drop every model's absolute accuracy (3–15 points), but rankings are strikingly preserved — the drop pattern suggests distribution gap more than adaptive overfitting. Roelofs et al. (2019), A Meta-Analysis of Overfitting in Machine Learning, NeurIPS, analysed thousands of Kaggle public/private splits: substantial adaptive overfitting is rare on large test sets, and concentrated where $n$ is small — consistent with the $\sqrt{\ln k / n}$ scaling.
The safe conclusion runs through the formula: a 10,000-example test set forgives thousands of peeks; a 200-example medical set is spent after a dozen serious ones.
Hygiene as protocol design
- Lockbox data: a slice untouched until publication; institutionally enforced in competitions via private leaderboards.
- Query accounting: report $k$ alongside the score; discount via the bound or via Bonferroni-style corrections — connects to is this improvement real?
- Rotating benchmarks: fresh test data over time (LiveBench-style, see the next lesson) trades comparability for validity.
What to learn next
- Did the model already see your test set? — the LLM era's version, where the training corpus is the whole internet.
- Report a range, not a single number — the honest way to present results after all this discipline.
- Overfitting and underfitting — the classical overfitting this lesson generalises to your workflow.