Data Leakage and Results That Are Too Good

Choosing features on the full dataset

Selecting the "best" features using all your data lets pure noise masquerade as signal — the selection step must live inside the validation loop.

On this page 5
  1. Why it exists
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

If you pick your best features by looking at all the data, the picking itself is cheating — even if everything afterwards is done honestly.

Flip 5,000 different coins ten times each. A few coins will land heads eight or nine times, by luck alone. Now keep those "lucky" coins, throw the rest away, and announce you have found special coins. Test them again and they behave like any coin — the luck does not repeat.

Choosing features works the same way. Among thousands of useless columns, some will look related to your labels by pure chance. Select those, and you have selected luck.

Why it exists

Modern datasets are wide. Gene studies measure 20,000 genes for 100 patients. Sensor rigs log thousands of channels. Marketing tables sprout hundreds of engineered columns. Nobody wants to feed all of that into a model. So a feature selection step picks the most promising columns: the ones that move together with the outcome.

The trap: with many columns and few rows, chance correlations are guaranteed. If selection sees all your data, it harvests those chance correlations. Then the later train/test split cannot save you — both sides now contain only columns pre-picked for flattering this exact dataset.

How it works

5,000 noise columns, 100 rows
        │
        ▼
"keep the 20 most correlated"  ← luck harvested HERE, using all rows
        │
        ▼
split → train → test           ← honest machinery, poisoned input
                                  score: excellent. truth: nothing.

The fix mirrors the last two lessons: the selection step must be re-done inside each training fold, blind to the fold's test rows.

A real example you have seen

Health headlines: "People who eat X live longer!" Studies that track hundreds of foods will find some food matching long life by chance. Follow-up studies then fail to confirm it. The famous illustration: a brain scan study found "activity" in a dead salmon, because checking thousands of brain locations guarantees a few false hits. Selecting features on the full dataset performs that salmon experiment on your own data.

Remember this

  • Among thousands of noise features, some always match your labels by luck.
  • Selecting on the full dataset harvests that luck into both train and test.
  • Selection is a learned step — it belongs inside the pipeline, refit per fold.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install scikit-learn

Verified with scikit-learn 1.7.2, numpy 1.26.4, CPU. Seeded, so your numbers should match.

85% cross-validated accuracy on pure noise

Every feature is random noise. Every label is a coin flip. Honest accuracy is 50%. The only difference between the two runs is where the selection happens.

lucky_coins.py
import numpy as np
from sklearn.feature_selection import SelectKBest, f_classif
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import cross_val_score
from sklearn.pipeline import make_pipeline

rng = np.random.default_rng(7)
X = rng.normal(size=(100, 5000))    # 5,000 features of pure noise
y = rng.integers(0, 2, 100)         # 100 coin-flip labels

# wrong: pick the 20 "best" features using every row, then cross-validate
X_picked = SelectKBest(f_classif, k=20).fit_transform(X, y)
leaky = cross_val_score(LogisticRegression(), X_picked, y, cv=5)

# right: the selector lives inside the pipeline, so it only sees training folds
pipe = make_pipeline(SelectKBest(f_classif, k=20), LogisticRegression())
honest = cross_val_score(pipe, X, y, cv=5)

print("select before CV:", round(leaky.mean(), 3))
print("select inside CV:", round(honest.mean(), 3))
Output
select before CV: 0.85
select inside CV: 0.54

Thirty-one points of pure fiction, from moving one line across a boundary.

The walkthrough

SelectKBest(f_classif, k=20) scores each feature by how strongly it separates the classes, then keeps the top 20. Applied to all 100 rows of 5,000 noise columns, the top 20 are the columns whose noise happened to align with these particular labels — the lucky coins.

Why the split cannot rescue it: the surviving 20 columns were chosen for correlating with all 100 labels, test labels included. Train and test rows both live inside that hand-picked coordinate system. The contamination happened at column level, before any row was split.

In the honest pipeline, each fold's selector picks its own top 20 using only that fold's 80 training rows. Its lucky columns are lucky for the training rows — and luck does not transfer to the 20 test rows. Score: chance, plus wobble.

The scale of the lie grows with width. More candidate columns means luckier winners. This is exactly the wide-data regime — genomics, text n-grams, sensor arrays — where selection is most tempting.

Common mistakes

Exploratory selection that silently becomes final. You explore correlations in a notebook on full data, "later" redo it properly, and never do. The exploration already spent your test set. Budget for it — or explore on a carved-out slice you then discard.

Selecting with the model's own importances, on full data. Fitting a forest on everything, keeping its top features, then cross-validating a new forest — same bug wearing model-flavoured clothes.

Dropping features by hand after eyeballing full-data plots. Human eyes harvesting luck are still harvesting luck. The rule covers you, too.

Believing regularisation makes selection unnecessary-but-harmless. L1 models select internally inside the fit — which is fine, because it happens per fold. External full-data selection stacked on top reintroduces the leak.

Try it yourself

Reduce the width from 5,000 to 50 columns and rerun. The leaky score drops sharply. Then push it to 20,000 (a few seconds of compute) and watch the fiction grow. You are plotting the price of luck against the number of lottery tickets.

What to learn next

Researcher — Mathematics and papers.

The expected size of the lie

For $p$ independent noise features and $n$ samples, the maximum absolute sample correlation with any label vector concentrates around

$$ \max_j |\hat\rho_j| \approx \sqrt{\frac{2 \ln p}{n}} $$

Where:

  • $p$ — number of candidate features.
  • $n$ — number of samples.
  • $\hat\rho_j$ — sample correlation of feature $j$ with the labels.

With $p = 5000$, $n = 100$: $\sqrt{2 \ln 5000 / 100} \approx 0.41$. The selected subset consists of ~20 such features, each spuriously "moderately predictive"; a linear model over them separates the training sample comfortably. The $\sqrt{\ln p}$ growth explains why the effect is gentle at 50 features and devastating at 20,000.

The canonical references

  • Ambroise and McLachlan (2002), Selection bias in gene extraction on the basis of microarray gene-expression data, PNAS: re-examined published cancer-classification results; gene selection performed outside CV produced near-zero error estimates on data where honest error was large. The paper that made this bug famous in bioinformatics.
  • Hastie, Tibshirani and Friedman (2009), The Elements of Statistical Learning, §7.10.2, "The Wrong and Right Way to Do Cross-validation": the textbook demonstration this lesson's demo reproduces — their simulated wrong-way error rate was 3% on a problem whose true error was 50%.
  • Cawley and Talbot (2010), JMLR: situates selection bias within the general theory of over-fitting in model selection.

Correct procedures, ranked by strictness

  1. Selection inside the resampling loop (the pipeline above): unbiased for the procedure "select then fit". Note the estimand: you are evaluating the recipe, and each fold may pick different features. Feature-stability across folds is itself informative — Meinshausen and Bühlmann (2010), Stability Selection, JRSS-B, make it the primary object.
  2. Nested CV when hyperparameters (including $k$) are tuned: outer loop untouched by any choice — Varma and Simon (2006).
  3. Sample splitting with a lockbox: selection and tuning on one portion; a final, single evaluation on an untouched portion. The lockbox connects to the adaptive-reuse theory in the next lesson.

Post-selection inference

A parallel statistics literature asks the sharper question: what are valid p-values and intervals after selection? Selective inference (Taylor and Tibshirani, 2015, PNAS; lee et al.'s lasso-conditional tests) conditions on the selection event to restore validity. For ML evaluation purposes, the pipeline discipline suffices; for scientific claims about which features matter, that literature is the standard.

The multiple-comparisons frame — thousands of hypotheses, guaranteed false discoveries, Benjamini–Hochberg control — is covered from the statistics side in hypothesis testing; the dead-salmon study (Bennett et al., 2009, IgNobel-winning fMRI poster) is its most quotable exhibit.

What to learn next