Data Leakage and Results That Are Too Good

Hunting a leak with ablations

When a score is suspiciously good, remove one thing at a time until the score collapses — the piece whose removal breaks the magic is your leak.

On this page 5
  1. Why it exists
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

To find what is secretly powering a suspicious score, switch things off one at a time and watch which switch kills it.

When an appliance keeps tripping your electricity, the electrician does not stare at wires and theorise. She unplugs everything, then reconnects one appliance at a time until the switch trips. The guilty appliance identifies itself by the crash it causes.

Leak-hunting works identically. Something in your pipeline carries the answer. Remove pieces one at a time; the score stays magically high — until you remove the guilty piece, and it crashes. Removing-one-thing-and-remeasuring has a name: an ablation.

Why it exists

The earlier lessons in this section gave you a suspect list: leaky features, duplicates, bad splits, peeking preparation. Fine — but a real pipeline has forty features, six preparation steps and three data sources. Suspicion needs a search procedure, or it is a mood.

Ablation is that procedure. It converts "something is wrong somewhere" into a short loop that points at which thing, usually within an afternoon.

How it works

score with everything        : 99%   ← suspicious
remove feature A, remeasure  : 99%   ← innocent
remove feature B, remeasure  : 99%   ← innocent
remove feature C, remeasure  : 64%   ← THE LEAK

There is a second, sneakier tool: keep everything, but shuffle the answers so labels no longer match rows. A model trained on scrambled answers should score no better than guessing. If it still scores well, your measurement machinery is broken — the score does not depend on truth at all.

Two switches, two suspects: ablation finds poisoned inputs; the shuffle test finds a poisoned scoreboard.

A real example you have seen

Finding which app drains your phone battery: close apps one by one, watch the battery graph. The one whose closing flattens the drain was the culprit. Same loop, different meter.

Remember this

  • Ablation: remove one piece, remeasure, repeat — collapse points at the leak.
  • Shuffle test: scrambled labels should give a chance-level score; anything more means broken measurement.
  • These tools locate leaks; the suspect list tells you what you found.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install scikit-learn pandas

Verified with scikit-learn 1.7.2, pandas 2.2.3, numpy 1.26.4, CPU. Seeded, so your numbers should match.

The hunt, end to end

A churn model scores a perfect 1.000 cross-validated. One of the three features is a leak — account_score comes from a table written after customers churned. Pretend you do not know that, and hunt.

ablation_hunt.py
import numpy as np
import pandas as pd
from sklearn.ensemble import RandomForestClassifier
from sklearn.model_selection import cross_val_score

rng = np.random.default_rng(9)
n = 400
visits = rng.normal(size=n)
spend = rng.normal(size=n)
churned = ((visits + rng.normal(0, 1.5, n)) < 0).astype(int)

# built from a table that already knows who churned
account_score = churned + rng.normal(0, 0.1, n)

X = pd.DataFrame({"visits": visits, "spend": spend,
                  "account_score": account_score})
model = RandomForestClassifier(random_state=9)

full = cross_val_score(model, X, churned, cv=5).mean()
print(f"all features           : {full:.3f}")
for col in X.columns:
    score = cross_val_score(model, X.drop(columns=col), churned, cv=5).mean()
    print(f"without {col:14s}: {score:.3f}")
Output
all features           : 1.000
without visits        : 1.000
without spend         : 1.000
without account_score : 0.640

Three switches flipped, one crash. visits and spend are innocent — the score survives their removal untouched. Removing account_score costs 36 points. That column is either your greatest discovery or your leak, and the prediction-time question settles which: would it exist before the churn happens? No. Leak found.

The shuffle test

Now verify the scoreboard itself. Shuffle the labels so no feature can honestly predict them, and remeasure. Append this to the same file, so rng, churned, X and model are already in scope:

ablation_hunt.py (continued)
y_shuffled = rng.permutation(churned)
score = cross_val_score(model, X, y_shuffled, cv=5).mean()
print(f"accuracy on shuffled labels: {score:.3f}")
Output
accuracy on shuffled labels: 0.530

Chance level — the evaluation machinery is honest. If this number had come out high, no feature ablation would matter: high-on-shuffled means duplicated rows, a broken split, or test data reaching training some other way. Run the shuffle test first; it is one line and it clears (or convicts) your entire measurement setup.

Reading ablation results with care

  • A crash to a believable score is the classic leak signature — 1.000 falling to 0.640, a plausible churn accuracy, is exactly how a leak hides behind an honest-looking remainder.
  • No single crash, still suspicious? Leaks can be spread across correlated features — remove them in groups by source table: everything from billing, everything from support logs. Source-level ablation catches what column-level misses.
  • Ablate pipeline steps, not only features. Rerun without the scaler, without the dedup, with GroupKFold instead of KFold. A score that drops when the split gets stricter is group or temporal leakage confessing.

Common mistakes

Using feature importance instead of ablation. Importances are computed by the possibly-poisoned model on the possibly-poisoned data — a hint, not a measurement. Ablation remeasures reality with the piece genuinely absent.

Removing the leak and keeping the experiment. After removing account_score, the honest baseline is 0.640 — rebuild expectations from there. Teams that anchored on 1.000 keep hunting for the "lost" performance and re-invent the leak.

Stopping at the first leak. Leaks travel in packs; one leaky source table usually poisoned several columns. Finish the loop over all features and sources.

Running ablations against the test set repeatedly. The hunt is itself many evaluations — run it against validation folds, or you trade a leak for test-set overfitting.

Try it yourself

Add spend_per_visit = spend / (visits + 3) — a harmless engineered feature — and a second leak, support_tickets = churned * rng.integers(1, 4, n). Rerun the hunt. Notice how the two leaks mask each other in single-column ablation (removing one leaves the other holding the score up), then catch them with a grouped ablation removing both together.

What to learn next

Researcher — Mathematics and papers.

Ablation as causal intervention on the pipeline

An ablation estimates the performance effect of an intervention $do(\text{remove } j)$ on the training pipeline — retraining without component $j$, not only zeroing its input at inference. The distinction matters: retraining lets remaining features absorb redundant signal, so ablation measures marginal contribution over the remaining set, i.e. $\Delta_j = S(\mathcal{C}) - S(\mathcal{C} \setminus {j})$.

Where:

  • $\mathcal{C}$ — the full component set (features, steps, sources).
  • $S(\cdot)$ — evaluation score of the retrained pipeline.
  • $\Delta_j$ — the crash size attributable to $j$.

With correlated leaks, $\Delta_j$ underestimates each individually (redundancy masking, as in the exercise); grouped ablations approximate coalition values. The complete solution concept is Shapley attribution over $2^{|\mathcal{C}|}$ subsets — principled and exponential; source-grouped ablation is the practical middle ground. Model-explanation variants (SHAP and LIME) attribute predictions; here we attribute evaluation scores, a different and coarser object.

The permutation test, properly

The shuffle test is a permutation test of the null "score independent of label assignment". Ojala and Garriga (2010), Permutation Tests for Studying Classifier Performance, JMLR, formalise two distinct nulls: (1) data and labels independent — the label-shuffling test here; (2) the classifier ignores feature dependency structure — column-wise permutations within classes. Repeating the shuffle $B$ times yields an exact p-value: $p = (#{S_{\pi} \geq S_{\text{obs}}} + 1)/(B + 1)$.

A high score under null (1) convicts the evaluation protocol — duplicated rows across splits, target information in the index, or split machinery reading labels. This is the only test in the section that isolates measurement corruption from input corruption, which is why it runs first.

The chemometrics tradition knows null (1) as y-scrambling / y-randomisation (Rücker, Rücker and Meringer, 2007, J. Chem. Inf. Model.), a mandatory QSAR validity check for decades — an example of a field institutionalising the leak hunt.

Learning-curve and dilution diagnostics

Two additional instruments complement ablation:

  • Learning curves: leakage via lookup (duplicates, memorised groups) produces near-ceiling accuracy at implausibly small training fractions; genuine signal grows smoothly. Plot $S(n_{\text{train}})$ — cheap and discriminating.
  • Noise dilution: progressively corrupt a suspect feature with noise and track the score. Genuine features degrade gracefully; a leak's score-contribution collapses along a sharp curve tracking its mutual information with $Y$ — a quantitative fingerprint of target leakage.

The hunt as standard practice

Kaufman, Rosset and Perlich (2012), TKDD, close their leakage treatment with exactly this workflow: detect by "too good" alarms, localise by component removal, confirm by provenance. Kapoor and Narayanan (2023), Patterns, operationalise it as model info sheets. The cultural point matters more than any single tool: in mature teams, a jump in offline score triggers an investigation, and the ablation loop is the investigation's first page — the same instinct reading a results table sceptically applies to other people's numbers.

What to learn next