Data leakage
In one sentence Data leakage is when information that will not exist at prediction time sneaks into training, producing scores that collapse in production.
Updated
Data leakage is when the training process gets information it will not have in real use — so test scores look brilliant and production results collapse.
It is the exam paper leaking the night before. Every student scores 95%, the school celebrates, and the board exam a month later reveals what was actually learned. The tragedy of leakage is that nothing looks wrong: training runs fine, the test score is superb, and the failure only appears when real money is on the line.
Leakage takes a few classic shapes, worth memorising because each has burned thousands of teams:
target leakage a feature secretly contains the answer
("discharge_medication" predicts the diagnosis — recorded after it)
split leakage preprocessing fitted on all data before the train/test split
(the scaler saw the test set's statistics)
time leakage random splits on time-ordered data
(training on Tuesday's data to predict Monday)
duplicate leakage the same record lands in both train and testThe defences are procedural, not clever. Split first, before any preprocessing, and fit every transform on training data only — scikit-learn pipelines exist to enforce this inside cross-validation. Split time-series by time. For each feature, ask one question: will this value exist, with this exact meaning, at the moment of prediction? And treat any result that seems too good as a leakage alarm first, a breakthrough second.
Where to go next
- Full lesson: Train/test split
- Related terms: train-test-split, cross-validation, feature-engineering, overfitting