Pipelines and why leakage disappears
A Pipeline chains your preprocessing and your model into one sealed estimator, which makes the most common scoring lie in machine learning impossible to commit.
- 8 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
A pipeline glues your data-cleaning steps and your model into one machine, so test data can never touch the learning steps.
Think of a water purifier bolted to a kitchen wall. Water enters at the top, passes through every filter in order, and comes out of one tap. You cannot accidentally drink from a middle stage, because the middle stages are sealed inside the body.
A scikit-learn Pipeline seals your steps the same way. Scaling, feature selection, and the model become one unit with one fit and one predict.
Why it exists
The most common way to fool yourself in machine learning is called leakage — letting information from your test data sneak into training. It is like a student who saw the answer key while preparing. Their mock score looks brilliant, and it means nothing.
Leakage rarely looks like cheating. It looks like tidy housekeeping: "I scaled the whole table first, then split it." That innocent order of operations lets the test rows influence the scaling. Your score quietly inflates, and no error message ever appears.
A pipeline removes the opportunity. The cleaning steps live inside the sealed machine, so they can only ever learn from whatever data fit receives.
How it works
one sealed estimator
┌──────────────────────────────┐
raw data → │ scale → select → model │ → prediction
└──────────────────────────────┘
fit(train) every inner step fits on train ONLY
predict(test) every inner step only applies what it learnedWhen a checking tool like cross-validation takes over, it splits your data into several folds. For each fold it gets a fresh, unfitted copy of the whole pipeline. The sealed body guarantees the held-out fold stayed unseen.
A real example you have seen
Mock exams work only when the practice papers do not contain the real questions. A coaching centre that drills students on leaked papers produces stellar mock scores and shocked faces on results day. Models leak the same way: brilliant in the lab, useless on real data.
Remember this
- Leakage means test information sneaking into training. It inflates scores silently.
- A Pipeline chains preprocessing and model into one estimator with one
fit. - Anything that learns from data belongs inside the pipeline, never before it.
What to learn next
- ColumnTransformer for mixed data — different preparation per column, still one sealed machine.
- Choosing a cross-validation strategy — the splitting side of an honest score.
- Overfitting and underfitting — the other way scores lie.
Developer — Code and libraries.
Setup
pip install scikit-learnTested against scikit-learn 1.7.
Watching leakage inflate a score
The data below is pure noise. Sixty patients, a thousand random gene readings, coin-flip labels. There is nothing to learn, so an honest score is around 0.5. Watch what the wrong order of operations reports.
import numpy as np
from sklearn.feature_selection import SelectKBest, f_classif
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import cross_val_score
from sklearn.pipeline import make_pipeline
rng = np.random.default_rng(0)
X = rng.normal(size=(60, 1000)) # 60 patients, 1000 random gene readings
y = rng.integers(0, 2, size=60) # coin-flip labels: nothing here is learnable
# The wrong way: pick the 20 "best" genes while looking at ALL the data
X_leaky = SelectKBest(f_classif, k=20).fit_transform(X, y)
leaky = cross_val_score(LogisticRegression(), X_leaky, y, cv=5).mean()
# The right way: selection lives inside the pipeline, refit fresh on each fold
pipe = make_pipeline(SelectKBest(f_classif, k=20), LogisticRegression())
honest = cross_val_score(pipe, X, y, cv=5).mean()
print(f"leaky accuracy: {leaky:.2f}")
print(f"honest accuracy: {honest:.2f}")leaky accuracy: 0.83 honest accuracy: 0.45
Read that again. 83% accuracy on coin flips. Nothing crashed, nothing warned. The only witness is the pipeline version telling the truth.
The walkthrough
Where the 83% came from. SelectKBest hunts for the 20 columns most correlated with the labels. With 1000 random columns, some correlate with the labels by luck. Because the selector looked at all 60 rows — including every future test fold — it hand-picked exactly the lucky columns. Cross-validation then "confirmed" a pattern that was baked in before splitting.
Why the pipeline version is honest. cross_val_score clones the pipeline for each fold and fits the clone on that fold's training part alone. The selector picks lucky columns from 48 training rows, and those columns are not lucky on the 12 unseen ones. Score collapses to a truthful coin-flip.
make_pipeline names steps for you. It lowercases the class names: selectkbest, logisticregression. The Pipeline([("select", ...), ("clf", ...)]) form lets you choose names. You need names for inspection and grid search:
pipe.fit(X, y)
print(pipe.named_steps["selectkbest"].get_support().sum())20
One object end to end. pipe.predict(X_new) pushes rows through selection, then the model. You cannot forget a step at prediction time, because the steps are not separate objects anymore.
Common mistakes
Scaling before splitting. StandardScaler().fit_transform(X) followed by train_test_split lets test rows shift the training mean. The inflation is small for scaling, so people call it harmless — until a distribution shift makes it matter. The pipeline habit costs nothing; keep the scaler inside.
Selecting features, then splitting. The mistake demonstrated above, and the one with the biggest inflation. Feature selection, target encoding, and oversampling are the three highest-risk steps. All three learn from labels, so all three must sit inside.
Imputing missing values from the full table. Filling gaps with a column's overall mean uses test rows to compute that mean. Put the imputer — the tool that fills missing values — inside the pipeline like anything else. See handling missing data for imputation itself.
Reporting the training-set score. pipe.score(X_train, y_train) measures memory, not skill. A pipeline stops leakage; it does not stop you grading the student on questions they studied. Judge on held-out data always.
Try it yourself
Change k from 20 to 100 and rerun. Predict first: will the leaky score rise or fall? Then add StandardScaler() as the first pipeline step and confirm the honest score barely moves — on data this random, honesty is stable.
What to learn next
- ColumnTransformer for mixed data — different preparation per column, still one sealed machine.
- Choosing a cross-validation strategy — the splitting side of an honest score.
- Overfitting and underfitting — the other way scores lie.
Researcher — Mathematics and papers.
Leakage, stated precisely
Let D = {(x_i, y_i)} be drawn i.i.d. from distribution P, split into folds. An evaluation is unbiased when the fitted function tested on fold F_k was produced by a procedure that received only D \ F_k. Leakage is any violation: the procedure includes every data-dependent choice — feature selection, scaling statistics, hyperparameter choice — not only the final weight fit.
The demo above is the selection-bias case analysed in Hastie, Tibshirani and Friedman, The Elements of Statistical Learning (2nd ed., §7.10.2, "The wrong and right way to do cross-validation"). With d independent null features and n samples, screening for the top k by association with y selects features whose sample correlation is extreme by chance, on the order of sqrt(log(d)/n). The downstream classifier then separates the folds it was screened on, and CV measures the screening artefact rather than generalisation.
Why Pipeline composes correctly
Pipeline.fit(X, y) runs fit_transform through steps 1..m−1 and fit on step m. The composed object satisfies the estimator contract, so any meta-estimator treats it atomically: clone rebuilds every step unfitted, and each CV fold receives an independent clone. The guarantee is structural — the held-out fold is not available to inner steps — rather than a runtime check.
Corollaries:
- Grid search over preprocessing becomes valid:
GridSearchCV(pipe, {"selectkbest__k": [10, 50]})re-selects per fold per candidate. Step-name double-underscore addressing reaches any depth. - Nested CV remains necessary for reporting: a score selected by CV is itself optimistically biased (Varma and Simon, 2006, Bias in error estimation when using cross-validation for model selection). The pipeline removes preprocessing leakage, not selection-of-the-best bias.
memory="cachedir"memoises fitted transformers keyed on parameters and input hash — useful when an expensive early step is shared across a grid.
Cost
Pipeline overhead is one Python call per step per fit or predict; asymptotically free. The real cost is refitting preprocessing per fold: k-fold CV multiplies preprocessing cost by k. That cost is the price of the honest estimate, and memory recovers much of it when grids repeat steps.
Scale of the problem in the literature
Kapoor and Narayanan (2023), Leakage and the reproducibility crisis in ML-based science, survey 294 papers across 17 fields and find leakage-driven overclaiming in every field surveyed, with preprocessing-before-split among the most common variants. Their taxonomy (no test set, preprocessing leakage, temporal leakage, non-independence) is a useful review checklist: the first two are exactly what Pipeline eliminates by construction.
What to learn next
- ColumnTransformer for mixed data — different preparation per column, still one sealed machine.
- Choosing a cross-validation strategy — the splitting side of an honest score.
- Overfitting and underfitting — the other way scores lie.