Imbalanced, Multi-class and Multi-label

Resampling inside cross-validation

Oversample before splitting and copies of your test rows sneak into training — here is the leak, measured, and the pipeline that prevents it.

On this page 5
  1. Why it exists
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Resampling must happen inside each cross-validation fold, on training data only — otherwise your scores are inflated by a hidden leak.

Think of a student who photocopies half the answer key into their practice notes. On the mock exam, some questions look strangely familiar. The mock score comes out brilliant, and the real exam does not.

Why it exists

Cross-validation means testing a model several times, each time hiding a different slice of data and training on the rest. The hidden slice acts as a mini exam. We covered the basic version in train test split.

Now add SMOTE. SMOTE creates new rare examples by blending real neighbours. Run it on the whole dataset first, and some blends are mixtures of rows that later land in the hidden slice. The exam questions were photocopied into the study notes.

This mistake is a form of data leakage — information from the test data quietly reaching the training data. It is one of the most common serious mistakes in published machine-learning work. Honest sentence: nearly everyone makes it once. The goal is making it only once.

How it works

WRONG:  all data → SMOTE → split into folds → score   (blends cross the wall)

RIGHT:  all data → split into folds
                   fold's training part → SMOTE → train
                   fold's hidden part   → untouched → score

The rule generalises far beyond SMOTE. Scaling, feature selection, filling missing values — anything learned from data must be learned inside the fold, from the training part only.

A real example you have seen

Medical-AI papers get retracted or corrected for exactly this. A model announces 99% accuracy detecting a disease from scans; reviewers find the augmentation or resampling ran before the split. Real-world performance falls apart in the clinic. Leakage reviews now form a standard part of serious ML auditing.

Remember this

  • Resample inside each fold, on the training part only.
  • The test slice must stay untouched and imbalanced — that is the world the model will meet.
  • The rule covers every learned preprocessing step, not only resampling.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install scikit-learn imbalanced-learn

Outputs verified with scikit-learn 1.7.2 and imbalanced-learn 0.14.2, CPU only.

Measure the leak

cv_leak.py
from imblearn.over_sampling import SMOTE
from imblearn.pipeline import Pipeline as ImbPipeline
from sklearn.datasets import make_classification
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import cross_val_score

X, y = make_classification(n_samples=1000, n_features=6, weights=[0.95, 0.05],
                           class_sep=0.5, random_state=42)

# WRONG: oversample first, cross-validate afterwards
X_bad, y_bad = SMOTE(random_state=42).fit_resample(X, y)
bad = cross_val_score(LogisticRegression(max_iter=1000), X_bad, y_bad,
                      cv=5, scoring="recall")

# RIGHT: SMOTE lives inside the pipeline, so it runs per fold, on train only
pipe = ImbPipeline([("smote", SMOTE(random_state=42)),
                    ("model", LogisticRegression(max_iter=1000))])
good = cross_val_score(pipe, X, y, cv=5, scoring="recall")

print(f"resample-then-CV (leaky):  recall = {bad.mean():.2f}")
print(f"CV-then-resample (honest): recall = {good.mean():.2f}")
Output
resample-then-CV (leaky):  recall = 0.75
CV-then-resample (honest): recall = 0.62

The walkthrough

Thirteen points of recall were imaginary. Same data, same model, same SMOTE. The leaky setup scores 0.75; the honest one 0.62. The gap is pure leakage, and it grows as datasets get smaller and imbalance gets worse.

Why the leaky version cheats twice. First, synthetic blends of test-fold rows appear in training folds. Second, the test folds themselves contain synthetic points, so the model is graded partly on easy interpolations rather than real cases.

imblearn.pipeline.Pipeline, not sklearn's. Scikit-learn's own Pipeline refuses samplers, because fit_resample changes the number of rows mid-stream. The imblearn version understands samplers and — the crucial part — applies them during fit only, never during predict or score.

This composes with tuning. Put the pipeline inside GridSearchCV and every hyperparameter combination is evaluated leak-free. You can even tune SMOTE itself: param_grid={"smote__k_neighbors": [3, 5, 10]}.

Common mistakes

Scaling before splitting. The same leak, milder dose: StandardScaler().fit(X) learns the test rows' statistics. Always fit preprocessing on training folds only. Pipelines make the right way the lazy way.

Stratification forgotten. With 5% positives and plain KFold, a fold can end up with almost no positives, making recall jump wildly between folds. cross_val_score stratifies automatically for classifiers; if you build folds by hand, use StratifiedKFold.

Reporting the mean without the spread. Print good.std() too. Rare-class metrics on small folds swing hard, and a mean of 0.62 ± 0.15 is a different claim from 0.62 ± 0.02.

Trusting random_state to hide the problem. The leak is structural. No seed choice fixes it; only fold-wise resampling does.

Try it yourself

Set weights=[0.99, 0.01] and rerun. The leaky-versus-honest gap widens as positives get scarcer. Then swap SMOTE for RandomUnderSampler and check whether undersampling leaks the same way — reason about why before running.

What to learn next

Researcher — Mathematics and papers.

Why the bias is mechanical

Let D be the dataset and S(D) the SMOTE-augmented version. A synthetic point x_new = x_i + λ(x_zi − x_i) depends on two real points. Under CV after augmentation, the event "x_i lands in a training fold while x_new lands in the test fold" has high probability: with k folds, any dependent pair splits across folds with probability (k−1)/k. Test items are then interpolants of training items, so the estimated risk is evaluated partly on a distribution concentrated inside the training set's convex hull — an optimistic bias, not noise. It does not vanish with more folds; (k−1)/k rises toward 1.

The general principle: cross-validation estimates the risk of the entire learning procedure. Any data-dependent transform is part of that procedure and must be refit per fold. Formal treatments: Stone (1974) on CV validity; Varma and Simon (2006), Bias in error estimation when using cross-validation for model selection; Kaufman et al. (2012), Leakage in data mining, KDD, which taxonomises leakage sources.

Measured magnitudes

Vandewiele et al. (2021), Overly optimistic prediction results on imbalanced data, audited 24 published studies on a preterm-birth dataset: several reported near-perfect AUCs traceable to oversampling before partitioning; correcting the order dropped results to modest values. Santos et al. (2018) report the same pattern across UCI datasets. The bias scales with imbalance ratio, SMOTE's k, and inversely with dataset size.

Nested cross-validation

Selecting the sampler, ratio, or k by the same CV loop that reports the final score introduces a second-order optimism (model-selection bias). The clean protocol is nested CV: an inner loop chooses hyperparameters, an outer loop — which never influenced any choice — reports performance. Cost multiplies: k_outer × k_inner × |grid| fits. For imbalance work with small positive counts, prefer repeated stratified k-fold (e.g. 5×5) to reduce variance of the outer estimate; see also leakage discussions in model evaluation.

Where resampling genuinely may live outside

Group- and time-structure change the rule's application, not the rule: with TimeSeriesSplit, resample within each training window; with grouped data (GroupKFold), resample within training groups so no patient contributes to both sides. The invariant is always the same — the test partition is sampled from the deployment distribution and is never an input to any fitted transform.

What to learn next