Choosing a cross-validation strategy
Cross-validation scores your model on several held-out slices, and choosing the wrong slicer — plain, stratified, grouped or time-ordered — quietly breaks the score.
- 7 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Cross-validation tests a model several times, each time on a different held-out slice, and the right way to slice depends on how your data hangs together.
Think of tasting a big pot of biryani. One spoonful from the top tells you almost nothing — the salt may have settled, the top may be drier. So you taste from several spots and judge from all of them together.
A single train-test split is one spoonful. Cross-validation takes several, by rotating which slice is held out for tasting.
Why it exists
One test score can be lucky or unlucky, especially with small data. Averaging over several held-out slices steadies the estimate.
But there is a second, sneakier problem: how you slice can break the test. If all the interesting rows land in one slice, some rounds test nothing. If related rows — the same patient, the same week — land on both sides, the model gets to recognise instead of generalise. Each situation has a matching slicer.
How it works
data: [ A ][ B ][ C ]
round 1: train on B+C, taste A
round 2: train on A+C, taste B
round 3: train on A+B, taste C
final score = average of the three tastesFour slicers cover most real situations:
- Plain: rotate slices in order. Fine when rows are unrelated and well mixed.
- Stratified: force every slice to contain rare cases in fair proportion.
- Grouped: keep all rows from one group — one patient, one shop — in the same slice.
- Time-ordered: always train on the past and taste the future, never the reverse.
A real example you have seen
Cricket selectors do not judge a batsman on one innings. They look at several matches — and importantly, across different conditions: home, away, spin, pace. Judging only on home pitches is the grouped-data mistake: great average, wrong conclusion.
Remember this
- Cross-validation = several rotating held-out slices, scores averaged.
- The slicer must match the data: rare classes → stratified, related rows → grouped, time order → past-only.
- A wrong slicer fails silently — the score looks fine and lies.
What to learn next
- Pipelines and why leakage disappears — the other half of an honest score.
- Model evaluation — which metric to average across those folds.
- Imbalanced data — when 3-in-12 becomes 3-in-10000.
Developer — Code and libraries.
Setup
pip install scikit-learnTested against scikit-learn 1.7.
Four splitters on the same 12 rows
import numpy as np
from sklearn.model_selection import GroupKFold, KFold, StratifiedKFold, TimeSeriesSplit
X = np.zeros((12, 1)) # features do not matter for splitting
y = np.array([0]*9 + [1]*3) # 12 loans, only 3 defaults
print("KFold - default labels in each test fold:")
for _, test in KFold(n_splits=3).split(X):
print(" ", y[test])
print("StratifiedKFold - same data:")
for _, test in StratifiedKFold(n_splits=3).split(X, y):
print(" ", y[test])
patients = np.array([1,1,1, 2,2,2, 3,3,3, 4,4,4]) # 3 scans per patient
print("GroupKFold - patients in each test fold:")
for _, test in GroupKFold(n_splits=2).split(X, y, groups=patients):
print(" ", patients[test])
print("TimeSeriesSplit - index ranges:")
for train, test in TimeSeriesSplit(n_splits=3).split(X):
print(f" train {train.min()}..{train.max()} test {test.min()}..{test.max()}")KFold - default labels in each test fold: [0 0 0 0] [0 0 0 0] [0 1 1 1] StratifiedKFold - same data: [0 0 0 1] [0 0 0 1] [0 0 0 1] GroupKFold - patients in each test fold: [2 2 2 4 4 4] [1 1 1 3 3 3] TimeSeriesSplit - index ranges: train 0..2 test 3..5 train 0..5 test 6..8 train 0..8 test 9..11
The walkthrough
Plain KFold walked into a wall. The data was sorted — defaults at the end, as exported tables often are — and KFold slices in order by default. Two of three folds contain zero defaults, so those rounds test nothing about catching defaults. KFold(n_splits=3, shuffle=True, random_state=0) fixes the sorted-data case; stratification fixes it structurally.
StratifiedKFold rationed the rare class. Each fold got exactly one default, matching the overall 3-in-12 rate. For classification, this is the sane default — and cross_val_score already stratifies when it sees a classifier.
GroupKFold never split a patient. All three scans of patient 2 travel together. Without this, the model sees a patient's scan in training and another scan of the same patient in testing. It scores well by recognising the patient — a leak the pipeline cannot catch, because it lives in the split itself.
TimeSeriesSplit only ever trains on the past. Training windows grow: 0..2, then 0..5, then 0..8, each tested on the slice after. No round tests the past using the future. For forecasting, any other splitter is fiction — see what is time series.
Plugging a splitter into scoring is one argument:
from sklearn.model_selection import cross_val_score
scores = cross_val_score(model, X, y, cv=StratifiedKFold(n_splits=3))Common mistakes
Trusting the unshuffled default on sorted data. The first output block is the warning. If your CSV arrived sorted by label, date or customer, plain in-order slicing is broken from the start.
Ignoring groups you know exist. Multiple rows per customer, per device, per hospital — cross-validated without groups, these inflate scores like the twin-photos problem. If a column identifies who generated the row, it probably belongs in groups.
Shuffling time away. shuffle=True on time-ordered data trains on Wednesday to predict Tuesday. Scores soar; the deployed model faces only the real direction and disappoints.
Reporting the mean without the spread. scores.mean() alone hides instability. [0.9, 0.6, 0.9] and [0.8, 0.8, 0.8] share a mean of 0.8 and deserve different levels of trust. Print scores.std() alongside, always.
Try it yourself
Extend to 6 patients with 2 scans each and run GroupKFold(n_splits=3). Verify no patient ever straddles folds. Then break it on purpose: run StratifiedKFold with the same data and count how many patients appear on both sides of some split.
What to learn next
- Pipelines and why leakage disappears — the other half of an honest score.
- Model evaluation — which metric to average across those folds.
- Imbalanced data — when 3-in-12 becomes 3-in-10000.
Researcher — Mathematics and papers.
What k-fold actually estimates
k-fold CV averages the loss of models trained on n(k−1)/k samples, so it estimates the expected prediction error at a smaller training size than n — a pessimistic bias that shrinks as k grows. Raising k toward n (leave-one-out) trims bias but the k fitted models become nearly identical, and the estimate inherits high variance for unstable learners. Kohavi (1995), A study of cross-validation and bootstrap for accuracy estimation and model selection, is the empirical basis for the folk default k ∈ {5, 10}.
A sharper caveat: Bengio and Grandvalet (2004), No unbiased estimator of the variance of k-fold cross-validation, prove exactly that. Fold scores are dependent (shared training data), so the naive std/√k error bar is not unbiased — treat printed spreads as diagnostics, not confidence intervals.
Exchangeability is the real assumption
Every i.i.d.-style splitter assumes test rows are exchangeable with future production rows. Structured data breaks this two ways:
- Group structure: rows share a latent factor (patient, site, session). Splitting within groups estimates within-group generalisation, while deployment demands across-group. The gap can be enormous in medical imaging — model-recognises-the-scanner is a documented failure mode.
GroupKFold(orStratifiedGroupKFold, 1.0+) restores the across-group estimand. - Serial dependence: for temporal data the estimand is h-step-ahead error. Forward-chaining (what
TimeSeriesSplitimplements) is the standard evaluation; Bergmeir and Benítez (2012), On the use of cross-validation for time series predictor evaluation, analyse when blocked/forward schemes are required. Agapparameter mitigates boundary contamination when features contain lagged windows.
Model selection on top of estimation
Selecting hyperparameters by CV and reporting the winning score reintroduces optimism — the winner partially fit the folds' noise. Varma and Simon (2006), Bias in error estimation when using cross-validation for model selection, quantify the effect; nested CV (an outer loop scoring the whole selection procedure) is the corrected design. In scikit-learn: cross_val_score(GridSearchCV(pipe, grid, cv=inner), X, y, cv=outer).
Cost
k-fold multiplies training cost by k (plus refit-on-all for the deployed model). Grouped and stratified variants add only O(n log n) split computation. The expensive part is never the splitter — it is the honesty.
What to learn next
- Pipelines and why leakage disappears — the other half of an honest score.
- Model evaluation — which metric to average across those folds.
- Imbalanced data — when 3-in-12 becomes 3-in-10000.