Data Leakage and Results That Are Too Good

The same person in train and test

When several rows belong to one person and a random split scatters them, the model learns to recognise individuals instead of learning the task — split by group, not by row.

On this page 5
  1. Why it exists
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

If several rows come from the same person, splitting by row lets the model meet the same person on both sides of the exam.

A doctor claims she can diagnose a rare condition from handwriting. To test her, you show pages from patients she has already examined — different pages, same people. She recognises the handwriting, remembers each diagnosis, and scores brilliantly. Show her a stranger's page and she is lost.

She was never diagnosing. She was recognising. Models do the same when one person's rows land in both training and test.

Why it exists

Real datasets rarely have one row per person. A patient contributes five hospital visits. A speaker contributes fifty voice clips. One machine contributes a thousand sensor readings. Rows from the same source share a hidden fingerprint — the person's voice, the machine's vibration signature, the patient's anatomy.

A random split treats every row as independent and scatters each person's rows everywhere. The model then passes the test by matching fingerprints to remembered answers. This is group leakage: rows are different, sources are shared.

It is the sibling of duplicate contamination. There, identical rows crossed the split. Here, related rows cross it — harder to see, equally fatal.

How it works

patient A: rows a1 a2 a3      random split
patient B: rows b1 b2         ──────────────
                              train: a1 a3 b2   ← learns A's fingerprint
                              test : a2 b1      ← recognises A, recalls answer

group split:
                              train: a1 a2 a3   ← all of A together
                              test : b1 b2      ← B is a true stranger

The honest question your test must answer: does this work on people the model has never met? Only a group split asks it.

A real example you have seen

Face unlock on your phone trains on many photos of you. It works wonderfully — on you. The manufacturer cannot claim it recognises anyone by testing it on more photos of you. Any "works on new people" claim needs new people. Medical AI has repeatedly stumbled here, with scan-level splits inflating patient-level claims.

Remember this

  • Group leakage: multiple rows per person, scattered across the split.
  • The model passes by recognising individuals, not by learning the task.
  • Split so that each person's rows stay entirely on one side.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install scikit-learn

Verified with scikit-learn 1.7.2, numpy 1.26.4, CPU. Seeded, so your numbers should match.

Coin-flip labels, 75% accuracy

Sixty patients, five visits each. Each patient's label is a coin flip — there is nothing medical to learn. But each patient has a stable "signature" in their measurements, as real people do.

group_leak.py
import numpy as np
from sklearn.ensemble import RandomForestClassifier
from sklearn.model_selection import GroupKFold, KFold, cross_val_score

rng = np.random.default_rng(5)
patients, visits = 60, 5
pid = np.repeat(np.arange(patients), visits)

signature = rng.normal(0, 1, (patients, 4))[pid]      # stable per-patient traits
X = signature + rng.normal(0, 0.3, (patients * visits, 4))  # small visit-to-visit change
y = rng.integers(0, 2, patients)[pid]                 # one coin flip per patient

model = RandomForestClassifier(random_state=5)
naive = cross_val_score(model, X, y, cv=KFold(5, shuffle=True, random_state=5))
grouped = cross_val_score(model, X, y, cv=GroupKFold(5), groups=pid)
print("plain KFold accuracy:", naive.round(2), " mean", round(naive.mean(), 3))
print("GroupKFold  accuracy:", grouped.round(2), " mean", round(grouped.mean(), 3))
Output
plain KFold accuracy: [0.75 0.78 0.72 0.72 0.8 ]  mean 0.753
GroupKFold  accuracy: [0.63 0.42 0.45 0.62 0.37]  mean 0.497

Plain K-fold — here, cross-validation, scoring on five different held-out slices — reports 75% on labels that are pure chance. GroupKFold reports the truth: 50%.

The walkthrough

The leak mechanism is row similarity, not row identity. No two rows are equal; visit noise differs. But rows from one patient sit close together in feature space, and the forest maps that neighbourhood to the remembered label. Deduplication would find nothing here.

GroupKFold(5) with groups=pid guarantees no patient straddles a fold boundary. Every test patient is a stranger. That single argument is the entire fix.

Notice the honest scores wobble — 0.37 to 0.63. With 12 test patients per fold, coin-flip accuracy swings widely. Honest evaluation on grouped data has fewer effective samples than the row count suggests: 60 independent patients, not 300 independent rows. Wide wobble on grouped data is a sign the evaluation is working, not failing.

Finding your groups

The hard part in practice is realising a grouping exists. Ask what generated multiple rows:

  • People: patients, users, speakers, authors, students.
  • Devices: sensors, cameras, hospital scanners, vehicles.
  • Places and batches: farms, stores, manufacturing lots, lab plates.
  • Sessions: one user session producing many events.

If any such column exists — or could be reconstructed — evaluate with it as groups and compare against the plain split. A large gap is the leak's confession.

Common mistakes

Splitting scans instead of patients. The most published version of this bug. Three X-rays of one pneumonia patient, scattered across splits, teach scanner-and-anatomy recognition.

Grouping in the split but not in the tuning. Hyperparameter search with an ungrouped inner loop leaks the same way, one level down. Pass the same groups into GridSearchCV via its cv argument. The general trap is the next lesson, leakage that survives cross-validation.

Assuming anonymised data has no groups. Removing the ID column removes your ability to group correctly, not the underlying fingerprints. Get the IDs back, or construct proxies.

Stratifying instead of grouping. StratifiedKFold balances label proportions; it does nothing about identity. Since scikit-learn 1.3 you can have both: StratifiedGroupKFold.

Try it yourself

Raise the visit noise from 0.3 to 3.0 and rerun. Watch the leaky score fall towards 0.5 as fingerprints drown. Then explain why heavy per-row noise reduces group leakage — and why relying on that would be a terrible plan.

What to learn next

Researcher — Mathematics and papers.

Exchangeability, violated

Cross-validation estimates generalisation error assuming test points are exchangeable with future deployment points. With grouped data the deployment question is usually performance on unseen groups, so the exchangeable unit is the group, not the row. Row-level CV estimates a different quantity: performance on new rows from known groups. Both are legitimate targets — the crime is reporting one as the other.

Formally, with random effects: $x_{ij} = \mu_{g(i)} + \varepsilon_{ij}$, where $\mu_g$ is a group-level effect and $\varepsilon_{ij}$ row noise. Row-level splits allow the model to estimate $\mu_g$ from training rows and reuse it at test time; group-level splits force marginalisation over unseen $\mu_g$.

Where:

  • $g(i)$ — the group (person, device) that generated row $i$.
  • $\mu_g$ — the group's stable signature.
  • $\varepsilon_{ij}$ — within-group variation.

The demo sets the label to a function of $g$ alone, making the row-level estimate pure recognition — accuracy 0.75 against a true group-level value of 0.5.

Evidence from applied fields

  • Saeb et al. (2017), The need to approximate the use-case in clinical machine learning, GigaScience: record-wise versus subject-wise CV in digital-health studies; record-wise splitting produced dramatically optimistic accuracy in the surveyed literature.
  • Roberts et al. (2017), Cross-validation strategies for data with temporal, spatial, hierarchical, or phylogenetic structure, Ecography: the ecology community's systematic treatment — block CV as the general answer to dependence structure.
  • Kapoor and Narayanan (2023), Patterns: group leakage ("no separation between distributions") recurs across their 17-field survey of leakage-affected science.

Spatial autocorrelation is the continuous cousin: nearby locations share environment the way visits share a patient. Spatial block CV holds out contiguous regions rather than random pixels.

Effective sample size

With $G$ groups of $m$ rows and intraclass correlation $\rho$, the effective number of independent units is approximately

$$ n_{\text{eff}} = \frac{Gm}{1 + (m - 1)\rho} $$

  • $\rho$ — the fraction of variance attributable to the group effect.

As $\rho \to 1$, $n_{\text{eff}} \to G$: sixty patients, not three hundred rows. This governs both the honest evaluation's variance (the wobble in the output) and how many rows a paper's confidence intervals may legitimately claim. Ignoring it is a quieter cousin of the leak itself.

Tooling

scikit-learn: GroupKFold, GroupShuffleSplit, LeaveOneGroupOut, StratifiedGroupKFold (1.3+), and groups threading through cross_val_score and search estimators. For two crossed grouping factors (e.g., speakers × phrases), no stock splitter suffices; construct splits so both factors are unseen — the "zero-shot" evaluation design used in speaker verification and recommender cold-start, covered from the recommender side in the cold-start problem.

What to learn next