Reproducibility and Running Experiments

Is this improvement real?

A paired test over matched seeds tells you whether your one-point gain is a genuine improvement or a lucky draw — and most one-point gains are luck.

On this page 5
  1. Why it exists
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Before celebrating a small improvement, ask: could pure luck have produced a gap this size? Often the honest answer is yes.

Two friends throw a die ten times each. One totals 38, the other 33. Is the first friend "better at dice"? Nobody would say so — five points of difference is exactly what luck produces in ten throws. Now replace dice with models and totals with scores: model B beats model A by half a point. The question is identical, and so is the danger of a silly answer.

Model scores wobble with the seed, the split, the shuffle. Any two honest runs differ a little. So "B scored higher than A" is not yet news. The news question is: is the gap bigger than the wobble?

Why it exists

Careers, papers and product decisions hang on comparisons: new model versus old, our method versus theirs. The wobble in a typical evaluation is often a point or two — the same size as many celebrated improvements.

Without a wobble-aware check, teams ship sideways changes and papers publish coin-flips as progress. The next person cannot build on the "gain", because it was never there. The check exists to protect you from your own hopeful eyes.

How it works

ask luck directly:

gap you observed:            model B is 1.3 points ahead of model A
wobble of the gap, measured
across repeated fair matches:  give or take 2.5 points

verdict: a 1.3-point gap sits well inside a 2.5-point wobble
         → luck produces gaps like this routinely
         → do not celebrate yet

The fair-match part matters. Compare A and B on the same splits with the same seeds, like judging two batsmen on the same pitches and the same bowlers. Matched comparisons cancel shared luck, so the wobble that remains belongs to the difference itself.

A real example you have seen

Medicine does this by law. A drug must beat a sugar pill by more than patient-to-patient luck can explain — otherwise pharmacies would stock lucky sugar. The phrase you hear in news reports, "statistically significant", means exactly: the gap was too big for luck to be the boring explanation.

Remember this

  • Every score has wobble; a gap smaller than the wobble is not yet a result.
  • Compare on matched seeds and splits, so luck cancels fairly.
  • "Not significant" means "not proven yet" — collect more evidence or claim less.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install scikit-learn scipy

Verified with scikit-learn 1.7.2, scipy 1.14.1, numpy 1.26.4, CPU. Seeded, so your numbers should match.

A paired test over ten matched seeds

Same data, same seeds, both models per seed — then two tools: a paired t-test (which asks how surprising the mean gap is, given its wobble) and a bootstrap interval (which shows the range of gaps consistent with the evidence).

is_it_real.py
import numpy as np
from scipy.stats import ttest_rel
from sklearn.datasets import make_classification
from sklearn.ensemble import GradientBoostingClassifier, RandomForestClassifier
from sklearn.model_selection import train_test_split

X, y = make_classification(n_samples=300, n_features=10, n_informative=4,
                           random_state=42)
forest, boosting = [], []
for seed in range(10):
    Xtr, Xte, ytr, yte = train_test_split(X, y, random_state=seed)
    forest.append(RandomForestClassifier(random_state=seed).fit(Xtr, ytr).score(Xte, yte))
    boosting.append(GradientBoostingClassifier(random_state=seed).fit(Xtr, ytr).score(Xte, yte))

diff = np.array(boosting) - np.array(forest)
t, p = ttest_rel(boosting, forest)
print(f"mean difference: {diff.mean():+.4f}")
print(f"paired t-test  : t = {t:.2f}, p = {p:.3f}")

rng = np.random.default_rng(0)
boots = np.array([rng.choice(diff, diff.size, replace=True).mean()
                  for _ in range(10000)])
lo, hi = np.percentile(boots, [2.5, 97.5])
print(f"95% bootstrap interval for the difference: [{lo:+.4f}, {hi:+.4f}]")
Output
mean difference: +0.0013
paired t-test  : t = 0.10, p = 0.921
95% bootstrap interval for the difference: [-0.0227, +0.0253]

The walkthrough

The verdict reads: boosting leads by 0.0013 — one tenth of one point. The p-value of 0.921 says: if the models were exactly equal, luck would produce a gap at least this big 92% of the time. The bootstrap interval says the true difference plausibly lies anywhere from −2.3 to +2.5 points. Verdict: indistinguishable. (These are the same two models whose per-seed scores ranged over 17 points in the seed-variance lesson — full circle.)

Why paired matters: ttest_rel tests the per-seed differences, so everything the two models share on a seed — an easy split, a hard one — subtracts out. An unpaired test (ttest_ind) would drown the signal in shared split-luck and need far more seeds for the same answer.

The interval says more than the p-value. A p-value answers one yes/no question. The interval shows what sizes of effect remain believable — here, anything within about ±2.5 points. If a two-point gain would matter for your product, this experiment cannot yet see gains that size: run more seeds, or use more data.

Reading a small p-value. Had p come out at 0.01 with a mean gap of +0.02, you would say: luck rarely makes gaps this big, the gain is probably real — and then still ask whether two points matters for the cost. Significance is not importance.

Common mistakes

Testing after peeking at many comparisons. Run twelve variants, test the best against the baseline: that p-value is broken — the winner was selected for its gap (best-of-many strikes again). Either correct for the twelve looks (multiply p by 12 — the crude Bonferroni fix) or confirm the winner on fresh seeds.

Ten CV folds treated as ten independent runs. Cross-validation folds share training rows, which makes the classic test overconfident. Prefer independent seeds/splits as above; when folds are all you have, use the corrected test in the researcher block.

"p = 0.06, so the improvement is fake." Not proven is not disproven. The interval tells you what you can still believe; act accordingly, or gather more evidence.

Significance without size. With enough seeds, a 0.05-point gain becomes significant — and remains a 0.05-point gain. Report both: the size (with its interval) and the p-value.

Try it yourself

Handicap the forest: RandomForestClassifier(n_estimators=10, max_depth=3, random_state=seed). Rerun. Watch the mean difference, the p-value and the interval all change together — and practise writing the one-sentence verdict for the new output.

What to learn next

Researcher — Mathematics and papers.

What the paired t-test assumes

For per-seed differences $d_i = S^B_i - S^A_i$, $i = 1..k$, the statistic is

$$ t = \frac{\bar{d}}{s_d / \sqrt{k}} $$

Where:

  • $\bar{d}$ — mean difference; $s_d$ — sample standard deviation of the $d_i$.
  • $k$ — number of matched runs (here 10); degrees of freedom $k - 1$.

Assumptions: differences approximately normal (mild, by CLT-in-miniature; the bootstrap cross-checks it) and — critically — independent across $i$. Independent seeds with disjoint splits approximate this; overlapping training sets do not.

The overlap problem and its corrections

  • Dietterich (1998), Approximate Statistical Tests for Comparing Supervised Classification Learning Algorithms, Neural Computation: the field's foundational audit. The resampled t-test has badly inflated Type-I error; recommended alternatives are 5×2cv paired t (five 2-fold CVs — folds within a replication share no training data) and McNemar's test for single-split comparisons of two fitted classifiers on the same test items.
  • Nadeau and Bengio (2003), Inference for the Generalization Error, Machine Learning: overlapping train sets induce positive correlation $\rho$ between fold estimates that does not vanish with more folds; their corrected resampled t-test replaces $s_d^2/k$ with $s_d^2 (1/k + n_2/n_1)$, where $n_2/n_1$ is the test/train size ratio. The default choice when your unit of repetition is a CV fold.
  • Bouthillier et al. (2021), MLSys: modern measurement of the variance components and a practical recipe — randomise everything per run, then paired comparisons over runs, essentially the protocol coded above.

Beyond two models

Comparing many models across many datasets: Demšar (2006), Statistical Comparisons of Classifiers over Multiple Data Sets, JMLR — Wilcoxon signed-ranks for two models (rank-based, no normality), Friedman test plus Nemenyi critical-difference diagrams for several; still the standard citation for benchmark tables. Multiple comparisons against a baseline want familywise control (Holm–Bonferroni) — the testing-side twin of the selection effects in feature-selection leakage.

Bootstrap details and honest alternatives

The percentile bootstrap over per-seed differences (10⁴ resamples) is distribution-free and gives the interval practitioners actually need. With $k = 10$, intervals are wide and slightly undercover; BCa corrections help marginally, more seeds help more. For per-example paired comparisons on one test set (large $n$, two prediction vectors), bootstrap over test examples instead — resolution improves with $n$ rather than $k$, and McNemar (exact, on the discordant pairs $b, c$: under the null $b \sim \mathrm{Binomial}(b+c, \tfrac{1}{2})$, so the two-sided $p = \min!\big(1,\; 2 \cdot 2^{-(b+c)} \sum_{j \le \min(b,c)} \binom{b+c}{j}\big)$ — the factor 2 is the second tail, and dropping it is the commonest way this test is quoted wrongly) is the classical equivalent.

Bayesian alternatives (Benavoli et al., 2017, Time for a Change: a Tutorial for Comparing Multiple Classifiers through Bayesian Analysis, JMLR) replace the p-value with posterior probabilities of {A better, B better, practically equivalent} given a ROPE (region of practical equivalence) — directly answering the product question "does a gap this small matter?", at the cost of choosing the ROPE. Whatever the machinery, the report that serves readers is constant: effect size, uncertainty, protocol, and the number of comparisons made before this one.

What to learn next