Reproducibility and Running Experiments

Report a range, not a single number

One training run is one roll of the dice — run several seeds and report the spread, or your "improvement" may be luck wearing a lab coat.

On this page 5
  1. Why it exists
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

A single training run is one dice roll — judge models by several rolls, and report the spread.

Would you pick a cricket team based on one innings per player? One player scores 80, another scores 12, and you conclude the first is better. But batting has luck in it: one good ball, one dropped catch. Selectors look at a whole season — the average, the consistency, the worst day — because one innings mostly measures luck.

Training a model is an innings. Randomness enters through the starting weights, the shuffling of data, and which rows land in the test split. Change the seed — the number that steers all that randomness — and the same code produces a different score.

Why it exists

Machine learning culture loves single numbers. Leaderboards rank by one score. Papers bold one number per row. Team chats announce "we got 94.2!". Single numbers are easy to compare and easy to celebrate.

But re-running with a different seed may give 92.8. Then the gap between your model and the old one is smaller than the gap between two runs of the same model. Any comparison that ignores the spread cannot tell improvement from luck.

How it works

same code, ten seeds:

model A:  87 84 89 86 88 85 87 90 86 88   → typically 84–90
model B:  91 83 85 86 88 84 87 85 86 85   → typically 83–91

one-run story : "B wins, 91 vs 87!"
ten-run story : the ranges overlap almost completely — no real winner

The honest report is a middle value plus a spread: "around 87, ranging 84 to 90 over ten seeds". The spread is not decoration. It is the yardstick that decides whether any gap means anything.

A real example you have seen

Opinion polls report "42%, plus or minus 3%". The pollsters are not being humble — they are telling you which differences are real. A rival on 44% is within the wobble; a rival on 52% is not. Model scores need the same plus-or-minus, for the same reason.

Remember this

  • One run measures your model and your luck, blended.
  • Run 3–10 seeds; report the middle and the spread.
  • A gap smaller than the spread is a coin toss, not a result.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install scikit-learn

Verified with scikit-learn 1.7.2, numpy 1.26.4, CPU. Seeded, so your numbers should match. Runtime is a few seconds.

Two models, ten seeds, one honest verdict

The seed controls both the split and the model's internal randomness. Everything else is identical.

ten_seeds.py
import numpy as np
from sklearn.datasets import make_classification
from sklearn.ensemble import GradientBoostingClassifier, RandomForestClassifier
from sklearn.model_selection import train_test_split

X, y = make_classification(n_samples=300, n_features=10, n_informative=4,
                           random_state=42)
scores = {"forest": [], "boosting": []}
for seed in range(10):
    Xtr, Xte, ytr, yte = train_test_split(X, y, random_state=seed)
    scores["forest"].append(
        RandomForestClassifier(random_state=seed).fit(Xtr, ytr).score(Xte, yte))
    scores["boosting"].append(
        GradientBoostingClassifier(random_state=seed).fit(Xtr, ytr).score(Xte, yte))

for name, s in scores.items():
    s = np.array(s)
    print(f"{name:9s} mean {s.mean():.3f}   worst {s.min():.3f}   best {s.max():.3f}")
wins = sum(f > b for f, b in zip(scores["forest"], scores["boosting"]))
print(f"forest wins {wins} of 10 seeds")
Output
forest    mean 0.867   worst 0.787   best 0.920
boosting  mean 0.868   worst 0.800   best 0.960
forest wins 3 of 10 seeds

The walkthrough

The means differ by 0.001. These models are, on this data, indistinguishable. Now look at what single runs could have told you: boosting's best seed scores 0.960 and the forest's worst scores 0.787. Two unlucky choices of seed and you would report a seventeen-point difference between equals. That is the whole lesson in two numbers.

The seed goes into both the split and the model. Both are real sources of wobble, and using the loop index for both keeps each run's world self-consistent. For a deeper look at what a seed controls, see seeds and reproducible runs.

forest wins 3 of 10 is a usefully humbling statistic. If one model were genuinely better, it should win most seeds. A 3–7 or 5–5 split of wins is what "no real difference" looks like in practice.

How many seeds? Three is the defensible minimum, five is comfortable, ten is generous. The cost is linear; the embarrassment saved is not. For expensive runs, report all the seeds you can afford — and say how many that was.

Reporting formats that work

  • In text: 86.7 ± 3.3 (mean ± std over 10 seeds) — the forest numbers above — and say which spread measure you used.
  • In tables: a ± column, or worst/best columns. Bold nothing that overlaps.
  • In plots: error bars or per-seed dots. Dots are honest and cheap.

Common mistakes

Reporting the best seed. Selecting the luckiest run and publishing it is test-set overfitting with extra steps. If you must pick a seed for deployment, pick it on validation data — and report test spread across seeds anyway.

Comparing your multi-seed mean against a paper's single run. Their number contains luck of unknown sign. Note the mismatch, or re-run their method under your protocol. Why your reimplementation is three points worse digs into this.

Averaging over seeds but not over splits. With small datasets the split matters more than the model's internal randomness. Vary both — as the code above does — or use cross-validation inside each seed.

Treating ± as proof. Overlapping ranges suggest no difference; they do not prove it, and narrow non-overlap is not automatic victory. When the decision matters, run the test in is this improvement real?

Try it yourself

Increase n_samples from 300 to 3000 and rerun. Watch both spreads shrink. Data size buys you certainty — which is why small-data claims need more seeds, not fewer.

What to learn next

Researcher — Mathematics and papers.

Decomposing the variance

For a score $S$ produced by a stochastic training procedure on sampled data,

$$ \operatorname{Var}(S) = \underbrace{\operatorname{Var}{\text{data}}\big(\mathbb{E}[S \mid D]\big)}{\text{sampling of train/test}} + \underbrace{\mathbb{E}{\text{data}}\big(\operatorname{Var}[S \mid D]\big)}{\text{optimisation noise}} $$

Where:

  • $S$ — the evaluation score of one run.
  • $D$ — the realised data split.
  • The first term — variance from which rows landed where.
  • The second — variance from initialisation, shuffling, dropout, nondeterministic kernels.

Bouthillier et al. (2021), Accounting for Variance in Machine Learning Benchmarks, MLSys, measure both terms across vision and NLP tasks: data-split variance typically dominates optimisation noise, and many published architecture gaps sit inside the combined spread. Randomising everything (splits included) across runs, as the demo does, samples the full variance rather than the flattering half.

Field evidence

  • Reimers and Gurevych (2017), Reporting Score Distributions Makes a Difference, EMNLP: LSTM NER results where seed-induced spread exceeded many claimed SOTA gaps; they argue for distribution reporting as standard.
  • Henderson et al. (2018), Deep Reinforcement Learning that Matters, AAAI: RL results where different sets of five seeds of one algorithm differ significantly — the seed problem at its most extreme.
  • Picard (2021), torch.manual_seed(3407) is all you need, arXiv:2109.08203: a joke title over serious data — scanning $10^4$ seeds on CIFAR-10 finds spreads wide enough to manufacture publishable "gains".
  • Dodge et al. (2019), Show Your Work, EMNLP: expected-best-validation curves as a function of tuning budget — the reporting format that makes luck and budget explicit.

Standard errors and what ± should mean

The standard error of a mean over $k$ seeds is $s/\sqrt{k}$, with $s$ the sample standard deviation. Report the pair $(k, s)$ or the pair (mean, SEM) — but never a bare ±, since readers cannot tell std from SEM from range. For comparisons, per-seed pairing (same seed, both models, same split) cancels the shared data-variance component and sharpens the contrast — the mechanism exploited by the paired tests in statistical significance in ML.

Note the caveat attached to resampled estimates: overlapping train sets across CV folds violate independence, biasing naive SEMs — Nadeau and Bengio (2003), Inference for the Generalization Error, Machine Learning, give the corrected variance. Seeds with disjoint randomness (the protocol here) avoid that particular trap; the correction matters once folds enter.

The cost argument, pre-empted

The standard objection — "we cannot afford ten runs" — has a standard reply: then you cannot afford the claim. Choose a claim size measurable within budget: larger test sets, paired seeds, or coarser conclusions ("comparable", not "better"). A wrong bolded number costs more than nine retrainings; Sculley et al. (2018), Winner's Curse?, ICLR workshop, catalogue the field-level price.

What to learn next