Reproducibility and Running Experiments

When you cannot reproduce your own number

Last week's 94% refuses to come back — here is the systematic hunt through seeds, data, code and environment that finds where the difference crept in.

On this page 5
  1. Why it exists
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

When an old result will not come back, something changed — your job is to find which of four things: randomness, data, code, or environment.

You made a wonderful cup of tea last Tuesday. Today, same leaves, same method — and it tastes different. Before blaming ghosts, you list what could have changed: the water, the milk, the boiling time, the leaves' age. Then you check each one. The difference is always somewhere, because tea does not change by magic.

Neither do models. A score that will not reproduce is not mysterious. It is a "spot the difference" puzzle between two runs. The differences hide in exactly four places.

Why it exists

This situation is a rite of passage. You reported 94% last week. Today, preparing to present, you rerun the script: 91%. Panic arrives with three questions: Was last week wrong? Is today wrong? What do I tell people?

The panic comes from treating the run as one indivisible thing. It is not. It is randomness plus data plus code plus environment, and each piece can be checked separately.

How it works

last week: 94%          today: 91%
        └──── what changed? ────┘

1. RANDOMNESS  seed not fixed? different shuffle, split, weights
2. DATA        rows added? file re-exported? different filter?
3. CODE        any edit since — including "harmless" ones?
4. ENVIRONMENT library updated? different machine? GPU vs CPU?

check them in that order: cheapest first, guiltiest first.

Randomness is checked in a minute and is the culprit most often. Data drifts silently in shared systems. Code you can interrogate with version control. Environment is the rarest and the sneakiest.

A real example you have seen

"But it worked yesterday!" — the universal cry, from homework to banking apps. Support teams ask the same four questions in the same order: did you change anything? (code) — did an update install? (environment) — is it the same file? (data) — does it happen every time? (randomness). The hunt below is that call script, made rigorous.

Remember this

  • An unreproducible number is a difference-hunt across four suspects: randomness, data, code, environment.
  • Check randomness first — cheapest to test, most often guilty.
  • The prevention is writing down all four at run time, so future-you can compare.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install scikit-learn

Verified with scikit-learn 1.7.2, numpy 1.26.4, Python 3.10, CPU.

Suspect 1, demonstrated: the unfixed seed

One missing random_state and "the" accuracy of this script is a fiction — it never had one:

five_identical_runs.py
from sklearn.datasets import make_classification
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import train_test_split

X, y = make_classification(n_samples=200, n_features=8, random_state=0)
for run in range(5):
    Xtr, Xte, ytr, yte = train_test_split(X, y)   # no random_state
    acc = LogisticRegression(max_iter=1000).fit(Xtr, ytr).score(Xte, yte)
    print(f"run {run}: accuracy {acc:.3f}")
Output
run 0: accuracy 0.940
run 1: accuracy 0.920
run 2: accuracy 0.960
run 3: accuracy 0.960
run 4: accuracy 0.940

These five numbers came from one execution on the author's machine — your five will differ, and differ again on every rerun. That is the bug speaking. Last week's 94% and today's 91% may both be honest draws from this lottery. The diagnosis test: run today's script five times. If the spread covers both numbers, the mystery dissolves — the fix is fixing seeds, then reporting the spread.

The hunt, suspect by suspect

Randomness. Fix every seed (random_state everywhere; in deep learning, seed all libraries). Then run twice. Identical output? Randomness was the difference. Still varying with seeds fixed? A hidden stream remains — data loader workers, hash-ordered sets, OS-level parallelism.

Data. Compare row counts, then a content fingerprint — a hash of the sorted data, or pandas' pd.util.hash_pandas_object(df).sum(). Shared tables and re-exported CSVs drift without telling you. This is why data versioning exists.

Code. git diff <last-week's-commit> — provided you committed and recorded which commit produced the number. The run-record pattern with a git hash makes this a lookup instead of an argument with your memory.

Environment. Capture and compare a manifest:

manifest.py
import json
import platform
import sys
import numpy
import sklearn

manifest = {
    "python": platform.python_version(),
    "numpy": numpy.__version__,
    "sklearn": sklearn.__version__,
    "os": platform.system(),
    "command": " ".join(sys.argv),
}
print(json.dumps(manifest, indent=2))
Output
{
  "python": "3.10.11",
  "numpy": "1.26.4",
  "sklearn": "1.7.2",
  "os": "Windows",
  "command": "manifest.py"
}

The values above are the author's machine — yours will differ, which is precisely the point: print this at the start of every training run, into the log. Library minor versions change defaults, solvers and random streams; scikit-learn documents such changes per release, which is why pinning versions (sklearn==1.7.2, not sklearn) belongs in every project. For bit-level environment capture, Docker is the heavier hammer.

Common mistakes

Rerunning once and declaring a crisis. With any unfixed randomness, two runs should differ. Establish the spread (five runs) before comparing anything to anything.

Hunting code first. Code is where you look first because git makes it visible — but randomness and data are guilty more often. Cheap checks first.

"I only changed something unrelated." Unrelated edits move random streams: adding one augmentation consumes random draws and shifts everything downstream, changing results without any real change. Seed-stream isolation (separate generators per component) contains this.

Blaming the GPU too early. GPU nondeterminism is real but usually small per step. Rule out the boring suspects before invoking exotic ones — and when you do get there, torch.use_deterministic_algorithms(True) is the test.

Try it yourself

Fix the split (random_state=0) and rerun five times — identical numbers. Now change max_iter from 1000 to 50 and run again: a different, stable number. You have manufactured a clean example of suspect 3 — a code change masquerading as noise — and seen how fixed seeds unmask it.

What to learn next

Researcher — Mathematics and papers.

A taxonomy with measurements

Distinguish reproducibility levels: same code + same data (repeatability), same code + fresh randomness (replicability under noise), independent reimplementation (reproducibility proper — the concern of reimplementing a paper). This lesson is the first two.

Pham et al. (2020), Problems and Opportunities in Training Deep Learning Software Systems: An Analysis of Variance, ASE: identical training code rerun with all controllable seeds fixed still varied — implementation-level nondeterminism (cuDNN algorithm selection, floating-point atomics, parallel reduction order) produced accuracy differences up to several percent on real systems, and 83% of surveyed practitioners were unaware of the magnitude. Zhuang et al. (2022), Randomness in Neural Network Training: Characterizing the Impact of Tooling, MLSys, decompose the stack's contributions and measure the cost of enforcing determinism (often 10–30% slowdown, hardware-dependent).

The fixed points, formally

A run computes $R = f(\text{code}, \theta, D, \omega, s)$ with $s$ the tuple of random streams and $\omega$ the execution environment. Bit-identical repetition requires fixing all five coordinates and $f$ being a deterministic function — which parallel floating-point arithmetic violates: addition reordering across threads changes results within $\varepsilon$-rounding, and training dynamics amplify $\varepsilon$ over steps (the practical chaos noted in the PyTorch seeds lesson).

Where:

  • $s$ — every random stream: splitters, initialisers, shufflers, augmenters, dropout.
  • $\omega$ — interpreter, library versions, BLAS backend, thread counts, hardware.
  • $\varepsilon$ — floating-point rounding; non-associativity is the root cause.

Python-specific $\omega$ trap: PYTHONHASHSEED randomises string hashing per process, so iteration order of sets and unordered dicts of strings varies across executions, silently reordering any pipeline that iterates a set. Fix it in the environment, or never let set order reach computation.

Environment capture in increasing strictness

  1. Manifest logging (the snippet): names and versions — catches the common case.
  2. Lockfiles (pip freeze, poetry/uv locks, conda-lock): pins the full dependency closure.
  3. Container images (Docker): pins system libraries and BLAS; the unit of exchange in most reproduction requests.
  4. Deterministic-kernel modes: torch.use_deterministic_algorithms(True), CUBLAS_WORKSPACE_CONFIG, disabling cuDNN benchmark — trades speed for bitwise stability within one hardware class.

Cross-hardware bitwise identity is generally unattainable (different GPUs schedule differently; TF32 vs FP32 paths differ); the defensible claim is statistical reproduction — same distribution over seeds, per Bouthillier et al. (2021), MLSys. State claims accordingly: "94.0 ± 0.8 over 5 seeds, torch 2.5.1, one A100" is reproducible; a bare "94.2" is a coordinate without a map.

Pineau et al. (2021), JMLR, close the loop institutionally: the NeurIPS checklist's environment and seed items exist because reviewers could not reproduce results missing them — the same four suspects, enforced at publication time.

What to learn next