Reproducibility and Running Experiments
Configs instead of constants scattered everywhere
Put every knob of an experiment in one config file, and save that config with every result — so any number you ever produced can be traced and re-run.
- 7 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Every setting of an experiment belongs in one file — and that file gets saved with the result it produced.
Think of a grandmother who cooks by instinct: a pinch of this, "enough" of that, taste and adjust. The food is wonderful. It is also unrepeatable — even she cannot make the same dish twice, and nobody can learn it from her. A written recipe card fixes that: every quantity named, every step listed.
A training script can have numbers scattered through it. A learning rate — the size of each training step — on line 12, a tree depth on line 87, a threshold typed into a command. That is instinct cooking. A config file — one small file listing every setting by name — is the recipe card.
Why it exists
Three weeks from now, someone asks: "That 94% you mentioned — what settings produced it?" With scattered constants, the honest answer is a shrug. The script has changed since. The command you typed is lost to history. The number is now an orphan: real, but unrepeatable.
Experiments multiply this pain. You will run dozens of variations. If each variation means editing the script in three places, you will forget one, and two runs that look comparable will silently differ.
How it works
without configs: with configs:
train.py (edited daily) config.json ──→ train.py (never edited)
│ │ │
▼ ▼ ▼
"94%" ← from which version? result saved WITH its config
→ any result re-runnable, foreverThe rule has two halves. Settings live in the config, not the code. And every result is written down together with the exact config that made it.
A real example you have seen
A pharmacy will not accept "the usual, roughly" — prescriptions name the medicine, dose, and schedule, and a copy stays on record. When something goes wrong, the record answers "what exactly was taken?". Config files are prescriptions for experiments — boring on good days, priceless on bad ones.
Remember this
- Every knob — seed, sizes, rates, paths — lives in one config file.
- Every result is saved with its config, not near it.
- If reproducing a number requires memory, the experiment is not finished.
What to learn next
- Making sense of two hundred runs — what config-keyed records make possible.
- Experiment tracking — the industrial version of runs.jsonl.
- When you cannot reproduce your own number — the failure this lesson quietly prevents.
Developer — Code and libraries.
Setup
pip install scikit-learnVerified with scikit-learn 1.7.2, numpy 1.26.4, Python 3.10, CPU. Seeded, so your numbers should match.
The pattern, complete in two files
First, the config — every knob, named:
{
"n_estimators": 200,
"max_depth": 4,
"seed": 0,
"test_size": 0.25
}Then a script that reads it, trains, and appends a run record joining config to result:
import json
import pathlib
from sklearn.datasets import load_iris
from sklearn.ensemble import RandomForestClassifier
from sklearn.model_selection import train_test_split
cfg = json.loads(pathlib.Path("config.json").read_text())
X, y = load_iris(return_X_y=True)
Xtr, Xte, ytr, yte = train_test_split(X, y, test_size=cfg["test_size"],
random_state=cfg["seed"])
model = RandomForestClassifier(n_estimators=cfg["n_estimators"],
max_depth=cfg["max_depth"],
random_state=cfg["seed"]).fit(Xtr, ytr)
accuracy = round(model.score(Xte, yte), 4)
record = {"config": cfg, "accuracy": accuracy}
with open("runs.jsonl", "a") as f:
f.write(json.dumps(record) + "\n")
print("accuracy:", accuracy)
print("recorded:", json.dumps(record))accuracy: 0.9737
recorded: {"config": {"n_estimators": 200, "max_depth": 4, "seed": 0, "test_size": 0.25}, "accuracy": 0.9737}The walkthrough
The script contains no experiment-specific numbers. Change max_depth to 8? Edit the config, not the code. The script becomes stable infrastructure; the config becomes the experiment. Version control now shows experiments as one-line config diffs instead of scattered code edits.
runs.jsonl is the humble hero. One JSON line per run, config and result welded together. Grep it, load it into pandas, never lose an orphan number again. This is a two-line experiment tracker — when runs multiply, graduate to a real one (experiment tracking, MLflow); the principle stays identical.
The seed is a config value. The ten-seed protocol from the previous lesson becomes a shell loop that rewrites one field — no code edits between runs.
What belongs in the config? Everything you might ever vary, plus everything needed to re-run: data paths, split sizes, model settings, seeds, preprocessing choices. When in doubt, put it in. An unused config key is free; a hardcoded constant costs a future afternoon.
Levelling up, when you need it
- Defaults plus overrides: a base config, with per-experiment files overriding a few keys. Beyond that, tools like Hydra manage config composition — worth it around the time configs start importing each other.
- Record the environment too: library versions and the git commit make the record fully re-runnable. When you cannot reproduce your own number shows the manifest pattern.
- Validate on load: assert required keys exist and values are sane before training, not ninety minutes in. A dataclass with type hints does this almost for free.
Common mistakes
Command-line arguments as the config. --lr 0.001 --depth 8 --seed 3 works until the command disappears with your terminal history. Arguments are fine as overrides, but the resolved, final settings must be written into the run record either way.
Editing the config after the run. The record's copy is the truth; a mutated config.json rewrites history. Append-only records — as runs.jsonl is — make this mistake harmless.
Half-configs. Nine knobs in the file, the tenth hardcoded — the tenth is always the one that mattered. Grep your script for numeric literals as a config-completeness check.
Configs that contain code. When the config imports libraries and computes values, it inherits code's irreproducibility. Keep it data: JSON, TOML, YAML.
Try it yourself
Run the script three times, editing max_depth to 2, then 4, then 8 between runs. Load runs.jsonl with pandas (pd.read_json("runs.jsonl", lines=True)) and print a depth-versus-accuracy table. You have built the input for making sense of two hundred runs.
What to learn next
- Making sense of two hundred runs — what config-keyed records make possible.
- Experiment tracking — the industrial version of runs.jsonl.
- When you cannot reproduce your own number — the failure this lesson quietly prevents.
Researcher — Mathematics and papers.
Configuration as provenance
Formally, a run is a function $R = f(\text{code}, \theta, D, \omega)$: code version, configuration $\theta$, data snapshot $D$, and environmental state $\omega$ (library versions, hardware, nondeterminism). Reproducibility requires capturing all four coordinates; the config file pins $\theta$, and the run-record pattern welds $(\theta, R)$ so the mapping is never lost. The remaining coordinates are handled by version control (code), data versioning ($D$), and environment manifests ($\omega$).
Pineau et al. (2021), Improving Reproducibility in Machine Learning Research, JMLR — the NeurIPS reproducibility program — found missing experimental specification (exactly $\theta$) among the dominant reproduction blockers; their checklist item "range of hyperparameters considered and final values" is the config file as publication requirement. Gundersen and Kjensmo (2018), State of the Art: Reproducibility in AI, AAAI, surveyed 400 papers: a minority documented enough of $\theta$ to attempt reproduction at all.
Tooling lineage
- Sacred (Greff et al., 2017, The Sacred Infrastructure for Computational Research, SciPy): early formalisation — config injection, automatic capture of $\theta$, code state and randomness per run.
- Hydra (Yadan, 2019, Facebook Research): hierarchical config composition with command-line overrides; the de-facto standard for multi-run sweeps in deep learning; its
multirunwrites one resolved config per job — the append-only record at scale. - MLflow (Zaharia et al., 2018, IEEE Data Eng. Bull.) and Weights & Biases: run records as a queryable store;
runs.jsonlwith a UI and lineage. - OmegaConf-style resolved configs: the crucial semantics is resolution before recording — the stored config must be the final merged values, not the layered inputs.
Design tensions worth knowing
- Expressiveness vs. auditability: config languages accrete interpolation, inheritance and code hooks until configs need debugging. The discipline that survives audits: configs are data; derivations happen in code and the derived values are logged too.
- Granularity of the seed: one global seed hides the fact that data shuffling, initialisation and augmentation are separate random streams. Recording per-component seeds (a config sub-block) enables the isolation experiments of seed variance.
- Config drift vs. schema: long-lived projects accumulate dead keys and renamed knobs; a versioned schema (even a dataclass) with explicit migration keeps historical records interpretable — the same argument as database migrations.
The through-line: the config system is the experiment's provenance layer, and its quality bounds every claim built on top. A result whose $\theta$ cannot be produced on demand is, for scientific purposes, an anecdote.
What to learn next
- Making sense of two hundred runs — what config-keyed records make possible.
- Experiment tracking — the industrial version of runs.jsonl.
- When you cannot reproduce your own number — the failure this lesson quietly prevents.