Experiment tracking
Experiment tracking is writing down what you changed, what data you used and what score you got, automatically, so a good result can be repeated next month.
- 13 min read
- 3 reading levels
- Published
Read these first
On this page 9
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
Experiment tracking means recording what you changed, what data you used and what score you got — every single time, automatically.
The analogy you have already lived
You once cooked something that came out perfect. You had been improvising. A little extra of this, a shortcut on that, something you added at the end.
Then you tried to make it again next week and it was ordinary. You could not remember what you did differently. The good version is gone forever.
That is exactly what happens to a model. You changed six things across a long evening. Something worked. On Monday, you cannot find it again.
Why it exists
A machine learning project is not one experiment. It is a few hundred.
You change the depth of a tree. You add a column. You drop the rows with missing values instead of filling them. You try a different split of training and test data. Each change gives you a number, and you keep the best number.
Three things go wrong, and they go wrong to everyone:
You cannot repeat the win. The best score came from some combination you no longer remember.
You cannot compare fairly. Your Thursday score used slightly different data from your Tuesday score. The improvement was in the data, not the model — but the number does not say that.
You cannot explain it to anyone. Six months later somebody asks why you picked this model. You open the folder and find model_final, model_final_v2 and model_final_v2_new.
Every one of these has one cure: write it down while it happens, not afterwards.
How it works
A run is one attempt: one set of settings, one dataset, one score. A tracker stores every run and lets you sort them.
you change something
|
v
train and score ------> [ tracker ] stores:
| * the settings you used
v * a fingerprint of the data
keep going * the library versions
* the score
* where the model file wentThe fingerprint is the part people leave out and later regret. It is a short code computed from the data itself. If the data changes by one row, the code changes completely. So you can always tell a genuinely better model from a quietly edited dataset.
What a good record looks like
For every run you want four things, and you want them recorded by the code, not by you:
- Settings — every knob you turned, including the ones you left at their default.
- Data fingerprint — which exact data went in.
- Environment — which versions of which libraries.
- Result — the scores, and where the trained model file was saved.
If you record these four, any run can be rebuilt. If you miss one, some runs cannot.
Where you have already seen this idea
- A gym app logging every set and weight, so you can see this month against last month.
- A bank statement — every transaction, dated, in order, whether or not you wanted it.
- Version history in a document, letting you go back to the paragraph you deleted.
- A hospital chart at the foot of the bed, recorded by whoever was on duty.
Nobody enjoys filling these in. Everybody is glad they exist.
The honest part
Tracking feels like a waste of time in week one. You have six experiments and you remember all six.
By week four you have two hundred, and the memory is gone. The cost of adding tracking later is that everything before that day is unusable. That is the real reason to start on day one, and it is a discipline problem more than a tooling problem.
Remember this
- A run is one attempt: settings, data, environment, score.
- Record the data fingerprint, or you will compare two numbers that were never comparable.
- Start tracking on day one, because you cannot backfill what you never wrote down.
What to learn next
- MLflow — the tool that does all of this, with a UI on top.
- Model evaluation — which numbers are worth logging in the first place.
- Docker for ML — freezing the environment your logged runs refer to.
Developer — Code and libraries.
Build one before you install one
Tools like MLflow and Weights & Biases do this for you. They make far more sense once you have written the small version yourself. Then you know what each button is doing.
A tracker is a table. Here is a complete one, on SQLite, from the standard library.
Setup
pip install scikit-learn numpysqlite3, hashlib and json ship with Python. Nothing here downloads a dataset and nothing needs a GPU.
import hashlib
import json
import sqlite3
import numpy as np
import sklearn
from sklearn.model_selection import cross_val_score
from sklearn.tree import DecisionTreeClassifier
# 300 fake customers, 6 measurements each. RandomState(0) is the legacy
# generator, and NumPy guarantees its stream never changes — so your numbers
# below will match this page exactly.
rng = np.random.RandomState(0)
X = rng.normal(size=(300, 6))
y = (X[:, 0] + 0.8 * X[:, 1] - 0.5 * X[:, 2] + rng.normal(0, 0.7, 300) > 0).astype(int)
def fingerprint(arr):
"""A short hash of the data. Without it you cannot tell a better model
from a quietly edited dataset."""
return hashlib.sha256(np.ascontiguousarray(arr).tobytes()).hexdigest()[:12]
db = sqlite3.connect("runs.db")
db.execute("""CREATE TABLE IF NOT EXISTS runs (
id INTEGER PRIMARY KEY, params TEXT, data_hash TEXT,
sklearn TEXT, accuracy REAL)""")
def log_run(params, data_hash, accuracy):
db.execute("INSERT INTO runs (params, data_hash, sklearn, accuracy)"
" VALUES (?, ?, ?, ?)",
(json.dumps(params, sort_keys=True), data_hash,
sklearn.__version__, accuracy))
db.commit()
def evaluate(params):
model = DecisionTreeClassifier(random_state=0, **params)
return float(cross_val_score(model, X, y, cv=5).mean())
data_hash = fingerprint(X)
for depth in [1, 2, 3, 5, 10]:
params = {"max_depth": depth}
log_run(params, data_hash, evaluate(params))
print("data fingerprint:", data_hash, " scikit-learn:", sklearn.__version__)
print("\nleaderboard")
for params, h, acc in db.execute(
"SELECT params, data_hash, accuracy FROM runs ORDER BY accuracy DESC"):
print(f" {acc:.4f} {h} {params}")
# Six weeks later somebody re-runs the winner and gets a better number.
X[7, 0] = 4.2 # one outlier quietly removed from the data
log_run({"max_depth": 3}, fingerprint(X), evaluate({"max_depth": 3}))
print("\nsame settings, run again after the data was edited")
rows = list(db.execute("SELECT data_hash, accuracy FROM runs"
" WHERE params = ? ORDER BY id",
(json.dumps({"max_depth": 3}, sort_keys=True),)))
for h, acc in rows:
print(f" {acc:.4f} {h}")
print(" same settings, different fingerprint:", rows[0][0] != rows[1][0])data fingerprint: e9e44ea60f0a scikit-learn: 1.7.2
leaderboard
0.7767 e9e44ea60f0a {"max_depth": 5}
0.7700 e9e44ea60f0a {"max_depth": 10}
0.7533 e9e44ea60f0a {"max_depth": 3}
0.7367 e9e44ea60f0a {"max_depth": 2}
0.6567 e9e44ea60f0a {"max_depth": 1}
same settings, run again after the data was edited
0.7533 e9e44ea60f0a
0.7600 0b955f70a12c
same settings, different fingerprint: TrueDelete runs.db before re-running, or the leaderboard keeps growing. Your scikit-learn version will print differently from the one above. The accuracy figures should still match. If a future release changes them, the logged version is how you would find out.
The last three lines are the whole lesson
Somebody re-ran max_depth=3 and got 0.7600 instead of 0.7533. In a meeting that reads as a small improvement.
It was not an improvement. One value in the training data was changed. The settings were identical, the score moved, and the only evidence is that the fingerprint went from e9e44ea60f0a to 0b955f70a12c.
Without that column you would have shipped a conclusion that was never true. This is the most common way a team fools itself. It never feels like cheating while it is happening.
Line by line, the parts that matter
np.random.RandomState(0) rather than np.random.default_rng(0). The legacy generator carries a strict stream-compatibility promise from NumPy, so this file produces the same numbers on your machine as on mine. Use default_rng in real projects for its better statistical properties; use RandomState when a document has to reproduce exactly.
np.ascontiguousarray(arr).tobytes() before hashing. A NumPy array can be a view with a stride pattern, and .tobytes() on a non-contiguous array is not what you expect. Forcing contiguity first makes the fingerprint depend on values, not on memory layout.
json.dumps(params, sort_keys=True) — without sort_keys, {"a":1,"b":2} and {"b":2,"a":1} are different strings, and your "have I run this before?" lookup silently misses.
cross_val_score(..., cv=5) with StratifiedKFold's default of no shuffling. Deterministic on purpose. Turn shuffling on and you must log the shuffle seed as well, or the score is not repeatable.
Storing sklearn.__version__ in every row. It looks like clutter until the day a library upgrade shifts every score by half a point and you need to know which runs are on which side of the change.
What to log that this example does not
- The git commit of the code:
subprocess.run(["git", "rev-parse", "HEAD"], capture_output=True, text=True).stdout.strip(), plus a flag for whether the working tree was dirty. - The path of the saved model file, so a leaderboard row leads to an artefact.
- The full metric set, not one number. Accuracy alone hides which class you got worse at — see model evaluation.
- How long it ran and on what hardware. Cost is a result too.
Common mistakes
Logging only the winners. Failed runs are how you learn which directions are dead. They cost nothing to store.
Logging the metric but not the seed. Two runs differing by 0.4 points may differ only by random initialisation. Without the seed you cannot tell a real gain from noise.
Tuning against the test set. Every time you look at the test score and change something, you leak a little information into your choices. Keep a validation split for tuning and touch the test split once. See train, test and validation splits.
Filenames as a tracking system. model_v3_lr001_final.pkl encodes three facts and loses the other twenty. A row in a table costs the same and holds everything.
Try it yourself
Add a notes column and store one sentence about why you tried each setting. Six weeks from now that column will be the most valuable thing in the table. The numbers tell you what happened. The note tells you what you were thinking.
What to learn next
- MLflow — the tool that does all of this, with a UI on top.
- Model evaluation — which numbers are worth logging in the first place.
- Docker for ML — freezing the environment your logged runs refer to.
Researcher — Mathematics and papers.
Reproducibility has layers, and only some are in your control
Ordered from easy to nearly impossible:
- Same code, same data, same machine, same seed. Achievable, and the bar every logged run should clear.
- Same code, same data, different machine. Blocked by BLAS thread counts changing floating-point reduction order, by different CPU instruction sets, and by library minor versions.
- Same code, same data, GPU.
atomicAdd-based kernels are non-deterministic by construction; cuDNN algorithm selection is autotuned per device. PyTorch offerstorch.use_deterministic_algorithms(True)plusCUBLAS_WORKSPACE_CONFIG=:4096:8, at a real throughput cost, and some operations have no deterministic implementation at all. - Same result across a distributed rerun. Gradient all-reduce order varies with network timing. Bitwise reproduction generally requires deterministic reduction trees.
Log the seed, the library versions, the CUDA and cuDNN versions, the device name and the thread environment variables. That does not buy determinism. It buys the ability to explain a discrepancy, which is what you actually need.
Search strategy
Bergstra and Bengio (2012), Random Search for Hyper-Parameter Optimization, showed that random search dominates grid search in high dimensions. The mechanism is low effective dimensionality: only a few hyperparameters matter for a given problem, and a grid with $k$ points per axis spends $k^{n}$ trials while sampling only $k$ distinct values of the axis that mattered. Random search with $T$ trials samples $T$ distinct values on every axis.
Beyond random search:
- Bayesian optimisation — model $p(\text{score} \mid \text{config})$ with a Gaussian process or a TPE (Bergstra et al., 2011) and maximise expected improvement. Wins when a trial is expensive relative to the modelling overhead.
- Successive halving and Hyperband (Li et al., 2017) — allocate a small budget to many configurations, discard the worst fraction, repeat. Treats hyperparameter search as a non-stochastic infinite-armed bandit problem. ASHA (Li et al., 2018) is the asynchronous variant, and is what most modern schedulers actually run.
- Population Based Training (Jaderberg et al., 2017) — mutate and copy weights between concurrently training runs, producing a schedule rather than a fixed configuration.
The multiple-comparisons trap in model selection
Selecting the best of $m$ configurations on a validation set of size $n$ gives an optimistically biased estimate of generalisation. Informally, the expected maximum of $m$ noisy estimates each with standard error $\sigma$ exceeds the true best by roughly $\sigma\sqrt{2\ln m}$.
Concretely: with $n = 2000$ validation examples and accuracy near $0.85$, $\sigma \approx \sqrt{0.85 \cdot 0.15 / 2000} \approx 0.008$. Selecting the best of 100 configurations inflates the apparent score by about $0.008 \cdot \sqrt{2 \ln 100} \approx 0.024$ — over two accuracy points of pure selection bias, before any real improvement.
Dwork et al. (2015), The reusable holdout: Preserving validity in adaptive data analysis, gives a differential-privacy-based mechanism for answering many adaptive queries against one holdout with provable generalisation guarantees. Rarely deployed; the underlying warning should still change how you read a leaderboard.
Reporting budget, not only the best number
Dodge et al. (2019), Show Your Work: Improved Reporting of Experimental Results, argues that a single best number is unreportable without the search budget that produced it. They propose reporting expected validation performance as a function of the number of trials, estimated from the observed distribution of trial scores. Two methods reported at "best of 5" and "best of 500" are not comparable, and the difference is frequently larger than the reported gain.
What a run record should contain
A defensible schema, drawn from what actually gets asked six months later:
| Group | Fields |
|---|---|
| Identity | run id, parent run id, git commit, dirty-tree flag, user, start and end time |
| Inputs | dataset URI, content hash, row count, split definition, preprocessing version |
| Configuration | full resolved config including defaults, all seeds, framework versions |
| Environment | OS, Python, CUDA and cuDNN, device model, thread and BLAS environment |
| Results | every metric with its split, per-epoch curves, confusion matrix, calibration |
| Artefacts | model URI and hash, plots, logs, and the resource cost of producing them |
The "full resolved config including defaults" row is the one people skip. A default that changes in a library upgrade is invisible in a config file that only records overrides.
Papers
- Bergstra and Bengio, Random Search for Hyper-Parameter Optimization, JMLR 2012 — jmlr.org/papers/v13/bergstra12a.html
- Li et al., Hyperband: A Novel Bandit-Based Approach to Hyperparameter Optimization, JMLR 2018 — arxiv.org/abs/1603.06560
- Li et al., A System for Massively Parallel Hyperparameter Tuning (ASHA), MLSys 2020 — arxiv.org/abs/1810.05934
- Jaderberg et al., Population Based Training of Neural Networks, 2017 — arxiv.org/abs/1711.09846
- Dwork et al., The reusable holdout, Science 2015 — science.org/doi/10.1126/science.aaa9375
- Dodge et al., Show Your Work, EMNLP 2019 — arxiv.org/abs/1909.03004
- Pineau et al., Improving Reproducibility in Machine Learning Research, JMLR 2021 — the NeurIPS reproducibility programme, and what it measured.
What to learn next
- MLflow — the tool that does all of this, with a UI on top.
- Model evaluation — which numbers are worth logging in the first place.
- Docker for ML — freezing the environment your logged runs refer to.