Golden outputs and regression tests
A golden file freezes today's known-good output so tomorrow's "small refactor" cannot quietly change it — the same reason a tailor keeps your exact measurements on paper instead of guessing again each visit.
- 9 min read
- 3 reading levels
- Published
Read these first
On this page 7
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
A golden output is a saved, known-correct result that every future run gets compared against.
The analogy you have already lived
You have gone back to a tailor for a second visit. They pulled out a card with your exact measurements from last time, instead of guessing again. Your shoulders have not changed, so the card is trusted, checked against, and only updated on purpose when something really has changed.
A golden file is that measurement card for your code. It stores today's correct output. Every future run gets checked against it, not re-derived from scratch and hoped to be the same.
Why it exists
Code gets refactored. Someone tidies a function, renames a variable, swaps one regular expression for what looks like an equivalent one. It runs. Nothing crashes. It looks the same.
But "looks the same" and "produces the exact same output" are different claims, and only a saved, exact comparison can tell them apart. A golden test is that exact comparison, run automatically, on every change.
How it works
step 1 (done once, by a human)
input -> today's pipeline -> output -> SAVE as "golden"
step 2 (done automatically, forever after)
same input -> current pipeline -> new output
|
v
compare to the saved golden
|
match? ----no----> STOP. Something changed. Is that on purpose?
|
yes
v
all goodThe golden file is never regenerated by the test itself. Regenerating it automatically on failure would defeat the entire point — it would always agree with whatever the code currently does.
A real example you have seen
A spreadsheet's "recalculate" button never changes an old, saved answer unless you actually edit a cell. If it started silently drifting old totals on its own, you would lose trust in every number in the sheet immediately. Golden tests protect that same trust for a machine learning pipeline.
Remember this
- A golden file stores one known-good output, checked in like any other file, never regenerated automatically.
- The test compares today's output to the saved one, not to a freshly recomputed "correct" answer.
- A failing golden test means the output changed. That might be a bug, or it might be an intended improvement — a human decides which, and updates the golden file on purpose.
What to learn next
- Testing output you cannot predict — what to do when the output is not the same every time, so exact golden comparison stops working.
- CI/CD for machine learning — running this exact regression test on every push, automatically.
- Behavioural tests for models — checking specific abilities on purpose, the layer above a golden set.
Developer — Code and libraries.
Setup
pip install pytestThe pipeline step being frozen
A small, deterministic text-cleaning function — the kind that runs before almost every text model sees a row. Deterministic means the same input always gives the same output, which is exactly what makes golden testing possible here.
"""Deterministic text-cleaning used before every model in this pipeline sees a row."""
import re
_WHITESPACE = re.compile(r"\s+")
_PUNCT = re.compile(r"[^\w\s]")
def clean(text: str) -> str:
text = text.strip().lower()
text = _PUNCT.sub("", text)
text = _WHITESPACE.sub(" ", text)
return textFreezing the golden file
Run once, by a human, and checked into version control. This script is never part of the test suite itself.
"""Run once by a human to freeze today's known-good output. Never run inside a test."""
import json
from clean import clean
INPUTS = [
" Hello, World! ",
"Great product!!! Would buy AGAIN.",
"N0T-a-typical-sentence...",
"Multiple spaces here",
"Already clean",
]
golden = {text: clean(text) for text in INPUTS}
with open("golden_clean.json", "w") as f:
json.dump(golden, f, indent=2, sort_keys=True)
print(json.dumps(golden, indent=2))python make_golden.py{
" Hello, World! ": "hello world",
"Great product!!! Would buy AGAIN.": "great product would buy again",
"N0T-a-typical-sentence...": "n0tatypicalsentence",
"Multiple spaces here": "multiple spaces here",
"Already clean": "already clean"
}Notice the third line: hyphens are stripped as punctuation with no space left behind, so "N0T-a-typical-sentence..." becomes one run-on word, "n0tatypicalsentence". That may or may not be what you want — but the golden file does not judge, it only remembers. Whether this is correct is a decision a human makes once, when freezing the file.
The regression test
import json
import pytest
from clean import clean
with open("golden_clean.json") as f:
GOLDEN = json.load(f)
@pytest.mark.parametrize("raw_text,expected", sorted(GOLDEN.items()))
def test_matches_the_frozen_output(raw_text, expected):
assert clean(raw_text) == expectedpytest test_clean_regression.py -q..... [100%] 5 passed in 0.02s
Now break something on purpose
Someone tries to tidy the punctuation regex, meaning to keep only letters and spaces, and writes [^a-z\s] instead of [^\w\s]:
# in clean.py
_PUNCT = re.compile(r"[^a-z\s]")....F [100%]
================================== FAILURES ===================================
_ test_matches_the_frozen_output[N0T-a-typical-sentence...-n0tatypicalsentence] _
raw_text = 'N0T-a-typical-sentence...', expected = 'n0tatypicalsentence'
@pytest.mark.parametrize("raw_text,expected", sorted(GOLDEN.items()))
def test_matches_the_frozen_output(raw_text, expected):
> assert clean(raw_text) == expected
E AssertionError: assert 'ntatypicalsentence' == 'n0tatypicalsentence'
E
E - n0tatypicalsentence
E ? -
E + ntatypicalsentence
test_clean_regression.py:13: AssertionError
=========================== short test summary info ===========================
FAILED test_clean_regression.py::test_matches_the_frozen_output[N0T-a-typical-sentence...-n0tatypicalsentence]
1 failed, 4 passed in 0.28s[^a-z\s] — letters and spaces only — silently also strips digits, because \w (the original pattern) meant "letters, digits, and underscore" and a-z does not. Four of the five golden cases had no digits, so they still pass. Only the one case with a 0 in it catches the change. This is exactly why a golden set needs a handful of varied examples, not one — a single easy case would have missed this entirely.
Updating a golden file on purpose
When the new output is genuinely the intended one, re-run make_golden.py, review the diff in the saved JSON file like any other code change, and commit it with a reason written in the commit message. That review step — a human looking at exactly what changed and why — is the entire point. A test that regenerates its own golden file on failure has no review step, and stops meaning anything.
Exact match versus tolerance
String and integer outputs, as above, compare with plain equality. Floating-point outputs from a model rarely will, because the exact bits can shift by a tiny amount between library versions or hardware. Use a tolerance instead:
import math
assert math.isclose(current_score, golden_score, rel_tol=1e-4)Set the tolerance loose enough to absorb harmless floating-point noise, and tight enough that a real change in behaviour still fails the test. There is no universal number — start tight, and loosen it only when a passing change genuinely needs it.
Common mistakes
Storing the golden file somewhere it will not be reviewed. If it lives outside version control, nobody sees it change, and the whole safety net disappears silently.
One golden example. As shown above, the digit-stripping bug was invisible in four of five cases. Cover the boundary cases you actually care about — empty strings, punctuation, numbers, unusual characters — not only the easy ones.
Regenerating on every failure "to unblock CI". This is the single most common way golden tests stop protecting anything. If regenerating feels urgent, that pressure is exactly why the review step needs to survive it.
Testing golden outputs for something that is meant to change often. A golden test fits a deterministic, stable transform. A model's raw prediction on live, growing data is the wrong target — that belongs in monitoring, not a fixed golden file.
Try it yourself
Add a sixth input to INPUTS containing an emoji, regenerate the golden file, and look at what clean() actually does to it. Decide, on purpose, whether that is the behaviour you want — that decision, made once and written down, is the entire value of this pattern.
What to learn next
- Testing output you cannot predict — what to do when the output is not the same every time, so exact golden comparison stops working.
- CI/CD for machine learning — running this exact regression test on every push, automatically.
- Behavioural tests for models — checking specific abilities on purpose, the layer above a golden set.
Researcher — Mathematics and papers.
Snapshot testing as a general technique
Golden-file testing is a specific case of snapshot testing: capture the output of a function once, store it, and diff future outputs against the stored snapshot. It is common in UI testing (comparing rendered component output) and long predates its use in machine learning; the ML-specific contribution is applying it to numeric and probabilistic outputs, where the tolerance question above has no analogue in exact string or pixel comparison.
Where this sits on the ML Test Score
Breck et al. (2017) name a golden set explicitly under Model tests: a small, hand-curated, version-controlled set of examples with known-correct outputs, re-evaluated on every model change. Their observation, borne out here, is that a golden set catches a distinct class of bug from held-out accuracy — it fails on a specific, named case, not a shift in an aggregate number, which is what makes debugging it fast.
Golden sets versus property-based tests
A golden set checks specific, chosen inputs. A property-based test (see the researcher section of unit testing machine learning code) checks a general property across many generated inputs. They are complementary, not competing: property tests find unknown failure modes across a wide input space; golden sets pin down known cases you have specifically decided must never regress, including ones a random generator would rarely hit — the single digit inside a longer string above is exactly this kind of case.
Diffing structured outputs
For outputs richer than a scalar or a short string — a JSON object, a ranked list, an image — exact equality is often too strict (irrelevant key ordering, floating-point jitter) and too weak (misses a semantically important change hidden in a large diff). Structural diffing tools (deepdiff in Python) report the specific path and value that changed, which is what makes a large golden object practical to review rather than an unreadable wall of red and green lines.
Papers and prior art
- Breck et al., The ML Test Score, IEEE Big Data 2017 — the golden-set recommendation this lesson implements directly.
- Fowler, Mocks Aren't Stubs, 2007 — the broader test-double vocabulary that snapshot and golden testing sit alongside.
What to learn next
- Testing output you cannot predict — what to do when the output is not the same every time, so exact golden comparison stops working.
- CI/CD for machine learning — running this exact regression test on every push, automatically.
- Behavioural tests for models — checking specific abilities on purpose, the layer above a golden set.