Reproducibility and Running Experiments
Designing an ablation that proves your claim
An ablation study removes your system's components one at a time to show each one earns its place — the difference between claiming your additions help and proving it.
- 8 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
To prove an ingredient matters, cook the dish without it — an ablation study does exactly that for every piece of your system.
Your dal tastes wonderful, and you credit the extra garlic. Your friend suspects the new pressure cooker. There is one honest way to settle it: cook the dal again without the garlic, everything else identical. If it still tastes wonderful, the garlic was a passenger, not a driver.
An ablation study — from a word meaning "removal" — is this cooking experiment for machine learning. You built a system with three additions and it scores well. Which additions earned their place? Remove each one alone, retrain, remeasure. The score drop after each removal is that piece's contribution.
Why it exists
Systems grow by accumulation. You add a feature, a trick, a preprocessing step — the score goes up, the addition stays. After six months you have twelve additions and a good score, and no idea which additions matter. Complexity you cannot justify is complexity you must maintain forever.
Worse: when you tell others "our clever feature X improved results", you are making a causal claim. Between your before and after, other things also changed — more data, a tuned threshold, a library update. Without ablation, "X helped" is a story. With it, it is a measurement.
How it works
full system (all pieces) : ████████████ 86
without piece A : ████████ 62 ← A carries a lot
without piece B : ███████████ 84 ← B helps a little
without piece C : ████████████ 86 ← C does nothingRules of the game: remove one piece at a time, keep everything else frozen, retrain properly each time, and use the same measurement throughout. The result is a table where every row answers one clean question.
A real example you have seen
Doctors finding a food allergy use an elimination diet: remove one food, watch, restore, remove the next. Removing five foods at once tells you nothing about which one was guilty. The one-at-a-time discipline is the entire method — in medicine and in model-building alike.
Remember this
- An ablation removes one component, retrains, and measures the drop.
- One change per experiment; everything else frozen.
- The output is a table that justifies each piece of your system — or retires it.
What to learn next
- Is this improvement real? — the statistics for gaps your table leaves in doubt.
- Hunting a leak with ablations — the same instrument pointed at suspicion instead of pride.
- Reading a results table sceptically — applying today's standards to other people's tables.
Developer — Code and libraries.
Setup
pip install scikit-learnVerified with scikit-learn 1.7.2, numpy 1.26.4, CPU. Seeded, so your numbers should match. Runs in seconds.
Ablating a two-trick system
The claim to prove: "our interaction feature and our class weighting both help on this imbalanced problem." The system has exactly two tricks, so the ablation table has four rows — full, minus each trick, minus both.
import numpy as np
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import cross_val_score
rng = np.random.default_rng(1)
n = 2000
x1, x2 = rng.normal(size=n), rng.normal(size=n)
y = ((x1 > 0) ^ (x2 > 0)).astype(int) # XOR pattern: not linear
y[rng.random(n) < 0.05] ^= 1 # 5% label noise
keep = (y == 0) | (rng.random(n) < 0.25) # positives become rare
x1, x2, y = x1[keep], x2[keep], y[keep]
print("positive share:", round(y.mean(), 3), "of", len(y), "rows")
def f1(features, **kw):
model = LogisticRegression(max_iter=2000, **kw)
return cross_val_score(model, features, y, cv=5, scoring="f1").mean()
raw = np.column_stack([x1, x2])
eng = np.column_stack([x1, x2, x1 * x2])
table = {
"full recipe ": f1(eng, class_weight="balanced"),
"- interaction feature ": f1(raw, class_weight="balanced"),
"- class weights ": f1(eng),
"- both ": f1(raw),
}
for name, score in table.items():
print(f"{name} F1 {score:.3f}")positive share: 0.208 of 1224 rows full recipe F1 0.862 - interaction feature F1 0.318 - class weights F1 0.615 - both F1 0.000
The walkthrough
Each row is one clean sentence. Removing the interaction feature costs 0.544 of F1 — it is the load-bearing wall (the XOR pattern is invisible to a linear model without it). Removing class weights costs 0.247 — real, substantial, smaller. The claim "both help" is now proven, with sizes attached.
The - both row earns its place. F1 of exactly 0.000 — without either trick, the model never predicts the rare class at all. This row also exposes an interaction between the tricks: the summed individual drops (0.544 + 0.247) do not equal the drop from removing both (0.862). Components can overlap or reinforce; the four-row design catches what two three-row designs would miss.
F1 is the right metric for this claim — with 21% positives, plain accuracy would hide the collapse in the last row behind a respectable-looking 0.79. Choose the metric before running the table, and keep it fixed for every row (model evaluation covers the choices).
Retraining per row is non-negotiable. Zeroing a feature at prediction time ablates a trained model's input, a different (and usually misleading) experiment. Each row above refits from scratch, via cross_val_score.
Designing a table that convinces
- Ablate claims, not code lines. Each row should correspond to a sentence in your announcement. If no sentence mentions it, it goes in one grouped "minor tricks" row.
- Add seeds when gaps are small. Each cell above is a 5-fold mean; for gaps under a few points, run several seeds per cell and show spreads — ranges, not points.
- Include the do-nothing row. A majority-class or constant-prediction baseline anchors the whole table to reality.
- Order rows by drop size in your final write-up. Readers should meet the load-bearing wall first.
Common mistakes
Removing two things at once, attributing to one. The classic. One row, one removal — or the row proves nothing.
Ablating without retuning. Sometimes removal changes what tuning would choose (a model without dropout wants a different learning rate). For rigorous claims, re-tune per row within the same budget; say so either way.
Testing every row against the test set. A twelve-row table is twelve peeks. Build tables on validation folds; confirm the final system once on test. The arithmetic of peeking lives in test-set overfitting.
Publishing only the flattering rows. An ablation that omits the row where removal helped is advertising. The embarrassing row is the most informative one you have.
Try it yourself
Add a third "trick": a useless feature x1 + x2 appended to eng. Extend the table with - useless feature. Confirm its removal costs nothing — then enjoy deleting it from the system with a table to justify the deletion.
What to learn next
- Is this improvement real? — the statistics for gaps your table leaves in doubt.
- Hunting a leak with ablations — the same instrument pointed at suspicion instead of pride.
- Reading a results table sceptically — applying today's standards to other people's tables.
Researcher — Mathematics and papers.
Ablations as factorial experiments
Leave-one-out ablation is a fractional factorial design: with $m$ binary components, the full design has $2^m$ cells; LOO measures the $m$ main-effect contrasts at the corner where all other components are on. The demo's four-row table is the full $2^2$ factorial, which is why it can expose the interaction term:
$$ \text{interaction} = S_{11} - S_{10} - S_{01} + S_{00} $$
Where:
- $S_{11}$ — score with both components (0.862).
- $S_{10}, S_{01}$ — one component each (0.318, 0.615).
- $S_{00}$ — neither (0.000).
Here $0.862 - 0.318 - 0.615 + 0.000 = -0.071$: mild negative interaction — the components partially substitute for each other. LOO alone cannot see this; as $m$ grows, targeted two-component cells around the suspected interactions are the affordable middle ground between LOO ($m+1$ runs) and full factorial ($2^m$).
Attribution semantics
The LOO drop measures a marginal contribution at a specific coalition — the same object as in leak-hunting ablations, pointed at a different hypothesis. Order-free attribution over all coalitions is the Shapley value; with correlated/substituting components, LOO systematically understates each (redundancy splitting the credit). Report the table, not a single "contribution percentage".
Statistical treatment: each cell is an estimate with variance; adjacent rows share data, so paired designs (same folds and seeds per cell) sharpen contrasts, and cell differences want the machinery of the next lesson — with multiple-comparison discipline once tables grow (Bonferroni over the $m$ claimed effects is the blunt, honest default).
The literature on ablation culture
- Lipton and Steinhardt (2018), Troubling Trends in Machine Learning Scholarship: names the failure this lesson prevents — papers where gains attributed to the headline idea actually came from tuning and "technical debt of wins"; empirical rigour sections (ablations) as the remedy.
- Melis, Dyer and Blunsom (2018), On the State of the Art of Evaluation in Neural Language Models, ICLR: carefully tuned LSTM baselines matched newer architectures — the field-scale cost of missing ablations against tuning.
- Musgrave, Belongie and Lim (2020), A Metric Learning Reality Check, ECCV: a decade of claimed gains largely dissolved under equalised training and evaluation — accumulated unablated confounds.
- Meyes et al. (2019), Ablation Studies in Artificial Neural Networks, arXiv: the term's neuroscience lineage (lesion studies) and network-internal ablation (units, layers) — a different granularity of the same logic.
Ablations under a compute budget
For expensive systems, per-cell retraining is the binding cost. Budget-honest options: reduced-scale proxies (ablate at 10% scale, confirm the top rows at full scale — with the caveat that effects can be scale-dependent), early-stopped cells under equal steps (Dodge et al., 2019, Show Your Work: report budget alongside), and ablating within one seed while confirming the largest effect across seeds. State which shortcut you took; an ablation whose protocol is hidden inherits the credibility problem it was meant to solve — the same transparency standard as reading a results table demands of others.
What to learn next
- Is this improvement real? — the statistics for gaps your table leaves in doubt.
- Hunting a leak with ablations — the same instrument pointed at suspicion instead of pride.
- Reading a results table sceptically — applying today's standards to other people's tables.