Registries, Artifacts and Environments
Model lineage and traceability
Model lineage records which code, data and parent model produced a given model, so you can trace any prediction back to its source.
- 8 min read
- 3 reading levels
- Published
Read these first
On this page 8
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
Model lineage is the recorded trail of exactly which code, which data, and which earlier model produced a given model.
The analogy you have already lived
You have looked at a family tree. It does not only list names. It shows who came from whom, generation by generation. You can trace any person back to their parents, and theirs, all the way to the root.
A model rarely appears from nothing either. It was trained by some code, on some data, and often started from an earlier version rather than from scratch. Lineage is that family tree, for models instead of people.
Why it exists
Imagine a model in production starts giving strange answers. Someone needs to find out why, fast.
Without lineage: which code produced it? Nobody remembers exactly. Which data? Maybe last month's, maybe this month's — nobody is sure. Was it trained from scratch, or fine-tuned from an older model that had its own problems? Guesswork, under pressure, is how a ten-minute investigation becomes a two-day one.
With lineage, the model's record names the exact code version, the exact dataset, and its exact parent, if it has one. The investigation starts with facts instead of guesses.
How it works
run 1: code=a1b2c3d data=data-v1-h9f2 parent=none accuracy=0.83
|
v (fine-tuned on more data)
run 2: code=e4f5a6b data=data-v2-3c8a parent=run 1 accuracy=0.88
|
v (a preprocessing bug fixed, retrained)
run 3: code=9d8c7b6 data=data-v2-3c8a parent=run 2 accuracy=0.91
the model in production today is run 3 -- trace it back, and the
full chain of decisions that produced it is right thereA real example you have seen
A courier tracking number does exactly this for a package. Every stage it passed through, in order, each one timestamped and attributable. When a package goes missing, that trail is what makes finding where it went possible, instead of an unanswerable mystery.
The honest part
Lineage only helps if it was recorded at the time of training, not reconstructed afterward from memory. A team that skips logging "just this once, because it is urgent" is most likely to need that missing record later, during the next incident.
Remember this
- Lineage records which code, which data, and which parent model produced a given model. It is the full trail, not only the final result.
- It turns "why did this model do that?" from a guess into a lookup.
- It is only useful if recorded at training time — reconstructing it afterward, from memory, rarely works.
What to learn next
- Pinning ML dependencies — recording which library versions alongside the code and data this lesson already tracks.
- What goes inside a model artifact — the dataset hash used above, and where it actually comes from.
- Experiment tracking — the tool that usually generates the raw run records a lineage system like this one is built from.
Developer — Code and libraries.
Setup
No installs needed beyond the standard library.
A lineage log, and walking a chain back to its root
import sqlite3
conn = sqlite3.connect(":memory:")
conn.execute("""
CREATE TABLE runs (
id INTEGER PRIMARY KEY,
model_name TEXT,
code_commit TEXT,
dataset_hash TEXT,
parent_id INTEGER,
accuracy REAL
)
""")
def log_run(model_name, code_commit, dataset_hash, accuracy, parent_id=None):
cur = conn.execute(
"INSERT INTO runs (model_name, code_commit, dataset_hash, parent_id, accuracy)"
" VALUES (?, ?, ?, ?, ?)",
(model_name, code_commit, dataset_hash, parent_id, accuracy),
)
conn.commit()
return cur.lastrowid
def trace_lineage(run_id):
"""Walks parent_id pointers back to the run that started from nothing."""
chain = []
current = run_id
while current is not None:
row = conn.execute(
"SELECT id, model_name, code_commit, dataset_hash, parent_id, accuracy"
" FROM runs WHERE id = ?", (current,),
).fetchone()
if row is None:
break
chain.append(row)
current = row[4]
return list(reversed(chain)) # oldest first
# a base model, trained from scratch
base = log_run("loan-scorer", "a1b2c3d", "data-v1-h9f2", accuracy=0.83)
# a fine-tune of the base, on more data
tuned = log_run("loan-scorer", "e4f5a6b", "data-v2-3c8a", accuracy=0.88, parent_id=base)
# a fine-tune of the fine-tune, after a bug fix in preprocessing
fixed = log_run("loan-scorer", "9d8c7b6", "data-v2-3c8a", accuracy=0.91, parent_id=tuned)
print(f"the model in production today is run {fixed}. Its full lineage:\n")
for run_id, name, commit, data_hash, parent_id, accuracy in trace_lineage(fixed):
print(f" run {run_id}: code={commit} data={data_hash}"
f" accuracy={accuracy:.2f} parent={parent_id}")the model in production today is run 3. Its full lineage: run 1: code=a1b2c3d data=data-v1-h9f2 accuracy=0.83 parent=None run 2: code=e4f5a6b data=data-v2-3c8a accuracy=0.88 parent=1 run 3: code=9d8c7b6 data=data-v2-3c8a accuracy=0.91 parent=2
Purely deterministic SQL — this reproduces exactly, every run, on any machine.
Walking through it
parent_id is the entire mechanism. Every other feature — the trace, the family tree, the full history — falls out of one self-referencing foreign key. Nothing more exotic is needed to represent a lineage graph, as long as it stays a simple chain rather than branching.
trace_lineage walks backward, then reverses. Following parent_id naturally produces the chain from newest to oldest. Reversing it at the end presents the story in the order it actually happened — far more natural for a human investigating an incident.
dataset_hash, not a dataset name. A name like "training_data.csv" can point at completely different content over time if the file gets overwritten. A hash — like the one built in packaging a model artifact — pins down the exact content, not only a label that could drift.
Common mistakes
Recording lineage as a note, not a queryable field. A free-text comment saying "based on last month's retrain" cannot be searched, joined, or traced automatically at 2 a.m. during an incident. Store it as structured data, the way parent_id is here, not as prose.
Losing lineage across tool boundaries. A model trained with one tool, exported, then fine-tuned with a different tool entirely, easily loses its parent_id in the handoff unless someone deliberately carries it across. Lineage breaks at every seam where responsibility changes hands, unless someone owns stitching it back together.
Confusing "retrained on new data" with "an entirely new lineage." Both are meaningfully different histories. A model fine-tuned from a known-good parent inherits that parent's properties, good and bad; a model trained fresh does not. Losing that distinction loses real information about what a model is likely to have inherited.
Try it yourself
Add a notes column to the runs table, and log a short reason for each run — "first version", "added Q2 data", "fixed the income-rounding bug". Print it alongside each row in trace_lineage. The trace then reads as a short, real story of why the model changed at each step, not only what changed.
What to learn next
- Pinning ML dependencies — recording which library versions alongside the code and data this lesson already tracks.
- What goes inside a model artifact — the dataset hash used above, and where it actually comes from.
- Experiment tracking — the tool that usually generates the raw run records a lineage system like this one is built from.
Researcher — Mathematics and papers.
Lineage as a directed acyclic graph
The chain shown above is the simplest case: one parent per model. Real pipelines often branch and merge — one model fine-tuned into two separate specialised versions, or one model built by combining several upstream models (an ensemble, or a distilled student model with several teacher models). The general structure is a directed acyclic graph (DAG), not a linear chain: multiple parents per node are legitimate, and trace_lineage above would need to become a graph traversal rather than a simple pointer walk to handle that correctly.
Data lineage, feature lineage, and model lineage as separate layers
Full traceability in a production ML system spans three connected but distinct lineage graphs. Data lineage asks which raw data fed which feature computation — the concern of why data quality decides everything. Model lineage, this lesson, asks which features and code version produced which model. Prediction lineage asks which model version served which specific prediction, by logging every prediction with the model version that produced it. An incident investigation frequently needs to walk across all three, not only one.
Standards and tooling
- MLMD (ML Metadata), the metadata layer underneath Google's TFX pipelines, formalises artifacts, executions and their relationships as a typed graph specifically to support this kind of cross-pipeline lineage query at scale.
- W3C PROV, a general-purpose provenance data model originally designed for scientific and web data, has been adapted by several ML lineage tools as a common interchange format, so lineage recorded by one tool can, in principle, be understood by another.
Why this matters beyond debugging
Beyond incident response, lineage is what makes a regulatory or audit question answerable at all. Take "was any model currently in production trained on data from a user who has since asked to be forgotten?" That question is unanswerable without a queryable trail connecting models back to the datasets that built them. In regulated industries, this is a real, increasingly common requirement, not a hypothetical one.
What to learn next
- Pinning ML dependencies — recording which library versions alongside the code and data this lesson already tracks.
- What goes inside a model artifact — the dataset hash used above, and where it actually comes from.
- Experiment tracking — the tool that usually generates the raw run records a lineage system like this one is built from.