Registries, Artifacts and Environments

Model lineage and traceability

Model lineage records which code, data and parent model produced a given model, so you can trace any prediction back to its source.

On this page 8
  1. The short answer
  2. The analogy you have already lived
  3. Why it exists
  4. How it works
  5. A real example you have seen
  6. The honest part
  7. Remember this
  8. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

Model lineage is the recorded trail of exactly which code, which data, and which earlier model produced a given model.

The analogy you have already lived

You have looked at a family tree. It does not only list names. It shows who came from whom, generation by generation. You can trace any person back to their parents, and theirs, all the way to the root.

A model rarely appears from nothing either. It was trained by some code, on some data, and often started from an earlier version rather than from scratch. Lineage is that family tree, for models instead of people.

Why it exists

Imagine a model in production starts giving strange answers. Someone needs to find out why, fast.

Without lineage: which code produced it? Nobody remembers exactly. Which data? Maybe last month's, maybe this month's — nobody is sure. Was it trained from scratch, or fine-tuned from an older model that had its own problems? Guesswork, under pressure, is how a ten-minute investigation becomes a two-day one.

With lineage, the model's record names the exact code version, the exact dataset, and its exact parent, if it has one. The investigation starts with facts instead of guesses.

How it works

   run 1: code=a1b2c3d   data=data-v1-h9f2   parent=none    accuracy=0.83
              |
              v  (fine-tuned on more data)
   run 2: code=e4f5a6b   data=data-v2-3c8a   parent=run 1   accuracy=0.88
              |
              v  (a preprocessing bug fixed, retrained)
   run 3: code=9d8c7b6   data=data-v2-3c8a   parent=run 2   accuracy=0.91

   the model in production today is run 3 -- trace it back, and the
   full chain of decisions that produced it is right there

A real example you have seen

A courier tracking number does exactly this for a package. Every stage it passed through, in order, each one timestamped and attributable. When a package goes missing, that trail is what makes finding where it went possible, instead of an unanswerable mystery.

The honest part

Lineage only helps if it was recorded at the time of training, not reconstructed afterward from memory. A team that skips logging "just this once, because it is urgent" is most likely to need that missing record later, during the next incident.

Remember this

  • Lineage records which code, which data, and which parent model produced a given model. It is the full trail, not only the final result.
  • It turns "why did this model do that?" from a guess into a lookup.
  • It is only useful if recorded at training time — reconstructing it afterward, from memory, rarely works.

What to learn next

Developer — Code and libraries.

Setup

No installs needed beyond the standard library.

A lineage log, and walking a chain back to its root

model_lineage.py
import sqlite3

conn = sqlite3.connect(":memory:")
conn.execute("""
    CREATE TABLE runs (
        id INTEGER PRIMARY KEY,
        model_name TEXT,
        code_commit TEXT,
        dataset_hash TEXT,
        parent_id INTEGER,
        accuracy REAL
    )
""")


def log_run(model_name, code_commit, dataset_hash, accuracy, parent_id=None):
    cur = conn.execute(
        "INSERT INTO runs (model_name, code_commit, dataset_hash, parent_id, accuracy)"
        " VALUES (?, ?, ?, ?, ?)",
        (model_name, code_commit, dataset_hash, parent_id, accuracy),
    )
    conn.commit()
    return cur.lastrowid


def trace_lineage(run_id):
    """Walks parent_id pointers back to the run that started from nothing."""
    chain = []
    current = run_id
    while current is not None:
        row = conn.execute(
            "SELECT id, model_name, code_commit, dataset_hash, parent_id, accuracy"
            " FROM runs WHERE id = ?", (current,),
        ).fetchone()
        if row is None:
            break
        chain.append(row)
        current = row[4]
    return list(reversed(chain))  # oldest first


# a base model, trained from scratch
base = log_run("loan-scorer", "a1b2c3d", "data-v1-h9f2", accuracy=0.83)
# a fine-tune of the base, on more data
tuned = log_run("loan-scorer", "e4f5a6b", "data-v2-3c8a", accuracy=0.88, parent_id=base)
# a fine-tune of the fine-tune, after a bug fix in preprocessing
fixed = log_run("loan-scorer", "9d8c7b6", "data-v2-3c8a", accuracy=0.91, parent_id=tuned)

print(f"the model in production today is run {fixed}. Its full lineage:\n")
for run_id, name, commit, data_hash, parent_id, accuracy in trace_lineage(fixed):
    print(f"  run {run_id}: code={commit}  data={data_hash}"
          f"  accuracy={accuracy:.2f}  parent={parent_id}")
Output
the model in production today is run 3. Its full lineage:

  run 1: code=a1b2c3d  data=data-v1-h9f2  accuracy=0.83  parent=None
  run 2: code=e4f5a6b  data=data-v2-3c8a  accuracy=0.88  parent=1
  run 3: code=9d8c7b6  data=data-v2-3c8a  accuracy=0.91  parent=2

Purely deterministic SQL — this reproduces exactly, every run, on any machine.

Walking through it

parent_id is the entire mechanism. Every other feature — the trace, the family tree, the full history — falls out of one self-referencing foreign key. Nothing more exotic is needed to represent a lineage graph, as long as it stays a simple chain rather than branching.

trace_lineage walks backward, then reverses. Following parent_id naturally produces the chain from newest to oldest. Reversing it at the end presents the story in the order it actually happened — far more natural for a human investigating an incident.

dataset_hash, not a dataset name. A name like "training_data.csv" can point at completely different content over time if the file gets overwritten. A hash — like the one built in packaging a model artifact — pins down the exact content, not only a label that could drift.

Common mistakes

Recording lineage as a note, not a queryable field. A free-text comment saying "based on last month's retrain" cannot be searched, joined, or traced automatically at 2 a.m. during an incident. Store it as structured data, the way parent_id is here, not as prose.

Losing lineage across tool boundaries. A model trained with one tool, exported, then fine-tuned with a different tool entirely, easily loses its parent_id in the handoff unless someone deliberately carries it across. Lineage breaks at every seam where responsibility changes hands, unless someone owns stitching it back together.

Confusing "retrained on new data" with "an entirely new lineage." Both are meaningfully different histories. A model fine-tuned from a known-good parent inherits that parent's properties, good and bad; a model trained fresh does not. Losing that distinction loses real information about what a model is likely to have inherited.

Try it yourself

Add a notes column to the runs table, and log a short reason for each run — "first version", "added Q2 data", "fixed the income-rounding bug". Print it alongside each row in trace_lineage. The trace then reads as a short, real story of why the model changed at each step, not only what changed.

What to learn next

Researcher — Mathematics and papers.

Lineage as a directed acyclic graph

The chain shown above is the simplest case: one parent per model. Real pipelines often branch and merge — one model fine-tuned into two separate specialised versions, or one model built by combining several upstream models (an ensemble, or a distilled student model with several teacher models). The general structure is a directed acyclic graph (DAG), not a linear chain: multiple parents per node are legitimate, and trace_lineage above would need to become a graph traversal rather than a simple pointer walk to handle that correctly.

Data lineage, feature lineage, and model lineage as separate layers

Full traceability in a production ML system spans three connected but distinct lineage graphs. Data lineage asks which raw data fed which feature computation — the concern of why data quality decides everything. Model lineage, this lesson, asks which features and code version produced which model. Prediction lineage asks which model version served which specific prediction, by logging every prediction with the model version that produced it. An incident investigation frequently needs to walk across all three, not only one.

Standards and tooling

  • MLMD (ML Metadata), the metadata layer underneath Google's TFX pipelines, formalises artifacts, executions and their relationships as a typed graph specifically to support this kind of cross-pipeline lineage query at scale.
  • W3C PROV, a general-purpose provenance data model originally designed for scientific and web data, has been adapted by several ML lineage tools as a common interchange format, so lineage recorded by one tool can, in principle, be understood by another.

Why this matters beyond debugging

Beyond incident response, lineage is what makes a regulatory or audit question answerable at all. Take "was any model currently in production trained on data from a user who has since asked to be forgotten?" That question is unanswerable without a queryable trail connecting models back to the datasets that built them. In regulated industries, this is a real, increasingly common requirement, not a hypothetical one.

What to learn next