Registries, Artifacts and Environments

Model registries

A model registry is the one place that lists every trained version of a model, and which one is actually live.

Read these first

On this page 9
  1. The short answer
  2. The analogy you have already lived
  3. Why it exists
  4. What a registry actually tracks
  5. How it works
  6. A real example you have seen
  7. The honest part
  8. Remember this
  9. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

A model registry is a catalog of every version of every model you have trained. It records which one is currently in use.

The analogy you have already lived

You have used a library. Every book has a catalog card — title, author, edition, shelf number. The catalog does not hold the books themselves. It tells you exactly where each one is, and which edition to borrow.

Without that catalog, a library is only a room full of books nobody can reliably find. A registry is that catalog, for trained models instead of books.

Why it exists

A single model you train once, on your laptop, is easy to keep track of. There is only one of it, and you remember what it is.

That stops being true fast. A real project accumulates versions. One trained last month, one retrained after new data arrived, one trained by a teammate, one currently answering real users. "Which model is actually live right now?" becomes a question nobody can answer confidently, without a shared record.

What a registry actually tracks

For each version of a model, a registry typically records where its file lives, how well it scored, and who or what produced it. It also records its stage — a label like staging, production, or archived, describing what that version is currently allowed to do.

How it works

   train v1  ->  register  ->  stage: staging   (accuracy 0.83)
   train v2  ->  register  ->  stage: staging   (accuracy 0.88)
                                     |
                            promote v2 to production
                                     v
                              stage: production
   train v3  ->  register  ->  stage: staging   (accuracy 0.86, worse than v2)
                                     |
                          v3 stays in staging, or gets archived

The registry does not decide which model is best — a person, or an automated check, decides that. The registry's job is to make that decision recorded, visible, and reversible.

A real example you have seen

Any app store version history works the same way. Every past version of an app is still listed, and one is marked as current. Rolling back to an older version is a real, supported action, not a scramble to find an old file on someone's laptop.

The honest part

A registry is only as trustworthy as the discipline behind it. A model deployed by manually copying a file, skipping the registry entirely, breaks the one property a registry exists to guarantee. The record no longer matches reality. The tool cannot enforce that on its own — the team's habits do.

Remember this

  • A registry answers "which version, and where is it?" It covers every model you have ever trained, not only the current one.
  • Stage labels (staging, production, archived) record what each version is currently allowed to do.
  • A registry is only trustworthy if every deployment actually goes through it — a manual shortcut breaks the record silently.

What to learn next

Developer — Code and libraries.

Setup

No installs needed beyond the standard library. This builds a minimal registry from scratch, using SQLite, to show what a tool like MLflow's registry is doing underneath.

A working registry in under 40 lines

registry.py
import json
import sqlite3

conn = sqlite3.connect(":memory:")
conn.execute("""
    CREATE TABLE models (
        name TEXT, version INTEGER, path TEXT,
        accuracy REAL, stage TEXT, metadata TEXT,
        PRIMARY KEY (name, version)
    )
""")


def register(name, path, accuracy, metadata):
    cur = conn.execute("SELECT COALESCE(MAX(version), 0) FROM models WHERE name = ?", (name,))
    next_version = cur.fetchone()[0] + 1
    conn.execute(
        "INSERT INTO models VALUES (?, ?, ?, ?, 'staging', ?)",
        (name, next_version, path, accuracy, json.dumps(metadata)),
    )
    conn.commit()
    return next_version


def promote(name, version, stage):
    conn.execute("UPDATE models SET stage = ? WHERE name = ? AND version = ?", (stage, name, version))
    conn.commit()


def get_by_stage(name, stage):
    return conn.execute(
        "SELECT version, path, accuracy FROM models WHERE name = ? AND stage = ?",
        (name, stage),
    ).fetchone()


# three training runs of the same model, over time
v1 = register("loan-scorer", "runs/v1/model.joblib", 0.83, {"trained_by": "pranay", "rows": 400})
v2 = register("loan-scorer", "runs/v2/model.joblib", 0.88, {"trained_by": "pranay", "rows": 900})
v3 = register("loan-scorer", "runs/v3/model.joblib", 0.86, {"trained_by": "pranay", "rows": 1500})
print(f"registered versions: {v1}, {v2}, {v3}")

promote("loan-scorer", v2, "production")
print("promoted version", v2, "to production")
print("current production model:", get_by_stage("loan-scorer", "production"))

promote("loan-scorer", v3, "archived")  # v3 scored lower than v2
print("version 3 scored lower, sent to archived instead of production")

for v, acc, stage in conn.execute(
    "SELECT version, accuracy, stage FROM models WHERE name = ? ORDER BY version", ("loan-scorer",)
):
    print(f"  v{v}: accuracy={acc:.2f}  stage={stage}")
Output
registered versions: 1, 2, 3
promoted version 2 to production
current production model: (2, 'runs/v2/model.joblib', 0.88)
version 3 scored lower, sent to archived instead of production
  v1: accuracy=0.83  stage=staging
  v2: accuracy=0.88  stage=production
  v3: accuracy=0.86  stage=archived

Purely deterministic SQL, so this reproduces exactly, every run, on any machine.

Walking through it

PRIMARY KEY (name, version) is what makes next_version safe to compute with a simple MAX(version) + 1. The database itself refuses to let two rows collide on the same name-and-version pair.

Registering never touches stage directly. Every new version starts at 'staging' by design. Nothing reaches production without an explicit promote call. Production is a decision, not a default.

get_by_stage is the function a serving system actually calls. It does not need to know version numbers at all. It asks "whichever model is currently production," and the registry answers, regardless of which version that happens to be today.

Common mistakes

Treating "highest version number" as "best model." v3 scored lower than v2 above — being newer does not mean being better. A registry that auto-promotes the latest version without checking its score will happily ship a regression.

No metadata worth trusting later. The metadata column here stores who trained it and on how much data — thin, but real. Model lineage goes further, but even this much turns "who made this?" from a guess into a lookup.

Building a registry that only tracks the file path, not the score. A registry with no accuracy or evaluation metric attached cannot answer "is this actually the better model?" It can only answer "where is it?" Both questions matter.

Try it yourself

Add a demote(name, version) function that moves whatever is currently in 'production' back to 'staging'. Call it automatically inside promote, right before a new version is promoted. That way, there is never more than one production version for a given model name at once.

What to learn next

Researcher — Mathematics and papers.

The registry as a state machine

Formally, each model version's lifecycle is a small state machine: staging -> production, staging -> archived, production -> archived, and (in most real registries) production -> staging on a rollback. Modelling it this way makes illegal transitions — jumping straight from a version nobody has evaluated to production — a thing the system can refuse, rather than a discipline problem left entirely to humans.

What production registries add beyond storage

MLflow's Model Registry, covered in MLflow, layers three things on top of the pattern shown above. Webhooks fire on a stage transition, triggering a deployment pipeline automatically. Approval workflows require a second person to sign off before a promotion completes. Aliases are a movable pointer like @champion that always resolves to whichever version currently holds it, so downstream code never needs updating when the underlying version changes.

Registry versus artifact store versus experiment tracker

These three are often confused, because one tool frequently provides all of them. They answer different questions, though:

SystemAnswers
Experiment tracker (experiment tracking)What happened during training — metrics, parameters, over time?
Artifact storeWhere are the actual files — model weights, datasets — physically stored?
Model registryWhich version is which, and which one is authoritative right now?

A registry typically references artifacts by pointer, rather than storing them itself. It is usually populated from the best runs an experiment tracker recorded. The three form a pipeline, not three competing tools.

Governance as the real differentiator at scale

Vartak et al., ModelDB: A System to Manage Machine Learning Models, HILDA 2016, is among the earliest academic treatments of tracking trained models as first-class, queryable objects, rather than opaque files. Production registries — MLflow, SageMaker Model Registry, Vertex AI Model Registry — all descend from that lineage of thinking. The paper's central argument, that a model without recorded provenance is effectively unauditable, is precisely the governance problem model lineage picks up in more depth.

What to learn next