Registries, Artifacts and Environments
Model registries
A model registry is the one place that lists every trained version of a model, and which one is actually live.
- 8 min read
- 3 reading levels
- Published
Read these first
On this page 9
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
A model registry is a catalog of every version of every model you have trained. It records which one is currently in use.
The analogy you have already lived
You have used a library. Every book has a catalog card — title, author, edition, shelf number. The catalog does not hold the books themselves. It tells you exactly where each one is, and which edition to borrow.
Without that catalog, a library is only a room full of books nobody can reliably find. A registry is that catalog, for trained models instead of books.
Why it exists
A single model you train once, on your laptop, is easy to keep track of. There is only one of it, and you remember what it is.
That stops being true fast. A real project accumulates versions. One trained last month, one retrained after new data arrived, one trained by a teammate, one currently answering real users. "Which model is actually live right now?" becomes a question nobody can answer confidently, without a shared record.
What a registry actually tracks
For each version of a model, a registry typically records where its file lives, how well it scored, and who or what produced it. It also records its stage — a label like staging, production, or archived, describing what that version is currently allowed to do.
How it works
train v1 -> register -> stage: staging (accuracy 0.83)
train v2 -> register -> stage: staging (accuracy 0.88)
|
promote v2 to production
v
stage: production
train v3 -> register -> stage: staging (accuracy 0.86, worse than v2)
|
v3 stays in staging, or gets archivedThe registry does not decide which model is best — a person, or an automated check, decides that. The registry's job is to make that decision recorded, visible, and reversible.
A real example you have seen
Any app store version history works the same way. Every past version of an app is still listed, and one is marked as current. Rolling back to an older version is a real, supported action, not a scramble to find an old file on someone's laptop.
The honest part
A registry is only as trustworthy as the discipline behind it. A model deployed by manually copying a file, skipping the registry entirely, breaks the one property a registry exists to guarantee. The record no longer matches reality. The tool cannot enforce that on its own — the team's habits do.
Remember this
- A registry answers "which version, and where is it?" It covers every model you have ever trained, not only the current one.
- Stage labels (staging, production, archived) record what each version is currently allowed to do.
- A registry is only trustworthy if every deployment actually goes through it — a manual shortcut breaks the record silently.
What to learn next
- What goes inside a model artifact — what a registry entry actually points to.
- Model lineage and traceability — tracing a registered model back to the data and code that produced it.
- MLflow — a real, widely used tool implementing everything this lesson built by hand.
Developer — Code and libraries.
Setup
No installs needed beyond the standard library. This builds a minimal registry from scratch, using SQLite, to show what a tool like MLflow's registry is doing underneath.
A working registry in under 40 lines
import json
import sqlite3
conn = sqlite3.connect(":memory:")
conn.execute("""
CREATE TABLE models (
name TEXT, version INTEGER, path TEXT,
accuracy REAL, stage TEXT, metadata TEXT,
PRIMARY KEY (name, version)
)
""")
def register(name, path, accuracy, metadata):
cur = conn.execute("SELECT COALESCE(MAX(version), 0) FROM models WHERE name = ?", (name,))
next_version = cur.fetchone()[0] + 1
conn.execute(
"INSERT INTO models VALUES (?, ?, ?, ?, 'staging', ?)",
(name, next_version, path, accuracy, json.dumps(metadata)),
)
conn.commit()
return next_version
def promote(name, version, stage):
conn.execute("UPDATE models SET stage = ? WHERE name = ? AND version = ?", (stage, name, version))
conn.commit()
def get_by_stage(name, stage):
return conn.execute(
"SELECT version, path, accuracy FROM models WHERE name = ? AND stage = ?",
(name, stage),
).fetchone()
# three training runs of the same model, over time
v1 = register("loan-scorer", "runs/v1/model.joblib", 0.83, {"trained_by": "pranay", "rows": 400})
v2 = register("loan-scorer", "runs/v2/model.joblib", 0.88, {"trained_by": "pranay", "rows": 900})
v3 = register("loan-scorer", "runs/v3/model.joblib", 0.86, {"trained_by": "pranay", "rows": 1500})
print(f"registered versions: {v1}, {v2}, {v3}")
promote("loan-scorer", v2, "production")
print("promoted version", v2, "to production")
print("current production model:", get_by_stage("loan-scorer", "production"))
promote("loan-scorer", v3, "archived") # v3 scored lower than v2
print("version 3 scored lower, sent to archived instead of production")
for v, acc, stage in conn.execute(
"SELECT version, accuracy, stage FROM models WHERE name = ? ORDER BY version", ("loan-scorer",)
):
print(f" v{v}: accuracy={acc:.2f} stage={stage}")registered versions: 1, 2, 3 promoted version 2 to production current production model: (2, 'runs/v2/model.joblib', 0.88) version 3 scored lower, sent to archived instead of production v1: accuracy=0.83 stage=staging v2: accuracy=0.88 stage=production v3: accuracy=0.86 stage=archived
Purely deterministic SQL, so this reproduces exactly, every run, on any machine.
Walking through it
PRIMARY KEY (name, version) is what makes next_version safe to compute with a simple MAX(version) + 1. The database itself refuses to let two rows collide on the same name-and-version pair.
Registering never touches stage directly. Every new version starts at 'staging' by design. Nothing reaches production without an explicit promote call. Production is a decision, not a default.
get_by_stage is the function a serving system actually calls. It does not need to know version numbers at all. It asks "whichever model is currently production," and the registry answers, regardless of which version that happens to be today.
Common mistakes
Treating "highest version number" as "best model." v3 scored lower than v2 above — being newer does not mean being better. A registry that auto-promotes the latest version without checking its score will happily ship a regression.
No metadata worth trusting later. The metadata column here stores who trained it and on how much data — thin, but real. Model lineage goes further, but even this much turns "who made this?" from a guess into a lookup.
Building a registry that only tracks the file path, not the score. A registry with no accuracy or evaluation metric attached cannot answer "is this actually the better model?" It can only answer "where is it?" Both questions matter.
Try it yourself
Add a demote(name, version) function that moves whatever is currently in 'production' back to 'staging'. Call it automatically inside promote, right before a new version is promoted. That way, there is never more than one production version for a given model name at once.
What to learn next
- What goes inside a model artifact — what a registry entry actually points to.
- Model lineage and traceability — tracing a registered model back to the data and code that produced it.
- MLflow — a real, widely used tool implementing everything this lesson built by hand.
Researcher — Mathematics and papers.
The registry as a state machine
Formally, each model version's lifecycle is a small state machine: staging -> production, staging -> archived, production -> archived, and (in most real registries) production -> staging on a rollback. Modelling it this way makes illegal transitions — jumping straight from a version nobody has evaluated to production — a thing the system can refuse, rather than a discipline problem left entirely to humans.
What production registries add beyond storage
MLflow's Model Registry, covered in MLflow, layers three things on top of the pattern shown above. Webhooks fire on a stage transition, triggering a deployment pipeline automatically. Approval workflows require a second person to sign off before a promotion completes. Aliases are a movable pointer like @champion that always resolves to whichever version currently holds it, so downstream code never needs updating when the underlying version changes.
Registry versus artifact store versus experiment tracker
These three are often confused, because one tool frequently provides all of them. They answer different questions, though:
| System | Answers |
|---|---|
| Experiment tracker (experiment tracking) | What happened during training — metrics, parameters, over time? |
| Artifact store | Where are the actual files — model weights, datasets — physically stored? |
| Model registry | Which version is which, and which one is authoritative right now? |
A registry typically references artifacts by pointer, rather than storing them itself. It is usually populated from the best runs an experiment tracker recorded. The three form a pipeline, not three competing tools.
Governance as the real differentiator at scale
Vartak et al., ModelDB: A System to Manage Machine Learning Models, HILDA 2016, is among the earliest academic treatments of tracking trained models as first-class, queryable objects, rather than opaque files. Production registries — MLflow, SageMaker Model Registry, Vertex AI Model Registry — all descend from that lineage of thinking. The paper's central argument, that a model without recorded provenance is effectively unauditable, is precisely the governance problem model lineage picks up in more depth.
What to learn next
- What goes inside a model artifact — what a registry entry actually points to.
- Model lineage and traceability — tracing a registered model back to the data and code that produced it.
- MLflow — a real, widely used tool implementing everything this lesson built by hand.