Registries, Artifacts and Environments
What goes inside a model artifact
A model artifact bundles the model file with everything needed to use it correctly, so nothing gets left behind when it ships.
- 8 min read
- 3 reading levels
- Published
Read these first
On this page 9
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
A model artifact is the model file bundled together with everything needed to use it correctly. It is not the model file alone.
The analogy you have already lived
You have packed a tiffin box for a full day away from home. Rice alone is not lunch. You pack the rice, the curry, a spoon. If you are careful, a note of what is inside, for whoever else might open it.
Leave out the spoon, and the rice is still perfectly good rice. It is completely useless to whoever opens the box expecting to eat.
Why it exists
A trained model file, by itself, is not enough to use safely. Which columns does it expect, and in what order? Was the input scaled before training — and if so, by what exact numbers? Which library version wrote this file, and can today's version even read it back?
None of that lives inside the raw model file. Someone has to write it down, and package it with the model. Otherwise every one of those questions becomes a guess the next person has to get right from memory.
What goes in the box
At minimum, a proper artifact bundles the model file itself, and the exact list of input columns in the exact order it expects them. It also needs enough version information to know whether the environment reading it back matches the one that wrote it.
How it works
model file alone:
model.joblib <- works, until someone guesses wrong
a real artifact:
artifact.zip
|- model.joblib
\- metadata.json
{ "feature_names": ["income", "years"],
"python_version": "3.11.4",
"sklearn_version": "1.4.0",
"training_data_hash": "ef93fa2d5d9b" }A loader that checks the metadata before trusting the model can refuse a broken or incomplete artifact outright. It does not load it and produce a wrong answer silently.
A real example you have seen
A software installer does the same thing at a different layer. It checks your operating system version before installing, rather than installing blindly and crashing halfway. It checks that what it is about to do actually matches the environment it is running in.
The honest part
Writing the metadata correctly, every single time, is a discipline problem as much as a technical one. It is easy to add a new feature to a model and forget to update the stored feature list. That mistake stays invisible, until someone tries to use an old artifact months later.
Remember this
- A model file alone is not enough — the exact input format and the exact library versions matter equally.
- A real artifact bundles the model with its metadata: feature names, versions, and enough context to use it correctly.
- A loader that validates the bundle before trusting it can refuse a broken artifact outright, instead of silently misusing it.
What to learn next
- Model lineage and traceability — tracing an artifact back to the exact code and data that produced it.
- Pinning ML dependencies — making the version metadata shown above actually enforceable, not only informational.
- Model registries — where a packaged artifact like this one gets registered and tracked over time.
Developer — Code and libraries.
Setup
pip install scikit-learn pandas joblibPackaging a real artifact, and rejecting a broken one
import hashlib
import json
import sys
import zipfile
from pathlib import Path
import joblib
import numpy as np
import pandas as pd
import sklearn
from sklearn.linear_model import LogisticRegression
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler
REQUIRED_KEYS = {"feature_names", "python_version", "sklearn_version", "training_data_hash"}
def train_and_package(artifact_path):
rng = np.random.RandomState(0)
X = pd.DataFrame({"income": rng.uniform(5, 80, 200), "years": rng.uniform(0, 10, 200)})
y = (0.05 * X["income"] + 0.4 * X["years"] - 2 > 0).astype(int)
model = make_pipeline(StandardScaler(), LogisticRegression()).fit(X, y)
data_hash = hashlib.sha256(pd.util.hash_pandas_object(X).values.tobytes()).hexdigest()[:12]
metadata = {
"feature_names": list(X.columns),
"python_version": sys.version.split()[0],
"sklearn_version": sklearn.__version__,
"training_data_hash": data_hash,
}
with zipfile.ZipFile(artifact_path, "w") as z:
tmp = Path(artifact_path).with_suffix(".tmp") # joblib needs a real file path
joblib.dump(model, tmp)
z.write(tmp, "model.joblib")
tmp.unlink()
z.writestr("metadata.json", json.dumps(metadata, indent=2))
return metadata
def load_artifact(artifact_path):
with zipfile.ZipFile(artifact_path) as z:
metadata = json.loads(z.read("metadata.json"))
missing = REQUIRED_KEYS - metadata.keys()
if missing:
raise ValueError(f"artifact is missing required metadata: {sorted(missing)}")
with z.open("model.joblib") as f:
model = joblib.load(f)
return model, metadata
metadata = train_and_package("artifact.zip")
print("packaged artifact.zip with metadata:")
for k, v in metadata.items():
print(f" {k}: {v}")
model, loaded_meta = load_artifact("artifact.zip")
print("\nloaded model back successfully:", type(model).__name__)
# Now package a broken artifact, missing the feature list, and try to load it.
with zipfile.ZipFile("broken.zip", "w") as z:
tmp = Path("model.tmp")
joblib.dump(model, tmp)
z.write(tmp, "model.joblib")
tmp.unlink()
z.writestr("metadata.json", json.dumps({"python_version": sys.version.split()[0]}))
try:
load_artifact("broken.zip")
except ValueError as e:
print(f"\nbroken.zip correctly rejected: {e}")packaged artifact.zip with metadata: feature_names: ['income', 'years'] python_version: 3.10.11 sklearn_version: 1.7.2 training_data_hash: ef93fa2d5d9b loaded model back successfully: Pipeline broken.zip correctly rejected: artifact is missing required metadata: ['feature_names', 'sklearn_version', 'training_data_hash']
training_data_hash and feature_names are exact reproductions of this fixed-seed data — they will match on any machine. python_version and sklearn_version reflect whatever is installed where you run this, and will differ from the values shown here. That difference is, in fact, exactly the thing this lesson's validation step is designed to catch.
Walking through it
hashlib.sha256(pd.util.hash_pandas_object(X).values.tobytes()) produces a short fingerprint of the exact training data used. Two artifacts trained on data that differs by even one value will show different hashes. It is a cheap, precise way to answer "was this really trained on the data I think it was?"
The zip format is doing real work here, not only convenience. A single file that bundles the model and its metadata together cannot be separated by accident. Copy the artifact, and both pieces travel together, or neither does.
REQUIRED_KEYS - metadata.keys() is a set difference — every key the loader demands that the metadata does not actually have. It is what turns "missing a field" from a mysterious later crash into an immediate, named error.
Common mistakes
Shipping the model file and calling it done. This is the single most common shortcut, and it is exactly what this lesson exists to push back on. A .joblib or .pt file with no accompanying metadata is a liability, waiting for someone else's guess to be wrong.
Validating metadata only for presence, never for correctness. The loader above checks the required keys exist. A stricter loader would also check sklearn_version matches what is installed, and warn or refuse otherwise — the subject of pinning ML dependencies.
Hashing the wrong thing. Hashing the file path, or the row count, catches almost nothing — two very different datasets can share both. Hashing actual data content, as done above, is what makes the fingerprint mean anything.
Try it yourself
Add a preprocessing_notes field to metadata — free text describing any manual data cleaning that happened before training, outside the code. Nothing about the artifact format prevents this; the required-keys check only enforces what is mandatory, and extra context is always welcome in a metadata file.
What to learn next
- Model lineage and traceability — tracing an artifact back to the exact code and data that produced it.
- Pinning ML dependencies — making the version metadata shown above actually enforceable, not only informational.
- Model registries — where a packaged artifact like this one gets registered and tracked over time.
Researcher — Mathematics and papers.
Standard artifact formats, and what each one solves
- ONNX (Open Neural Network Exchange) standardises the model graph itself, so a model trained in PyTorch can run in a runtime that has never heard of PyTorch. It solves cross-framework portability, not metadata completeness — an ONNX file still benefits from the same bundled-metadata pattern shown above.
- MLflow's
MLmodelformat is closest to what this lesson built by hand: a directory with the model file(s), a YAML manifest describing the "flavor" (which framework can load it), input/output signature, and pip requirements, all read back together by MLflow's own loader. It formalises the same required-keys idea shown above, at production scale. - Model cards (Mitchell et al., Model Cards for Model Reporting, FAT* 2019) standardise a different, complementary kind of metadata — intended use, known limitations, evaluation across subgroups — aimed at human reviewers and downstream decision-makers rather than at a loader function.
Input/output signatures as executable documentation
A signature is a formal description of expected input and output shapes and types, checked at load or inference time rather than only documented in prose. It turns "the model expects three float columns in this order" from a comment someone might not read, into a check the system enforces automatically. MLflow, TensorFlow Serving's SavedModel format, and ONNX Runtime all support this pattern; it is the artifact-level analog of the pydantic validation shown in the developer section of model serving.
The provenance problem this partially solves
An artifact with complete, correct metadata answers "how do I use this correctly?" It does not, by itself, answer "where did this come from?" — which code, which data, which parent model. That deeper question is model lineage, covered next, and Gebru et al.'s Datasheets for Datasets (CACM 2021) makes the equivalent case for datasets that Mitchell et al. make for models: undocumented provenance is a systemic risk, not a small inconvenience.
What to learn next
- Model lineage and traceability — tracing an artifact back to the exact code and data that produced it.
- Pinning ML dependencies — making the version metadata shown above actually enforceable, not only informational.
- Model registries — where a packaged artifact like this one gets registered and tracked over time.