Versioning a model API
Versioning a model API means freezing the exact shape of a contract once real callers depend on it, the way you keep an old phone number forwarding for a while after getting a new one, so nobody who still has it suddenly cannot reach you.
- 11 min read
- 3 reading levels
- Published
Read these first
On this page 7
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
Versioning a model API means freezing the exact shape of a request and response once real callers depend on it. Any breaking change goes behind a new version instead.
The analogy you have already lived
You have changed your phone number and kept the old one forwarding calls for a while, before letting it go dead. Anyone who still had the old number could reach you during that window. Nobody who had saved your number in their phone was suddenly cut off the moment you switched.
An API version is that same courtesy, for software. v1 is the old number — still answering, exactly as it always did, for as long as it is promised to. v2 is the new one, for anyone ready to use it. Nobody calling v1 should ever be surprised by it suddenly behaving differently.
Why it exists
A model's serving code changes constantly — new fields, a renamed value, a more detailed response. Most of those changes feel harmless from inside the team making them.
They are not harmless to whatever is calling the API from outside. Picture a mobile app already published to millions of phones, or an integration another company wired up eight months ago. Or picture a script a colleague wrote and forgot about. Each one is reading a specific field, by name, expecting a specific type. Change that field under them, and every one of those callers breaks, all at once, usually without warning.
Versioning draws a clear line: v1's shape is a promise, not a suggestion. Anything that needs to break that promise ships as v2 instead of quietly editing v1.
How it works
caller hits /v1/predict
|
v
ALWAYS the same response shape -- frozen, the day it was published
{"repaid": 1, "confidence": 0.94}
caller hits /v2/predict
|
v
the current shape -- free to grow and change over time
{"decision": "approve", "probability": 0.9941, "model_version": "2026.03"}Both versions can be served by the exact same underlying model. Versioning is a promise about the shape of the answer, not necessarily a different model behind it.
A real example you have seen
A payments API that still honours requests in an older format, years after introducing a newer one, is protecting exactly this. Every business that integrated years ago keeps working, without needing to rewrite their code the moment the provider wants to improve something.
Remember this
- A version is a promise about shape — once real callers depend on it, that shape does not change underneath them.
- A breaking change (a renamed field, a changed type, a changed meaning) belongs in a new version, never edited into an old one.
- Retiring an old version needs a clear, advance-notice plan — not a surprise.
What to learn next
- KServe and Seldon — production platforms that handle versioned model serving, including much of this lesson's contract and routing logic, for you.
- Switching between model providers — a related but distinct kind of contract change, where the thing behind your API is not your own model at all.
- Model registries — tracking the model-version axis this lesson's researcher section separates from the contract-version axis.
Developer — Code and libraries.
Setup
pip install scikit-learn pandas numpy pytestTwo contracts, one model
"""Two versions of the same underlying model, served under two frozen,
independent contracts. v1 never changes shape once real callers depend on it.
"""
import numpy as np
import pandas as pd
from sklearn.linear_model import LogisticRegression
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler
COLUMNS = ["income", "years", "age"]
MODEL_VERSION = "2026.03"
def make_training_data(n=400, seed=0):
rng = np.random.RandomState(seed)
X = pd.DataFrame({
"income": rng.uniform(5, 80, n).round(1),
"years": rng.uniform(0, 10, n).round(1),
"age": rng.randint(21, 65, n).astype(float),
})
score = 0.05 * X["income"] + 0.35 * X["years"] + 0.01 * X["age"] - 3.5
y = (score + rng.normal(0, 0.8, n) > 0).astype(int)
return X, y
X, y = make_training_data()
MODEL = make_pipeline(StandardScaler(), LogisticRegression(max_iter=1000)).fit(X, y)
def _score(applicant: dict) -> float:
row = pd.DataFrame([applicant], columns=COLUMNS)
return float(MODEL.predict_proba(row)[0, 1])
def predict_v1(applicant: dict) -> dict:
"""FROZEN since launch. Exactly these two keys, these two types. Never add,
remove, or rename a field here.
"""
probability = _score(applicant)
return {
"repaid": int(probability >= 0.5),
"confidence": round(probability, 2),
}
def predict_v2(applicant: dict) -> dict:
"""The current version. Free to evolve -- this is where new fields go."""
probability = _score(applicant)
return {
"decision": "approve" if probability >= 0.5 else "decline",
"probability": round(probability, 4),
"model_version": MODEL_VERSION,
}
applicant = {"income": 60.0, "years": 7.0, "age": 34.0}
print("v1:", predict_v1(applicant))
print("v2:", predict_v2(applicant))v1: {'repaid': 1, 'confidence': 0.99}
v2: {'decision': 'approve', 'probability': 0.9941, 'model_version': '2026.03'}v2 renamed repaid to decision and changed it from an integer to a string, added a full-precision probability instead of a rounded confidence, and added a version tag v1 never had. Every one of those is a breaking change for anything still reading v1's shape — which is exactly why they live in v2 instead of being edited into v1.
A contract test that actually enforces the freeze
A promise that is only written in a design document gets broken eventually, by someone who never read the document. A promise enforced by a test gets caught the moment it is broken, by name.
from api import predict_v1, predict_v2
V1_APPLICANT = {"income": 60.0, "years": 7.0, "age": 34.0}
def test_v1_response_has_exactly_the_frozen_keys_and_types():
response = predict_v1(V1_APPLICANT)
assert set(response.keys()) == {"repaid", "confidence"}
assert isinstance(response["repaid"], int)
assert response["repaid"] in (0, 1)
assert isinstance(response["confidence"], float)
assert 0.0 <= response["confidence"] <= 1.0
def test_v1_repaid_field_is_never_renamed_to_decision():
# a real, common "helpful improvement" that would break every existing v1 caller
response = predict_v1(V1_APPLICANT)
assert "decision" not in response
assert "repaid" in response
def test_v2_is_free_to_carry_more_fields_than_v1():
response = predict_v2(V1_APPLICANT)
assert set(response.keys()) == {"decision", "probability", "model_version"}pytest test_contract.py -q... [100%] 3 passed in 0.86s
Now break something on purpose
A well-meaning "improvement": add a more precise probability field to v1, alongside the existing confidence — surely harmless, it only adds information.
# in api.py, predict_v1 becomes:
return {
"repaid": int(probability >= 0.5),
"confidence": round(probability, 2),
"probability": round(probability, 4), # "only adding more detail" -- feels harmless
}F.. [100%]
================================== FAILURES ===================================
___________ test_v1_response_has_exactly_the_frozen_keys_and_types ____________
def test_v1_response_has_exactly_the_frozen_keys_and_types():
response = predict_v1(V1_APPLICANT)
> assert set(response.keys()) == {"repaid", "confidence"}
E AssertionError: assert {'confidence'...ty', 'repaid'} == {'confidence', 'repaid'}
E
E Extra items in the left set:
E 'probability'
E Use -v to get more diff
test_contract.py:8: AssertionError
=========================== short test summary info ===========================
FAILED test_contract.py::test_v1_response_has_exactly_the_frozen_keys_and_types
1 failed, 2 passed in 1.28sThis is a genuinely debatable case, which is exactly why it is worth including: adding a field rarely breaks a caller reading specific known fields by name — but it is not free of risk either. Some client libraries validate a response against a strict schema and reject anything with an unexpected extra field; some callers iterate over every key in the response and would now see one they do not recognise. The contract test does not decide whether this change is acceptable — it makes sure the decision is made on purpose, by a human looking at a failing, precisely-named test, rather than by accident.
Common mistakes
Treating "nobody complained yet" as proof nothing broke. A caller that silently ignores an unexpected field, or crashes in a way nobody is monitoring, does not complain — it quietly fails instead. A contract test catches the change at the moment it is made, which is far cheaper than waiting for a bug report.
No version at all, ever. A service with one unversioned endpoint has no way to improve the contract without breaking every existing caller on that exact day. Even a service with exactly one real client benefits from v1 existing as a name, because it makes the next change a deliberate decision instead of a silent edit.
Versioning the model but not the API. Retraining a model and improving its accuracy is not, by itself, a contract-breaking change — the shape of predict_v1's response is unaffected by which weights are behind _score. Confusing "the model changed" with "the API changed" leads teams to either version too often or not carefully enough.
Retiring a version with no notice. Turning off v1 the moment v2 ships breaks every caller who has not migrated yet, which defeats the entire purpose of having versioned it in the first place. Publish a deprecation date, communicate it, and only remove v1 after that date has actually passed.
Putting the version somewhere inconsistent. Whether it lives in the URL path (/v1/predict), a header (API-Version: 1), or the request body, pick one convention and use it everywhere in your API — a service that versions some endpoints one way and others a different way is a constant source of integration mistakes for callers.
Try it yourself
Add a third contract test asserting that predict_v1's confidence field is always rounded to exactly two decimal places — not one, not three. Then change the rounding to round(probability, 3) and watch the test name tell you precisely what changed, without needing to read a diff to find out.
What to learn next
- KServe and Seldon — production platforms that handle versioned model serving, including much of this lesson's contract and routing logic, for you.
- Switching between model providers — a related but distinct kind of contract change, where the thing behind your API is not your own model at all.
- Model registries — tracking the model-version axis this lesson's researcher section separates from the contract-version axis.
Researcher — Mathematics and papers.
Contracts as a testable invariant, not documentation
The contract tests above formalise an API version as an invariant on the output space of a function: for predict_v1, the set of keys, their types, and their value ranges must remain fixed across arbitrary changes to the model or the code producing them. This reframes API versioning from a documentation and process discipline into something checkable by the same testing infrastructure covered throughout testing and CI for ML — a contract test belongs in the same CI gate as a model-quality gate, because a broken contract is exactly as much a release blocker as a broken model.
Semantic versioning applied to a probabilistic contract
Conventional semantic versioning (MAJOR.MINOR.PATCH) maps awkwardly onto a model API, because a change can be behaviourally significant without changing the shape of the contract at all — a retrained model with materially different accuracy is a MINOR or PATCH change to the API surface but can be a MAJOR change in practice for a caller relying on specific decision behaviour. A more complete versioning scheme for a model API tracks two independent axes: the contract version (this lesson's v1/v2, changed only on a shape-breaking change) and a separate model version (MODEL_VERSION above, changed on every retrain), exposed together so a caller — or an incident investigation — can distinguish "the API changed" from "the model behind it changed" as two different, independently-occurring events.
Deprecation as a first-class, observable state
RFC 8594 (The Sunset HTTP Header Field) and the informal Deprecation header convention give a machine-readable way to announce that an endpoint will stop working on a specific date, allowing well-behaved clients to detect and react to deprecation automatically rather than relying on callers reading a changelog. Combined with usage logging on the deprecated version (who is still calling v1, and how often), this turns "can we retire this yet" from a guess into a measured decision — retiring a version nobody has called in ninety days carries a fundamentally different risk profile than retiring one still serving significant real traffic.
Backward and forward compatibility as distinct properties
Backward compatibility: new server code correctly serves old client requests. Forward compatibility: old server code tolerates new client requests it does not fully understand (commonly by ignoring unknown fields rather than rejecting them). A well-designed contract — using a schema library that ignores unrecognised fields by default rather than rejecting them, for instance — buys forward compatibility for free and makes the "add a field" case from the developer section unambiguously safe, removing the debate the contract test above deliberately leaves open for a stricter schema.
Papers and standards
- Fielding, Architectural Styles and the Design of Network-based Software Architectures, PhD dissertation, 2000 — the REST constraints, including the case for versioning at the resource/representation level rather than embedding it invisibly in behaviour.
- RFC 8594, The Sunset HTTP Header Field, IETF 2019 — datatracker.ietf.org/doc/html/rfc8594
- Postel, TCP, RFC 761, 1980 — the origin of the "be conservative in what you send, liberal in what you accept" robustness principle underlying forward compatibility.
What to learn next
- KServe and Seldon — production platforms that handle versioned model serving, including much of this lesson's contract and routing logic, for you.
- Switching between model providers — a related but distinct kind of contract change, where the thing behind your API is not your own model at all.
- Model registries — tracking the model-version axis this lesson's researcher section separates from the contract-version axis.