Registries, Artifacts and Environments
Training and serving environment parity
Training and serving must compute the same features the same way, or a model can score identically well and still answer differently in production.
- 10 min read
- 3 reading levels
- Published
Read these first
On this page 8
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
Environment parity means training and serving compute every feature the exact same way. The same input then always produces the same answer, regardless of where it is scored.
The analogy you have already lived
You have rehearsed a school play on the classroom floor. Then you performed it on a real stage that was a different size. Steps that worked perfectly in rehearsal landed in the wrong spot on the actual night. The floor itself had quietly changed.
A model is rehearsed during training, on one specific setup for computing its inputs. If serving computes those same inputs a different way, even slightly, the model steps onto a different stage than the one it rehearsed on.
Why it exists
A model does not see raw data. It sees features — numbers computed from raw data. Think "income divided by age," or "a value scaled against the training set's average." Training and serving often use two completely separate pieces of code to compute those same features. One is written for a batch pipeline. One is rewritten for speed inside a live server.
Two pieces of code, each trying to do the "same" calculation, can drift apart. Nobody notices, until the numbers stop matching. This drift has a name: training-serving skew.
How it works
training: income / age, using the whole training set's stored average
|
v
model learns to expect features computed exactly that way
|
v
serving: a well-meaning rewrite computes "the same" feature
but using a different, quietly incorrect method
|
v
the model receives a number it never saw the meaning of during training
|
v
the answer changes -- not because the applicant changed,
but because the arithmetic feeding the model changedA real example you have seen
A weighing scale that gives a different number depending on which shop's floor it stands on has an environment problem, not a measurement problem. The object never changed weight. A model that scores the same person differently, depending on which server answered, has exactly this failure — computed instead of physical.
The honest part
This is one of the hardest bugs in machine learning to catch. Both the training pipeline and the serving code can look completely correct in isolation. The bug only exists in the difference between them. Nothing about either piece of code, read on its own, reveals it.
Remember this
- Training-serving skew happens when the same logical feature is computed two different ways, in two different places.
- A model can score perfectly well in training and still misbehave in production, purely from this mismatch. The model itself was never wrong.
- The safest fix is to share the exact same code between training and serving. Not two versions of "the same" logic.
What to learn next
- Model serving — the pipeline discipline (scaling inside the saved object) that prevents this bug in the first place.
- Pinning ML dependencies — closing the library-version half of the parity gap this lesson focuses on for feature logic.
- CI/CD for machine learning — where a parity check like this one belongs, so it runs on every change instead of being discovered in production.
Developer — Code and libraries.
Setup
pip install scikit-learn pandas joblibA real skew bug, and the exact numbers it produces
This trains a small model, then scores the same applicant two ways. First through the correctly saved pipeline. Then through a common, real mistake — serving code that recomputes normalisation statistics from whatever batch happens to be nearby, instead of reusing the frozen statistics from training.
import joblib
import numpy as np
import pandas as pd
from sklearn.linear_model import LogisticRegression
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler
rng = np.random.RandomState(0)
n = 400
X = pd.DataFrame({
"income": rng.uniform(5, 80, n).round(1),
"years": rng.uniform(0, 10, n).round(1),
})
y = (0.05 * X["income"] + 0.4 * X["years"] - 2 + rng.normal(0, 0.5, n) > 0).astype(int)
model = make_pipeline(StandardScaler(), LogisticRegression()).fit(X, y)
joblib.dump(model, "parity_model.joblib")
applicant = pd.DataFrame([{"income": 40.0, "years": 3.0}])
correct_score = model.predict_proba(applicant)[0, 1]
print(f"the applicant's true score, from the saved pipeline: {correct_score:.4f}\n")
def buggy_serving_score(applicant_row, batch):
"""A real mistake: re-fitting 'standardisation' on whatever else is in
the current batch, instead of reusing the training-time mean and std."""
combined = pd.concat([batch, applicant_row], ignore_index=True)
scaled = (combined - combined.mean()) / combined.std()
lr = model.named_steps["logisticregression"]
z = applicant_row.copy()
z.iloc[0] = scaled.iloc[-1]
return float(1 / (1 + np.exp(-(lr.intercept_[0] + z.values[0] @ lr.coef_[0]))))
def correct_serving_score(applicant_row):
"""Reuses the exact StandardScaler fitted during training. No refitting."""
return float(model.predict_proba(applicant_row)[0, 1])
quiet_batch = X.sample(5, random_state=1)
busy_batch = X.sample(5, random_state=2) * 1.8 # a batch of unusually high earners
print("buggy serving code, same applicant, two different surrounding batches:")
print(f" next to a quiet batch: {buggy_serving_score(applicant, quiet_batch):.4f}")
print(f" next to a busy batch: {buggy_serving_score(applicant, busy_batch):.4f}")
print("\ncorrect serving code, same two situations:")
print(f" next to a quiet batch: {correct_serving_score(applicant):.4f}")
print(f" next to a busy batch: {correct_serving_score(applicant):.4f}")the applicant's true score, from the saved pipeline: 0.9589 buggy serving code, same applicant, two different surrounding batches: next to a quiet batch: 0.9989 next to a busy batch: 0.5027 correct serving code, same two situations: next to a quiet batch: 0.9589 next to a busy batch: 0.9589
Deterministic with the fixed seeds throughout — this reproduces exactly on any machine, every run.
Walking through it
The exact same applicant scores 0.9989 in one batch and 0.5027 in another, with the buggy code. Nothing about the applicant changed between those two lines — only which other rows happened to be nearby when the "standardisation" was computed. A score that depends on unrelated people in the same batch is not a model behaving correctly.
The correct code gives the identical answer, 0.9589, regardless of the batch. It reuses the exact StandardScaler that was fitted once, during training, and saved inside the pipeline. That is the same discipline the training script in model serving already follows, keeping the scaler inside the saved pipeline rather than as a separate step.
The bug is subtle because both versions "compute a z-score." Read in isolation, buggy_serving_score looks like a reasonable, even faster, reimplementation. The defect only shows up by comparing its output against the correct version, side by side. That is exactly why this class of bug survives code review so often.
Common mistakes
Reimplementing preprocessing by hand "for speed" in the serving path. The pipeline object already knows exactly how to transform new data correctly — calling model.predict_proba on it, as correct_serving_score does, is both simpler and correct. A hand-written reimplementation is extra code that can drift, for a speed gain that is rarely worth the risk.
Computing any statistic — mean, std, min, max — from the incoming request or batch, instead of from training. Anything computed "live" from whatever data happens to be nearby depends on that data, by definition. Only statistics frozen at training time, and saved alongside the model as in packaging a model artifact, give a stable, reproducible answer.
Trusting that two pipelines are equivalent because their code "looks similar." As shown above, similar-looking code can compute genuinely different numbers. Test parity directly — score identical known inputs through both paths, and assert the outputs match — rather than reviewing the code by eye alone.
Assuming this only happens with hand-written features. Even automated feature pipelines can skew, if the batch job and the real-time job pull from different underlying tables, or apply a transformation library at two different versions. It is the version-pinning problem from pinning ML dependencies, applied specifically to feature computation.
Try it yourself
Change busy_batch to X.sample(5, random_state=2) * 5 — an even more extreme batch. Watch the buggy score move even further from the correct 0.9589, while the correct version, unaffected by any batch at all, stays exactly the same.
What to learn next
- Model serving — the pipeline discipline (scaling inside the saved object) that prevents this bug in the first place.
- Pinning ML dependencies — closing the library-version half of the parity gap this lesson focuses on for feature logic.
- CI/CD for machine learning — where a parity check like this one belongs, so it runs on every change instead of being discovered in production.
Researcher — Mathematics and papers.
Where skew actually enters real systems
Breck et al.'s The ML Test Score (also covered in the researcher section of CI/CD for machine learning) treats training-serving skew detection as a named, checkable property — Infra 5 in that rubric's numbering. Common real sources, beyond the normalisation-statistics bug demonstrated above:
- Different code paths entirely — a Python training pipeline and a Java or C++ serving path, each independently implementing "the same" feature logic, with no shared source of truth between them.
- Time-of-computation mismatch — a feature computed as "days since signup" behaves differently depending on whether it is computed once, in a nightly batch, versus live at request time, even given identical underlying code.
- Silent library version drift — the same feature-engineering code, running under two different versions of pandas or NumPy, can change behaviour at edge cases (a rounding rule, a NaN-handling default) without either version being "wrong" in isolation.
- Leakage of the label itself — the most severe case, where a feature available at training time (because it was computed after the outcome was already known) does not exist yet at serving time, forcing a different, non-equivalent computation in production.
Feature stores as the structural fix
A feature store stores feature definitions once, computed by one shared pipeline. It serves the same computed values to both training (in batch, historically) and serving (in real time, from a low-latency store). This eliminates the two-implementations problem at its root, rather than catching skew after the fact. Feast, an open-source feature store, formalises this pattern. So does Michelangelo, Uber's internally built ML platform, documented in Uber Engineering's public writeups. The trade-off is real infrastructure and operational cost, in exchange for a structural guarantee rather than a discipline-dependent one.
Detecting skew statistically, when a shared feature store is not feasible
Where two independent implementations remain, TensorFlow Data Validation's training-serving skew detector (also covered in the researcher section of CI/CD for machine learning) offers a safety net. It compares the statistical distribution of each feature as seen at training time against the same feature as seen at serving time, flagging features whose distributions diverge beyond a configured threshold. This is a monitoring-based check that catches skew introduced after deployment, complementing the artifact-level checks from pinning ML dependencies and what goes inside a model artifact.
References
- Breck et al., The ML Test Score, IEEE Big Data 2017 — research.google/pubs/pub46555
- Breck et al., Data Validation for Machine Learning, SysML 2019 — mlsys.org/Conferences/2019/doc/2019/167.pdf
- Sculley et al., Hidden Technical Debt in Machine Learning Systems, NeurIPS 2015 — names training-serving skew among the paper's catalog of ML-specific technical debt, several years before the tooling above existed to detect it automatically.
What to learn next
- Model serving — the pipeline discipline (scaling inside the saved object) that prevents this bug in the first place.
- Pinning ML dependencies — closing the library-version half of the parity gap this lesson focuses on for feature logic.
- CI/CD for machine learning — where a parity check like this one belongs, so it runs on every change instead of being discovered in production.