Monitoring Models in Production
Training-serving skew
Training-serving skew is when the same real event produces a different feature value depending on which code computed it, so a model gets a question it was never trained to answer.
- 8 min read
- 3 reading levels
- Published
Read these first
On this page 8
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
Training-serving skew is when the code that builds a feature for training quietly disagrees with the code that builds it for real predictions.
The analogy you have already lived
Learner drivers in many driving schools practice on an automatic car. No clutch, no gear stick, only an accelerator and a brake.
On test day, the examiner hands them a manual car instead. Same driver, same skill, same roads. A completely different set of pedals under their feet.
The driver did not get worse. The conditions they were tested under stopped matching the conditions they trained under. That gap is training-serving skew, aimed at machine learning.
Why it exists
A model does not see raw events. It sees features — numbers computed from those events, like "days since signup" or "average order value".
Two different pieces of code usually build those numbers. One runs offline, in a batch job, written by whoever trained the model. Another runs online, inside the live service, often written later, by someone else, in a different language.
Nobody plans for these two to disagree. They are supposed to compute the exact same thing. In practice, a rounding rule, a time zone, or a unit gets handled differently in each, and the two quietly drift apart.
How it works
the SAME real event
/ \
training code serving code
(batch job) (live service)
| |
account age: 30 days account age: 30000 days
| |
the model was never shown a feature value like thisThe model itself did not change between training and serving. The number it was handed for the exact same real event did.
A real example you have seen
A recommendation system trained offline computes "how many items you viewed today" using a clean daily log file, counted once a day.
The live service counts the same thing from a real-time event stream, which can double-count a retried request. The live number for a real, ordinary user can end up larger than any number the model ever saw in training.
The honest part
This is one of the hardest bugs in production machine learning to catch by reading code. Both pipelines can look correct in isolation.
The only reliable way to catch it: compare real feature values from both pipelines, side by side, on a schedule. Reading either pipeline alone will not reveal it.
Remember this
- Training-serving skew is two different pieces of code disagreeing about the same feature.
- It hides even when both pipelines look correct on their own.
- Catch it by comparing real feature values from both pipelines, not by reading code.
What to learn next
- Working with delayed labels — the next problem once your features are trustworthy.
- Logging every prediction — how you capture the real serving-time feature values this comparison needs.
- Model serving — the serving code this lesson assumes you already have running.
Developer — Code and libraries.
Setup
pip install pandas numpy scikit-learnA believable version of the bug
A training pipeline computes account age correctly from proper timestamps. A serving pipeline receives the same event as milliseconds, and divides by a divisor that was correct for a different upstream source.
import pandas as pd
import numpy as np
from sklearn.linear_model import LogisticRegression
rng = np.random.RandomState(0)
# ---- TRAINING PIPELINE (offline, batch, built by the data team) ----
# signup_time and event_time are proper pandas Timestamps in a warehouse table.
signup = pd.Timestamp("2026-01-01")
event_times = pd.to_datetime("2026-01-01") + pd.to_timedelta(
rng.randint(1, 400, 2000), unit="D"
)
def account_age_days_training(event_time, signup_time):
return (event_time - signup_time).days
train_age = [account_age_days_training(t, signup) for t in event_times]
# A believable rule: older accounts churn less.
y = (np.array(train_age) < 60).astype(int)
model = LogisticRegression().fit(np.array(train_age).reshape(-1, 1), y)
print("train accuracy:", round(model.score(np.array(train_age).reshape(-1, 1), y), 3))
# ---- SERVING PIPELINE (a second engineer, a second codebase, months later) ----
# The event arrives from a different system as epoch MILLISECONDS, not seconds.
signup_epoch_ms = int(signup.timestamp() * 1000)
def account_age_days_serving(event_epoch_ms, signup_epoch_ms):
# Written for an older upstream event that sent epoch SECONDS.
# This one sends epoch MILLISECONDS. The divisor was never updated.
return (event_epoch_ms - signup_epoch_ms) / 86400
sample_event = signup + pd.Timedelta(days=30) # a genuinely young account
sample_epoch_ms = int(sample_event.timestamp() * 1000)
correct = account_age_days_training(sample_event, signup)
skewed = account_age_days_serving(sample_epoch_ms, signup_epoch_ms)
print(f"same real event -- training feature: {correct} days")
print(f"same real event -- serving feature: {skewed:.1f} days (off by 1000x)")
pred_correct = model.predict([[correct]])[0]
pred_skewed = model.predict([[skewed]])[0]
print(f"prediction using the training-shaped feature: {pred_correct}")
print(f"prediction using the serving-shaped feature: {pred_skewed}")train accuracy: 1.0 same real event -- training feature: 30 days same real event -- serving feature: 30000.0 days (off by 1000x) prediction using the training-shaped feature: 1 prediction using the serving-shaped feature: 0
Every number above is exact, deterministic output from this script. A real production skew is rarely a clean 1000x. It is usually a few percent, which is why it survives so long unnoticed.
Line-by-line walkthrough
account_age_days_training uses proper pandas.Timestamp subtraction, which handles units correctly by construction.
account_age_days_serving reimplements the same idea by hand, in raw arithmetic, against a differently-shaped input. That reimplementation is the entire bug.
The model itself is never touched. It was trained once, correctly, and is applied unchanged to both feature values. The prediction flips only because the number reaching it flipped.
Common mistakes
Writing the serving feature code from memory, instead of sharing the training code. The single biggest fix available is making one function do both jobs — call the same account_age_days_training from both the batch job and the live service, and this entire bug class becomes impossible.
Never comparing real feature values between the two pipelines. Log the serving-time feature values alongside predictions, as covered in logging every prediction, and periodically diff a sample against a freshly recomputed training-side value for the same real events.
Testing only with hand-written example inputs. Hand-written test events tend to accidentally use round, well-behaved numbers. Test with real logged events, which contain the messy edge cases that reveal unit mismatches.
Assuming a passing unit test on each pipeline separately proves they agree. Both functions can have perfect unit tests and still disagree, because a unit test author usually assumes the same units the buggy code assumes.
Try it yourself
Add a third event exactly 60 days after signup, the boundary the model's rule flips on. Compute both feature versions for it, and check whether the two pipelines' predictions agree at that boundary too.
What to learn next
- Working with delayed labels — the next problem once your features are trustworthy.
- Logging every prediction — how you capture the real serving-time feature values this comparison needs.
- Model serving — the serving code this lesson assumes you already have running.
Researcher — Mathematics and papers.
Where skew enters a real feature pipeline
Sculley et al. (2015) name this among the "hidden technical debt" categories that make ML systems costly to maintain, under the broader heading of CACE (Changing Anything Changes Everything) and pipeline entanglement. Common concrete sources:
- Language or runtime mismatch — a Python/pandas offline pipeline reimplemented in Java or Go for low-latency serving, by a different team.
- Windowing mismatch — an offline aggregate computed over a clean, closed daily partition, against an online aggregate computed over an open, still-arriving stream.
- Null-handling mismatch — offline code that drops missing values silently, online code that must return a value for every request and substitutes a default.
- Point-in-time leakage in reverse — offline features accidentally allowed to see data that would not yet exist at serving time, producing training accuracy that serving can never match.
Detecting it without a labelled outcome
Because skew is a feature-value disagreement, it does not require waiting for labels the way concept drift does. The direct method: log a sample of production requests with their raw inputs, recompute features for those same inputs through the offline pipeline, and diff the two feature vectors.
$$\delta_j = \frac{1}{n}\sum_{i=1}^{n} \mathbb{1}\left[\,|f^{\text{train}}_j(x_i) - f^{\text{serve}}_j(x_i)| > \tau_j\,\right]$$
Where $f^{\text{train}}_j$ and $f^{\text{serve}}_j$ are the training-time and serving-time implementations of feature $j$, $\tau_j$ is a per-feature tolerance (zero for exact-match categorical features, an epsilon for floating point), and $\delta_j$ is the fraction of sampled requests where the two disagree beyond that tolerance. TensorFlow Data Validation's training-serving skew detector (Breck et al., 2019) implements a variant of this comparison as a first-class pipeline stage.
The structural fix
The design that eliminates this bug class rather than detecting it after the fact: a shared feature-transformation layer, compiled or interpreted identically in both the batch and serving paths, so $f^{\text{train}}_j \equiv f^{\text{serve}}_j$ by construction rather than by convention. Feature stores (Feast, Tecton) and frameworks like TFX's tf.Transform exist largely to guarantee this identity.
Papers
- Sculley et al., Hidden Technical Debt in Machine Learning Systems, NeurIPS 2015 — papers.nips.cc/paper/5656
- Breck, Polyzotis, Roy, Whang and Zinkevich, Data Validation for Machine Learning, SysML 2019
- Baylor et al., TFX: A TensorFlow-Based Production-Scale Machine Learning Platform, KDD 2017
What to learn next
- Working with delayed labels — the next problem once your features are trustworthy.
- Logging every prediction — how you capture the real serving-time feature values this comparison needs.
- Model serving — the serving code this lesson assumes you already have running.