Monitoring Models in Production

Training-serving skew

Training-serving skew is when the same real event produces a different feature value depending on which code computed it, so a model gets a question it was never trained to answer.

On this page 8
  1. The short answer
  2. The analogy you have already lived
  3. Why it exists
  4. How it works
  5. A real example you have seen
  6. The honest part
  7. Remember this
  8. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

Training-serving skew is when the code that builds a feature for training quietly disagrees with the code that builds it for real predictions.

The analogy you have already lived

Learner drivers in many driving schools practice on an automatic car. No clutch, no gear stick, only an accelerator and a brake.

On test day, the examiner hands them a manual car instead. Same driver, same skill, same roads. A completely different set of pedals under their feet.

The driver did not get worse. The conditions they were tested under stopped matching the conditions they trained under. That gap is training-serving skew, aimed at machine learning.

Why it exists

A model does not see raw events. It sees features — numbers computed from those events, like "days since signup" or "average order value".

Two different pieces of code usually build those numbers. One runs offline, in a batch job, written by whoever trained the model. Another runs online, inside the live service, often written later, by someone else, in a different language.

Nobody plans for these two to disagree. They are supposed to compute the exact same thing. In practice, a rounding rule, a time zone, or a unit gets handled differently in each, and the two quietly drift apart.

How it works

        the SAME real event
          /              \
   training code       serving code
   (batch job)          (live service)
         |                  |
   account age: 30 days   account age: 30000 days
         |                  |
   the model was never shown a feature value like this

The model itself did not change between training and serving. The number it was handed for the exact same real event did.

A real example you have seen

A recommendation system trained offline computes "how many items you viewed today" using a clean daily log file, counted once a day.

The live service counts the same thing from a real-time event stream, which can double-count a retried request. The live number for a real, ordinary user can end up larger than any number the model ever saw in training.

The honest part

This is one of the hardest bugs in production machine learning to catch by reading code. Both pipelines can look correct in isolation.

The only reliable way to catch it: compare real feature values from both pipelines, side by side, on a schedule. Reading either pipeline alone will not reveal it.

Remember this

  • Training-serving skew is two different pieces of code disagreeing about the same feature.
  • It hides even when both pipelines look correct on their own.
  • Catch it by comparing real feature values from both pipelines, not by reading code.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install pandas numpy scikit-learn

A believable version of the bug

A training pipeline computes account age correctly from proper timestamps. A serving pipeline receives the same event as milliseconds, and divides by a divisor that was correct for a different upstream source.

skew.py
import pandas as pd
import numpy as np
from sklearn.linear_model import LogisticRegression

rng = np.random.RandomState(0)

# ---- TRAINING PIPELINE (offline, batch, built by the data team) ----
# signup_time and event_time are proper pandas Timestamps in a warehouse table.
signup = pd.Timestamp("2026-01-01")
event_times = pd.to_datetime("2026-01-01") + pd.to_timedelta(
    rng.randint(1, 400, 2000), unit="D"
)


def account_age_days_training(event_time, signup_time):
    return (event_time - signup_time).days


train_age = [account_age_days_training(t, signup) for t in event_times]
# A believable rule: older accounts churn less.
y = (np.array(train_age) < 60).astype(int)

model = LogisticRegression().fit(np.array(train_age).reshape(-1, 1), y)
print("train accuracy:", round(model.score(np.array(train_age).reshape(-1, 1), y), 3))

# ---- SERVING PIPELINE (a second engineer, a second codebase, months later) ----
# The event arrives from a different system as epoch MILLISECONDS, not seconds.
signup_epoch_ms = int(signup.timestamp() * 1000)


def account_age_days_serving(event_epoch_ms, signup_epoch_ms):
    # Written for an older upstream event that sent epoch SECONDS.
    # This one sends epoch MILLISECONDS. The divisor was never updated.
    return (event_epoch_ms - signup_epoch_ms) / 86400


sample_event = signup + pd.Timedelta(days=30)  # a genuinely young account
sample_epoch_ms = int(sample_event.timestamp() * 1000)

correct = account_age_days_training(sample_event, signup)
skewed = account_age_days_serving(sample_epoch_ms, signup_epoch_ms)

print(f"same real event -- training feature: {correct} days")
print(f"same real event -- serving feature:  {skewed:.1f} days  (off by 1000x)")

pred_correct = model.predict([[correct]])[0]
pred_skewed = model.predict([[skewed]])[0]
print(f"prediction using the training-shaped feature: {pred_correct}")
print(f"prediction using the serving-shaped feature:   {pred_skewed}")
Output
train accuracy: 1.0
same real event -- training feature: 30 days
same real event -- serving feature:  30000.0 days  (off by 1000x)
prediction using the training-shaped feature: 1
prediction using the serving-shaped feature:   0

Every number above is exact, deterministic output from this script. A real production skew is rarely a clean 1000x. It is usually a few percent, which is why it survives so long unnoticed.

Line-by-line walkthrough

account_age_days_training uses proper pandas.Timestamp subtraction, which handles units correctly by construction.

account_age_days_serving reimplements the same idea by hand, in raw arithmetic, against a differently-shaped input. That reimplementation is the entire bug.

The model itself is never touched. It was trained once, correctly, and is applied unchanged to both feature values. The prediction flips only because the number reaching it flipped.

Common mistakes

Writing the serving feature code from memory, instead of sharing the training code. The single biggest fix available is making one function do both jobs — call the same account_age_days_training from both the batch job and the live service, and this entire bug class becomes impossible.

Never comparing real feature values between the two pipelines. Log the serving-time feature values alongside predictions, as covered in logging every prediction, and periodically diff a sample against a freshly recomputed training-side value for the same real events.

Testing only with hand-written example inputs. Hand-written test events tend to accidentally use round, well-behaved numbers. Test with real logged events, which contain the messy edge cases that reveal unit mismatches.

Assuming a passing unit test on each pipeline separately proves they agree. Both functions can have perfect unit tests and still disagree, because a unit test author usually assumes the same units the buggy code assumes.

Try it yourself

Add a third event exactly 60 days after signup, the boundary the model's rule flips on. Compute both feature versions for it, and check whether the two pipelines' predictions agree at that boundary too.

What to learn next

Researcher — Mathematics and papers.

Where skew enters a real feature pipeline

Sculley et al. (2015) name this among the "hidden technical debt" categories that make ML systems costly to maintain, under the broader heading of CACE (Changing Anything Changes Everything) and pipeline entanglement. Common concrete sources:

  • Language or runtime mismatch — a Python/pandas offline pipeline reimplemented in Java or Go for low-latency serving, by a different team.
  • Windowing mismatch — an offline aggregate computed over a clean, closed daily partition, against an online aggregate computed over an open, still-arriving stream.
  • Null-handling mismatch — offline code that drops missing values silently, online code that must return a value for every request and substitutes a default.
  • Point-in-time leakage in reverse — offline features accidentally allowed to see data that would not yet exist at serving time, producing training accuracy that serving can never match.

Detecting it without a labelled outcome

Because skew is a feature-value disagreement, it does not require waiting for labels the way concept drift does. The direct method: log a sample of production requests with their raw inputs, recompute features for those same inputs through the offline pipeline, and diff the two feature vectors.

$$\delta_j = \frac{1}{n}\sum_{i=1}^{n} \mathbb{1}\left[\,|f^{\text{train}}_j(x_i) - f^{\text{serve}}_j(x_i)| > \tau_j\,\right]$$

Where $f^{\text{train}}_j$ and $f^{\text{serve}}_j$ are the training-time and serving-time implementations of feature $j$, $\tau_j$ is a per-feature tolerance (zero for exact-match categorical features, an epsilon for floating point), and $\delta_j$ is the fraction of sampled requests where the two disagree beyond that tolerance. TensorFlow Data Validation's training-serving skew detector (Breck et al., 2019) implements a variant of this comparison as a first-class pipeline stage.

The structural fix

The design that eliminates this bug class rather than detecting it after the fact: a shared feature-transformation layer, compiled or interpreted identically in both the batch and serving paths, so $f^{\text{train}}_j \equiv f^{\text{serve}}_j$ by construction rather than by convention. Feature stores (Feast, Tecton) and frameworks like TFX's tf.Transform exist largely to guarantee this identity.

Papers

  • Sculley et al., Hidden Technical Debt in Machine Learning Systems, NeurIPS 2015 — papers.nips.cc/paper/5656
  • Breck, Polyzotis, Roy, Whang and Zinkevich, Data Validation for Machine Learning, SysML 2019
  • Baylor et al., TFX: A TensorFlow-Based Production-Scale Machine Learning Platform, KDD 2017

What to learn next