ML Code That Survives

Reproducing a failure you only see in production

A model that scores well offline and behaves badly live is almost never a broken model — it is the same inputs arriving in a different shape, and one comparison of the two feature paths finds it in minutes.

On this page 5
  1. Why it exists
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

When a model works offline and misbehaves live, the model is usually fine — the two sides arrange the same information differently.

A tailor takes your measurements and writes them on a slip: chest, waist, sleeve. Weeks later a different tailor reads the slip and stitches a shirt using their own order: waist, chest, sleeve.

Every number on the slip is correct. Nobody wrote anything down wrong. The shirt is unwearable.

That is what almost every "the model is broken in production" ticket turns out to be. The training code and the serving code both build a row of numbers, and they build it differently.

Why it exists

Offline, features are built once, for a whole table of rows, usually in a notebook or a training script.

Live, features are built one request at a time, in a different program. Often a different language, written months later by a different person.

Two programs, one job, no shared code. When they drift apart, nothing crashes. The model receives a row of the right length, full of numbers of the right kind, and it returns a confident answer that is wrong.

This gap has a name: training-serving skew. It is the single most common production failure in machine learning, and it is invisible to every offline test you have.

How it works

  TRAINING                          SERVING
  builds a whole table              builds one row per request

  [amount, mumbai, nagpur, pune]    [amount, mumbai, pune, nagpur]
                ^      ^                          ^      ^
                +------+--------------------------+------+
                  the last two are swapped

  no error, no crash, no warning.
  the model answers 0.9986 instead of 0.8319.

The reproduction technique is short enough to remember: run both feature builders on the same rows and compare them number by number. Whatever disagrees is your bug, and you found it without touching production.

A real example you have seen

A parcel arrives with the flat number and building number swapped. The address is complete. Every digit is correct. The delivery still fails, and the tracking page says "delivered".

Nobody can debug that by staring at the parcel. You compare the address that was written against the address that was read.

Remember this

  • The model is rarely the bug. The feature-building code on each side usually is.
  • Training-serving skew is silent — right shape, right types, wrong values.
  • Reproduce it by running both builders on the same rows and diffing the output.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install scikit-learn pandas

Verified with scikit-learn 1.7.2, pandas 2.2.3, numpy 1.26.4, Python 3.10, CPU. Runs in about three seconds. Seeded, so your numbers should match.

The bug, and the test that finds it

A fraud model with one numeric column and one category column. The serving code was written by hand, months later, and it lists the cities in a sensible-looking order that is not the training order.

skew_demo.py
import numpy as np, pandas as pd
from sklearn.linear_model import LogisticRegression

rng = np.random.default_rng(0)
n = 500
df = pd.DataFrame({
    "amount": rng.gamma(2.0, 50.0, n).round(2),
    "city": rng.choice(["mumbai", "pune", "nagpur"], n, p=[.6, .3, .1]),
})
df["fraud"] = ((df.amount > 150) ^ (df.city == "nagpur")).astype(int)

# ---------- offline training ----------
Xtr = pd.get_dummies(df[["amount", "city"]], columns=["city"])
TRAIN_COLUMNS = list(Xtr.columns)
model = LogisticRegression(max_iter=1000).fit(Xtr.to_numpy(dtype=float), df.fraud)
print("training column order:", TRAIN_COLUMNS)

request = {"amount": 220.0, "city": "pune"}

# ---------- how the serving code was written ----------
def featurise_serving(req):                      # hand-written, "obvious" order
    return np.array([[req["amount"],
                      req["city"] == "mumbai",
                      req["city"] == "pune",
                      req["city"] == "nagpur"]], dtype=float)

# ---------- how it should be built ----------
def featurise_correct(req):
    row = pd.get_dummies(pd.DataFrame([req]), columns=["city"])
    return row.reindex(columns=TRAIN_COLUMNS, fill_value=0).to_numpy(dtype=float)

print("serving  column order: ['amount', 'city_mumbai', 'city_pune', 'city_nagpur']")
print()
print(f"probability from the serving code:  {model.predict_proba(featurise_serving(request))[0,1]:.4f}")
print(f"probability it should have been:    {model.predict_proba(featurise_correct(request))[0,1]:.4f}")

# ---------- the test that would have caught it ----------
def skew_report(featurise, k=200):
    bad = 0
    for i in range(k):
        req = {"amount": float(df.amount[i]), "city": df.city[i]}
        if not np.allclose(featurise(req), Xtr.to_numpy(dtype=float)[i]):
            bad += 1
    return bad

print()
print(f"rows where the serving featuriser disagrees with training: {skew_report(featurise_serving)} of 200")
print(f"rows where the correct featuriser disagrees with training: {skew_report(featurise_correct)} of 200")
Output
training column order: ['amount', 'city_mumbai', 'city_nagpur', 'city_pune']
serving  column order: ['amount', 'city_mumbai', 'city_pune', 'city_nagpur']

probability from the serving code:  0.9986
probability it should have been:    0.8319

rows where the serving featuriser disagrees with training: 81 of 200
rows where the correct featuriser disagrees with training: 0 of 200

The walkthrough

Two probabilities, 0.9986 and 0.8319, from one model and one request. If the alerting threshold is 0.9, this bug flips the decision. Nothing in the serving code is wrong as code — it reads well, it has no bug a reviewer would circle, and it produces a four-element float array exactly as intended.

pd.get_dummies sorts categories alphabetically: mumbai, nagpur, pune. The human wrote them in frequency order. Both orders are reasonable. Only one matches the model.

Only 81 of 200 rows disagree. The mumbai rows — the majority class — are identical under both orders, because swapping two zeros changes nothing. So the bug is invisible on most traffic and fires on a minority segment. That is why it survived launch and arrives as a strange complaint from one region. Bugs that affect everything get caught in staging.

skew_report is the whole reproduction technique. It needs no production access, no traffic replay and no logs. Take rows you already have, push them through both paths, and compare arrays. Nought disagreements or a bug; there is no third answer.

The fix is reindex(columns=TRAIN_COLUMNS, fill_value=0). Saving the column list with the model turns a positional contract into a named one. Positional contracts between two programs written months apart always break eventually.

Where skew comes from, in rough order of frequency

  1. Two implementations of one featuriser. The case above. The structural fix is to ship one function, imported by both sides, rather than two that agree by inspection.
  2. Statistics recomputed at serve time. A scaler fitted on the request batch instead of loaded from training. Live means and standard deviations drift by the hour.
  3. A category the model never saw. A new city arrives, get_dummies produces a column that does not exist, and naive code either crashes or silently drops it.
  4. Type and precision differences. A float64 pipeline served as float32, or an integer column arriving as a string of digits. See dtype mismatch errors.
  5. Time-dependent features computed at the wrong moment. "Transactions in the last 30 days" measured from the request time offline and from midnight online.
  6. Missing values handled differently. Training filled with the median; serving sends the literal string "null" or a zero.

The debugging order that works

  1. Log the exact feature vector the model received in production, for a request known to be wrong. One row is enough. Without this, everything else is guesswork.
  2. Rebuild that same request offline with the training featuriser and diff the two vectors. The index that differs names the bug.
  3. If the vectors match, the model is genuinely the problem — and it is now a data-distribution question, not a code question. That is monitoring and drift, a different investigation.
  4. Write the skew check as a test so this specific bug cannot come back. Then add it to CI, since the featuriser will be edited again.

Two things make step 1 possible, and both must be built before the incident: log the feature vector alongside the prediction, and attach a version identifier for both the model and the featuriser. Reconstructing a feature vector after the fact turns a one-hour investigation into a one-week one — raw request logs, and a featuriser that has changed twice since.

Common mistakes

Retraining the model as the first response. It changes the one component that was working and destroys the evidence. It often appears to help briefly, because the new model is less sensitive to the corrupted column.

Testing the serving path only with fresh requests. Fresh requests have no ground truth, so you cannot tell a wrong answer from a surprising one. Replay rows whose correct feature vector you already have.

Comparing predictions instead of features. Two nearby probabilities can come from very different inputs, and a matching probability does not prove matching inputs. Compare the vectors — they are exact.

Fixing serving to match training without a test. The next edit to either side reopens the bug. The test is the deliverable; the fix is a consequence of it.

Assuming a crash is the worst case. The ValueError raised when scikit-learn checks feature names is the lucky outcome. Converting to a bare NumPy array to make that error go away is how a loud failure gets turned into a silent one.

Try it yourself

Add a fourth city, "thane", to the serving requests only, and run both featurisers on it. Watch what each one does with a category the model has never seen — one produces a wrong-width array, the other silently sends all-zero city flags. Then decide which behaviour you want in production, and write the assertion that enforces it.

What to learn next

Researcher — Mathematics and papers.

Skew as a systems property, not a bug class

Sculley et al. (2015), Hidden Technical Debt in Machine Learning Systems (NeurIPS), identify training-serving skew as a consequence of two structural properties rather than of individual carelessness. First, ML systems have data dependencies with no compiler, type system or linker to check them — a column reordering is not expressible as a type error, so no static tool can catch it. Second, the training and serving paths are typically implemented in different runtimes for legitimate reasons (batch Python for training, low-latency Java or Go for serving), which makes duplicated featurisation logic the default architecture and skew its expected failure mode.

Breck et al. (2017), The ML Test Score (IEEE Big Data), make this concrete in their rubric: "training and serving features compute the same values" is a named test, and "the model is tested via a canary process before it enters production serving" is its runtime complement. Their field data reports these among the least implemented items in surveyed teams, which is consistent with skew being simultaneously the commonest incident cause and the rarest thing tested for.

The architectural response with the strongest track record is to eliminate the duplication rather than to test around it. A feature store with a shared transformation definition and a guaranteed point-in-time-correct offline read (Uber Michelangelo; Feast; Tecton) makes the two paths executions of one specification. Polyzotis et al. (2018), Data Lifecycle Challenges in Production Machine Learning (SIGMOD Record), analyse the guarantees such systems must supply, of which point-in-time correctness is the subtle one: an offline join that leaks a value recorded after the prediction timestamp reproduces temporal leakage at infrastructure scale.

Detection without ground truth

The observability constraint that shapes production ML debugging is that labels arrive late or never. Any detector usable during an incident must therefore work on inputs and predictions alone.

  • Distributional comparison of live features against the training distribution, per feature, using population stability index, Kolmogorov–Smirnov distance or a discrete divergence. Detects categories that shifted, columns gone constant, and units changed upstream — the input-side view of monitoring and drift.
  • Direct skew assays: replay a sample of logged production requests through the offline featuriser and compare vectors elementwise. This is skew_report at production scale, and unlike distributional tests it localises the defect to a feature index rather than reporting an aggregate anomaly.
  • Shadow or canary serving: run the candidate path on live traffic without acting on its output and compare against the incumbent. It catches skew introduced by the deployment itself, which offline testing cannot observe at all.
  • Prediction-distribution monitoring: the output histogram is a cheap sufficient statistic that moves under most input corruptions. The example above illustrates its limitation — the corruption affected 40% of rows and shifted the affected predictions upward, which a coarse aggregate might absorb entirely.

Ovadia et al. (2019), Can You Trust Your Model's Uncertainty? (NeurIPS), supply the pessimistic result underlying all of these: under distribution shift, model confidence degrades unreliably across every method tested, so a model's own uncertainty is not a dependable skew alarm. Detection must come from comparing artefacts, not from asking the model whether it feels sure.

Reproducibility as the precondition for debugging

The investigative procedure above rests on an assumption worth naming: that a recorded request can be replayed to produce a deterministic feature vector. That holds only if the featuriser is a pure function of the request plus versioned, immutable reference data. Real featurisers routinely violate it — they read a mutable feature table, call the wall clock, or query a service whose state has since changed. Each violation removes one coordinate from the provenance tuple and converts a reproducible defect into an unreproducible report.

The engineering consequences are specific and worth building before an incident, not during one: log the materialised feature vector next to the prediction; version the featuriser and the model together and record both identifiers with each response; and make every time-dependent feature take an explicit as_of timestamp rather than reading the clock. These are the production analogue of the pinning discipline in refactoring a training script — in both cases the artefact that makes the investigation possible must exist before the question is asked.

What to learn next