Incident Response for ML Systems

When upstream data breaks

Most ML incidents are not the model's fault — they are a change in the data arriving from somewhere else, with nobody who touched your code involved.

On this page 9
  1. The short answer
  2. The analogy you have already lived
  3. Why it exists
  4. How it works
  5. The three usual shapes
  6. Where you have already felt this
  7. The honest part
  8. Remember this
  9. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

Upstream data breakage means the data feeding your model changed shape or meaning, without anyone touching your code.

The analogy you have already lived

Imagine your vegetable vendor has billed you onion prices per kilogram, for years. One day, with no warning, the same number on the invoice means price per piece instead.

Nothing on the paper says the meaning changed. The number still looks like a normal price. Your accounts, built on the old meaning, are now quietly wrong. It was never your mistake, and you cannot catch it by looking at your own books.

Why it exists

Your model does not live alone. It reads data some other team, company, or pipeline produces. You do not control when they change it.

They rename a column because it made more sense to them. They switch currencies. They fix what looked like a bug on their end — a value your feature quietly depended on. None of them know your model exists.

Your code did not change. The ground it stands on did.

How it works

Upstream team changes something
  (renames a column, changes units, starts sending nulls)
              |
              v
    Your pipeline reads it as if nothing changed
              |
              v
    Feature values are wrong, or missing, or the wrong scale
              |
              v
    The model scores confidently on nonsense inputs
              |
              v
    Nobody upstream knows. Nobody downstream is told.

This is the same shape as silent failures, with one difference worth naming. The bug did not originate in your code at all — it arrived from outside, fully formed.

The three usual shapes

A renamed or dropped column. The field your pipeline expects is not there any more, under that name.

A changed unit or scale. Rupees where thousands of rupees used to arrive. Seconds where milliseconds used to arrive. The type is right; the meaning is not.

A flood of missing values. An upstream job failed partway through. Instead of erroring, it sent nulls for the fields it could not compute.

Where you have already felt this

A weather app suddenly showing an impossible temperature because a sensor started reporting in a different unit. A price comparison site briefly showing a laptop at nine rupees because a retailer's feed dropped three zeros.

The honest part

You cannot stop another team from changing their data. What you can do is stop trusting it blindly. Check the shape of every arrival against what you agreed on, before it reaches your model.

Remember this

  • Upstream breakage means the data changed, not your code — and you rarely get advance warning.
  • The three usual shapes: a renamed column, a changed unit, or a flood of nulls.
  • Check incoming data at the door, every time, instead of trusting it arrived the way you expect.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install pandas

Checking a batch before it reaches the model

validate_batch.py
import pandas as pd

SCHEMA = {
    "income": {"dtype": "float", "min": 0, "max": 200},   # thousands per month
    "years":  {"dtype": "float", "min": 0, "max": 60},
    "age":    {"dtype": "float", "min": 18, "max": 100},
}
MAX_NULL_FRACTION = 0.02


def validate_batch(df: pd.DataFrame, schema: dict = SCHEMA) -> list[str]:
    """Returns a list of problems. An empty list means the batch is clean."""
    problems = []

    missing = [c for c in schema if c not in df.columns]
    if missing:
        problems.append(f"missing column(s): {missing}")
        return problems  # nothing else below is checkable without the columns

    for col, rules in schema.items():
        series = df[col]

        null_frac = series.isna().mean()
        if null_frac > MAX_NULL_FRACTION:
            problems.append(f"'{col}': {null_frac:.0%} missing, over the {MAX_NULL_FRACTION:.0%} limit")

        numeric = pd.to_numeric(series, errors="coerce")
        if numeric.isna().sum() > series.isna().sum():
            problems.append(f"'{col}': contains values that are not numbers")
            continue

        below = (numeric < rules["min"]).sum()
        above = (numeric > rules["max"]).sum()
        if below:
            problems.append(f"'{col}': {below} value(s) below {rules['min']}")
        if above:
            problems.append(f"'{col}': {above} value(s) above {rules['max']} "
                             f"(median is {numeric.median():.1f}, check for a unit change)")
    return problems


good = pd.DataFrame({"income": [60.0, 45.0, 30.0], "years": [7.0, 3.0, 1.5], "age": [34, 29, 22]})

# Upstream renamed a column without telling anyone.
renamed = pd.DataFrame({"monthly_income": [60.0, 45.0, 30.0], "years": [7.0, 3.0, 1.5], "age": [34, 29, 22]})

# Upstream started sending income in rupees, not thousands of rupees.
# Same column name, same type, values quietly 1000x bigger.
unit_change = pd.DataFrame({"income": [60000.0, 45000.0, 30000.0], "years": [7.0, 3.0, 1.5], "age": [34, 29, 22]})

for label, batch in [("Batch 1 (good)", good), ("Batch 2 (renamed column)", renamed),
                     ("Batch 3 (unit change)", unit_change)]:
    problems = validate_batch(batch)
    print(f"{label}:")
    if problems:
        for p in problems:
            print(f"  REJECTED - {p}")
    else:
        print("  accepted, forwarded to the model")
    print()
Output
Batch 1 (good):
  accepted, forwarded to the model

Batch 2 (renamed column):
  REJECTED - missing column(s): ['income']

Batch 3 (unit change):
  REJECTED - 'income': 3 value(s) above 200 (median is 45000.0, check for a unit change)

Notice validate_batch never had to know why the column was missing or the values were large. It only had to know what was agreed — the schema — and compare. That comparison is cheap, and it catches both bugs before either reaches the model.

Quarantine instead of discard

A rejected batch that vanishes teaches nobody anything. Logging it lets someone review the pattern later, and go talk to the team that owns the feed.

quarantine.py
import sqlite3
from datetime import datetime

conn = sqlite3.connect(":memory:")
conn.execute("""
    CREATE TABLE quarantine (
        received_at TEXT, source TEXT, row_count INTEGER, problems TEXT
    )
""")


def quarantine_batch(source: str, row_count: int, problems: list[str]) -> None:
    conn.execute(
        "INSERT INTO quarantine VALUES (?, ?, ?, ?)",
        (datetime.now().isoformat(timespec="seconds"), source, row_count, "; ".join(problems)),
    )
    conn.commit()


quarantine_batch("partner-feed-v2", row_count=3, problems=["missing column(s): ['income']"])
quarantine_batch("partner-feed-v2", row_count=3,
                  problems=["'income': 3 value(s) above 200 (median is 45000.0, check for a unit change)"])

print("Quarantine log:")
for row in conn.execute("SELECT received_at, source, row_count, problems FROM quarantine"):
    print(f"  [{row[0]}] {row[1]} ({row[2]} rows): {row[3]}")

count = conn.execute("SELECT COUNT(*) FROM quarantine WHERE source = ?", ("partner-feed-v2",)).fetchone()[0]
print(f"\n'partner-feed-v2' has {count} rejected batches on record, worth a message to that team.")
Output
Quarantine log:
  [2026-09-02T12:09:19] partner-feed-v2 (3 rows): missing column(s): ['income']
  [2026-09-02T12:09:19] partner-feed-v2 (3 rows): 'income': 3 value(s) above 200 (median is 45000.0, check for a unit change)

'partner-feed-v2' has 2 rejected batches on record, worth a message to that team.

The timestamp in your run will show the moment you ran it, not 12:09:19 — datetime.now() reads the real clock, which is the point: a quarantine log is only useful because it is honest about exactly when each bad batch arrived.

Common mistakes

Trusting a partner's data because it usually arrives correctly. "Usually" is exactly when a validation gate earns its cost — the day it does not.

Checking types but not ranges. renamed's bug would slip past a check that only confirms columns are numbers. A unit change like unit_change keeps the right type; only the range check catches it.

Fixing the batch instead of rejecting it. Guessing that monthly_income means the same as income and silently renaming it is the silent-failure mistake wearing a data-pipeline costume. Reject and ask, rather than guess and hope.

No owner notified. A quarantine table nobody reads is a graveyard, not a safety net. Alert whoever owns the pipeline the moment a batch lands there.

Try it yourself

Add a fourth synthetic batch where age arrives as text like "34 years" instead of a number, and confirm validate_batch catches it through the "contains values that are not numbers" branch rather than crashing.

What to learn next

Researcher — Mathematics and papers.

Schema validation as a contract

Formalising incoming data with a schema turns an implicit assumption — "the upstream data looks like this" — into an explicit, checkable contract. This is the same idea CI/CD's data tests apply to a model's own training data, applied here at the ingestion boundary instead, where it catches a breaking change before a single row reaches a model.

Breck et al., Data Validation for Machine Learning (SysML 2019), formalise three distinct checks that generalise the example above:

  • Single-batch validation — does this batch alone satisfy the schema (types, domains, required-ness, ranges)? validate_batch implements this.
  • Inter-batch validation — does this batch's distribution differ materially from recent history, even if each value individually passes the schema? A unit change with values still inside the declared range (rupees per month instead of thousands, both technically "numbers") would pass validate_batch and only be caught here.
  • Training-serving skew detection — do the feature distributions seen at training time and at serving time agree? Relevant when the same upstream source feeds both.

The gap between what validate_batch catches and full inter-batch validation is real: a unit change that stays within the declared numeric range, without crossing any hard boundary, needs a distributional comparison against recent history to catch, not a fixed schema alone.

Detecting a distributional shift without a fixed range

A schema with fixed min/max values, as used above, misses a shift that keeps every value inside the declared range but changes their distribution — for instance, a currency conversion that happens to leave values numerically plausible. Population Stability Index (PSI) and the Kolmogorov-Smirnov statistic compare a current batch's distribution against a reference window and flag a shift regardless of whether any single value crosses a hard limit; this is the mechanism monitoring and model drift applies on an ongoing schedule, and it is the correct complement to a fixed-range check at ingestion time.

Contract testing between teams

Where the upstream source is itself another internal team, consumer-driven contract testing (Pact and similar frameworks, originating in microservice API testing) formalises the schema as a contract both sides run in their own CI — the consumer publishes the shape it expects, and the producer's pipeline fails its own tests if a change would violate it. This shifts detection from "the incident, in production" to "a failed build, before the change ships" — the same shift CI/CD for machine learning makes for a model's own code.

Why this class of failure dominates ML incident volume

Multiple postmortem studies of production ML incidents (informally corroborated across large tech companies' internal incident reviews, though rarely published in detail due to their operational sensitivity) attribute a majority of ML-specific incidents to data issues rather than model or infrastructure code — consistent with Sculley et al.'s (2015) broader observation that ML systems carry disproportionate technical debt at their data boundaries relative to their modelling code. This is the practical argument for investing in ingestion-time validation before investing further in model sophistication.

What to learn next