Incident Response for ML Systems

Silent failures

A silent failure is a model that answers every request, with no error, while being wrong — because nothing about being wrong trips a health check.

On this page 9
  1. The short answer
  2. The analogy you have already lived
  3. Why it exists
  4. How it works
  5. The usual causes
  6. Where you have already seen the aftermath
  7. The honest part
  8. Remember this
  9. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

A silent failure is a model that keeps answering every request, with no error, while being wrong.

The analogy you have already lived

A shopkeeper's weighing scale can drift by two hundred grams and still show a confident number on the dial. It does not beep. It does not turn red. You find out only when you get home and the bag feels lighter than it should.

A model with a silent failure is that scale. It answers every request. Every answer looks reasonable. Nothing about the system tells you it is wrong.

Why it exists

Every check you have met so far — health checks, error rates, latency alerts — catches a loud failure. The server crashed. A request timed out. Those are easy to notice, because something visibly stopped working.

A silent failure is different in kind, not degree. The server is healthy. Every response has the right shape, the right fields, a plausible-looking number. The runbook from the previous lesson would pass every step. The triage signals would all read green.

The model is answering a question it was never trained to answer, and saying nothing about it.

How it works

Training time:
  the model learns about cities:  Mumbai, Delhi, Bengaluru

Months later, in production:
  a new city arrives:  "Pune"
       |
       v
  code written to be forgiving: unknown city -> treat as Mumbai
       |
       v
  every Pune request scored as if it were a Mumbai request
       |
       v
  HTTP 200. Valid-looking number. No error anywhere.

The bug is not the missing city. New cities are normal. The bug is the silent fallback — treating "I do not know" as "I will guess, and tell nobody."

The usual causes

A forgiving default. Code that maps an unrecognised value to something safe-looking instead of raising. Kind in ordinary software; dangerous in a model, because "safe-looking" and "correct" are not the same thing.

A stale model file. A deploy reports success, but the server still serves yesterday's model. A healthy process gives no hint of which model it loaded.

A cache serving old answers. A prediction cached an hour ago, returned again after the inputs — and the correct answer — have changed.

Where you have already seen the aftermath

A loan approved for someone the model was never trained to evaluate, discovered months later in a review. A translation that is confidently, fluently wrong. A recommendation engine still suggesting a product that went out of stock last week.

The honest part

You cannot watch for "wrong" directly. That needs the true outcome, which often arrives too late to matter. What you can watch is the shape of the model's inputs and outputs instead. A spike in values falling back to a default. An output distribution that suddenly looks unlike every day before it. Neither is proof of a silent failure. Each is the closest signal you get without waiting for the true answer.

Remember this

  • A silent failure returns HTTP 200 and a plausible number, while being wrong.
  • It usually comes from a forgiving default somewhere in the code, quietly guessing instead of asking for help.
  • Watch the shape of inputs and outputs. A spike in defaults is your earliest warning.

What to learn next

Developer — Code and libraries.

Setup

None. Pure Python standard library.

A forgiving default, quietly wrong

silent_bug.py
# Learned at training time. Three cities the model has ever seen.
CITY_INDEX = {"MUM": 0, "DEL": 1, "BLR": 2}


def encode_naive(city: str) -> int:
    """The common, quiet bug: an unknown value falls back to index 0
    instead of raising. Every unknown city silently becomes Mumbai."""
    return CITY_INDEX.get(city, 0)


def encode_safe(city: str) -> int | None:
    """Same encoder, but an unknown value is reported, not guessed."""
    return CITY_INDEX.get(city)  # None when the key is missing


def predict_naive(city: str, amount: float) -> float:
    city_idx = encode_naive(city)
    # Mumbai (index 0) has the lowest base fraud risk in this training
    # data, so misreading a city as Mumbai quietly under-scores its risk.
    base_risk = [0.10, 0.18, 0.15][city_idx]
    return round(min(0.99, base_risk + amount / 200_000), 4)


DAY1 = [("MUM", 4000), ("DEL", 9000), ("BLR", 3000), ("MUM", 12000), ("DEL", 2000)]

# A new dark store opened in Pune. Nobody told the model — "PUN" now
# arrives in the same feature, unrecognised.
DAY2 = [("MUM", 4000), ("PUN", 9000), ("BLR", 3000), ("PUN", 12000), ("PUN", 2000)]


def default_rate(rows: list[tuple[str, float]]) -> float:
    unknown = sum(1 for city, _ in rows if encode_safe(city) is None)
    return unknown / len(rows)


for label, rows in [("Day 1", DAY1), ("Day 2", DAY2)]:
    scores = [predict_naive(city, amt) for city, amt in rows]
    print(f"{label}: scores = {scores}")
    print(f"{label}: default_rate (fraction treated as unknown) = {default_rate(rows):.0%}\n")

print("Health check status both days: 200 OK, model_loaded=True, error_rate=0.0%")
print("Nothing paged. The fraud model has been under-scoring Pune since Day 2.")
Output
Day 1: scores = [0.12, 0.225, 0.165, 0.16, 0.19]
Day 1: default_rate (fraction treated as unknown) = 0%

Day 2: scores = [0.12, 0.145, 0.165, 0.16, 0.11]
Day 2: default_rate (fraction treated as unknown) = 60%

Health check status both days: 200 OK, model_loaded=True, error_rate=0.0%
Nothing paged. The fraud model has been under-scoring Pune since Day 2.

Look at the Pune rows on Day 2. 0.145, 0.16, 0.11 — these look like ordinary fraud scores. Nothing about a 0.145 on its own looks wrong. It is only wrong compared with what a correctly-recognised Pune request should have scored, which this code has no way to know, because it was never taught Pune existed.

Turning the one number you can see into an alert

You cannot watch "wrong" directly. You can watch how often the encoder falls back to a default — the closest signal you get without a true outcome.

silent_failure_alert.py
from silent_bug import DAY1, DAY2, default_rate


def silent_failure_alert(rows: list[tuple[str, float]], threshold: float = 0.20):
    dr = default_rate(rows)
    return ("PAGE_NOW" if dr > threshold else "OK"), dr


for label, rows in [("Day 1", DAY1), ("Day 2", DAY2)]:
    decision, dr = silent_failure_alert(rows)
    print(f"{label}: unknown-category rate {dr:.0%} -> {decision}")
Output
Day 1: unknown-category rate 0% -> OK
Day 2: unknown-category rate 60% -> PAGE_NOW

This is a small version of the same shape as on-call for ML systems's classify_alert — a threshold on one measurable number. The 20% threshold is a choice, not a law; set it low enough to catch a real shift, high enough that a handful of genuinely new, rare categories does not page someone every day.

Common mistakes

Silently defaulting unknown values "to be safe". It is the opposite of safe. encode_naive looks defensive and is actually the bug. Prefer encode_safe's None — an explicit "I do not know" a caller must handle.

Measuring the default rate only in aggregate. A 2% default rate site-wide can be 60% for one city and 0% everywhere else, exactly as in the example. Break the metric down by the input that is failing, not only the overall number.

Assuming a green health check means a correct model. It means the process is alive. Runbooks built only from health and latency checks will never catch this class of failure — they need a schema and default-rate check added deliberately.

No log of what the fallback value actually was. When encode_naive returns 0 for "PUN", log the raw unrecognised value. Without it, nobody can tell if the fallback is happening for one new city or fifty different typos.

Try it yourself

Change encode_naive to raise a ValueError for any city not in CITY_INDEX, and run predict_naive over DAY2. Watch the loud failure this produces, then decide honestly which you would rather debug at 2 a.m. — five clear errors, or a fraud model quietly wrong for an unknown fraction of one city's traffic.

What to learn next

Researcher — Mathematics and papers.

Silent failure as an unobserved-shift problem

Formally, this is covariate shift with no fallback coverage: the serving-time input distribution $P_{\text{serve}}(x)$ contains support outside the training distribution's domain, $x \notin \text{dom}(P_{\text{train}})$, and the encoding function is not a partial function that fails on that region but a total function that silently projects it onto an existing category. The failure is not that the model extrapolates poorly outside its training domain — every model does that to some degree — it is that the system gives no signal that extrapolation occurred at all.

Why aggregate metrics under-power detection

If a genuinely new category affects a small, systematically-clustered subpopulation (one city, one device type, one partner integration), an aggregate accuracy or drift metric computed over all traffic dilutes the signal by the inverse of that subpopulation's traffic share. Detecting it requires segment-level monitoring — computing the same drift or default-rate statistic within slices of traffic rather than only in aggregate — the approach the site's monitoring and model drift lesson introduces and this lesson's per-category default_rate applies at the smallest possible scale, one feature.

Detecting an unknown-category spike statistically

Treat the daily count of fallback-encoded requests as a count process and compare it against a historical baseline using a statistical process control approach: a control chart with limits set at the baseline mean plus some multiple of its standard deviation, flagging any day that exceeds them. This is the same family of technique as CUSUM (cumulative sum control chart, Page, 1954) used for detecting a shift in a process mean as early as possible after it occurs, applied here to "rate of encoder fallback" instead of a physical manufacturing measurement.

Prevention over detection

Breck et al., Data Validation for Machine Learning (SysML 2019) — also cited in CI/CD for machine learning — argues for a checked-in schema with an explicit, versioned domain for every categorical feature, and treats any value outside that domain at serving time as a schema violation to reject or quarantine, not a value to encode. Under that discipline encode_naive's behaviour is not a possible design choice; it is a validated-away class of bug, caught before a request ever reaches the model.

Sculley et al., Hidden Technical Debt in Machine Learning Systems (NeurIPS 2015), names this general shape "unstable data dependencies" and "correction cascades" — a downstream component silently compensating for an upstream change rather than surfacing it. The Pune example above is a correction cascade of exactly one line: .get(city, 0).

What to learn next