Incident Response for ML Systems

Triaging an ML incident

Triage is deciding how bad an incident is and what to do first, in the first few minutes, before anyone understands the full cause.

On this page 9
  1. The short answer
  2. The analogy you have already lived
  3. Why it exists
  4. How it works
  5. The fastest question in any incident
  6. Where you have already benefited from this
  7. The honest part
  8. Remember this
  9. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

Triage means deciding how bad an incident is, and what to do first, before you understand the full cause.

The analogy you have already lived

You have sat in a crowded hospital waiting room. Before any doctor sees you, a nurse asks a few quick questions — since when, how bad, are you bleeding. She decides who goes in first.

The nurse is not diagnosing you. There is no time for that yet, and ten more people are waiting. The nurse is deciding order and urgency. That is triage — for a model incident, the questions change, but the job is the same.

Why it exists

An incident starts with almost no information. A dashboard turned red. Someone said predictions look wrong.

You could spend twenty minutes understanding the full cause before doing anything. For a small, contained problem, that is fine. For a payments model failing for every user, twenty minutes is twenty minutes too many.

Triage buys time to think, by deciding two things fast: how many people does this hurt, and can it wait?

How it works

Incident noticed
       |
       v
  How much of the traffic is affected?
       |                    |
     a lot                a little
       |                    |
       v                    v
  Is it revenue or       Investigate on a
  safety sensitive?      normal timeline
       |          |
      yes         no
       |          |
       v          v
   SEV1: stop    SEV2/3: fix soon,
   the bleeding  no need to panic
   first, then
   investigate

Two questions do almost all the sorting: how big, and how sensitive. Together they decide "act now" or "act calmly, soon." Everything after that sort is investigation, covered by runbooks when one exists.

The fastest question in any incident

"What changed recently?"

Most incidents start within minutes of a deploy, a config change, or an upstream data update. Check the last thirty minutes of changes before anything else. It finds the cause more often than deep investigation does.

Where you have already benefited from this

An airline gate agent deciding who re-books first when a flight is cancelled. A bank's fraud team deciding which of a thousand alerts a human looks at in the next five minutes. Neither waits to understand everything before deciding what happens next.

The honest part

Triage decisions are made with incomplete information, and some will be wrong. A SEV3 that turns out to be SEV1 is not a personal failure. It is the cost of deciding fast over deciding right. Re-triage the moment new information arrives; do not defend the first guess.

Remember this

  • Triage decides urgency and order, not the cause — that comes after.
  • Two questions sort most incidents: how much of the traffic, and how sensitive is it.
  • Check what changed recently first — it is the fastest lead you have.

What to learn next

Developer — Code and libraries.

Setup

None. Both examples use only the Python standard library.

Sorting incidents by severity

triage.py
def triage(incident: dict) -> tuple[str, str]:
    """Return (severity, first_action) from a handful of incident signals."""
    pct = incident["affected_pct"]
    revenue = incident["revenue_sensitive"]
    trend = incident["trend"]
    workaround = incident["workaround_exists"]

    if pct >= 0.50 or (revenue and pct >= 0.10):
        severity = "SEV1"
    elif pct >= 0.10 or (revenue and trend == "worsening"):
        severity = "SEV2"
    elif pct >= 0.01:
        severity = "SEV3"
    else:
        severity = "SEV4"

    if severity in ("SEV1", "SEV2") and not workaround:
        action = "mitigate first (rollback / feature flag off), investigate after"
    elif severity in ("SEV1", "SEV2") and workaround:
        action = "apply the workaround now, investigate on a normal timeline"
    else:
        action = "investigate on a normal timeline, no immediate mitigation needed"

    return severity, action


INCIDENTS = [
    {"name": "checkout-fraud-check down",     "affected_pct": 1.00, "revenue_sensitive": True,  "trend": "worsening", "workaround_exists": False},
    {"name": "recsys-home scores stale",      "affected_pct": 0.35, "revenue_sensitive": False, "trend": "stable",    "workaround_exists": False},
    {"name": "loan-scorer p99 up slightly",   "affected_pct": 0.03, "revenue_sensitive": True,  "trend": "stable",    "workaround_exists": True},
    {"name": "churn-model batch job delayed", "affected_pct": 1.00, "revenue_sensitive": False, "trend": "stable",    "workaround_exists": True},
    {"name": "search typo-tolerance off",     "affected_pct": 0.005,"revenue_sensitive": False, "trend": "improving", "workaround_exists": False},
]

print(f"{'incident':<30}{'severity':<10}{'first action'}")
for inc in INCIDENTS:
    sev, action = triage(inc)
    print(f"{inc['name']:<30}{sev:<10}{action}")
Output
incident                      severity  first action
checkout-fraud-check down     SEV1      mitigate first (rollback / feature flag off), investigate after
recsys-home scores stale      SEV2      mitigate first (rollback / feature flag off), investigate after
loan-scorer p99 up slightly   SEV3      investigate on a normal timeline, no immediate mitigation needed
churn-model batch job delayed SEV1      apply the workaround now, investigate on a normal timeline
search typo-tolerance off     SEV4      investigate on a normal timeline, no immediate mitigation needed

Notice checkout-fraud-check down and churn-model batch job delayed are both SEV1 — both affect 100% of their traffic — but their first actions differ. One has no workaround and needs an immediate rollback. The other has a workaround, so the fix can wait for daylight while the workaround absorbs the impact.

What changed recently

recent_change.py
from datetime import datetime, timedelta

LOOKBACK = timedelta(minutes=30)


def find_suspect_changes(alert_time: datetime, changes: list[dict]) -> list[dict]:
    """Changes inside the lookback window before the alert, most recent first."""
    suspects = [c for c in changes if alert_time - LOOKBACK <= c["at"] <= alert_time]
    return sorted(suspects, key=lambda c: c["at"], reverse=True)


ALERT_TIME = datetime(2026, 8, 31, 14, 20)

CHANGES = [
    {"at": datetime(2026, 8, 31, 14, 5),  "what": "deployed loan-scorer v12"},
    {"at": datetime(2026, 8, 31, 11, 0),  "what": "rotated an API key"},
    {"at": datetime(2026, 8, 31, 13, 50), "what": "upstream 'income' feature pipeline re-ran"},
    {"at": datetime(2026, 8, 30, 9, 0),   "what": "deployed loan-scorer v11"},
]

suspects = find_suspect_changes(ALERT_TIME, CHANGES)
lookback_min = int(LOOKBACK.total_seconds() // 60)
print(f"Alert fired at {ALERT_TIME:%H:%M}. Changes in the {lookback_min} minutes before it:\n")
if not suspects:
    print("None found — this is not a 'something recently changed' incident.")
for c in suspects:
    minutes_before = int((ALERT_TIME - c["at"]).total_seconds() // 60)
    print(f"  {c['at']:%H:%M}  ({minutes_before:>2} min before)  {c['what']}")
Output
Alert fired at 14:20. Changes in the 30 minutes before it:

  14:05  (15 min before)  deployed loan-scorer v12
  13:50  (30 min before)  upstream 'income' feature pipeline re-ran

Two suspects, not one. find_suspect_changes does not decide which caused the incident — it narrows a search that would otherwise cover the entire deploy history down to two events worth checking first. The API key rotation at 11:00, well outside the window, is correctly excluded.

Common mistakes

Investigating before mitigating on a SEV1. Understanding the exact cause feels more satisfying than a blunt rollback. On a severe, revenue-sensitive incident, roll back or fail over first. Understand it after the bleeding stops.

Widening the lookback window until something matches. If nothing appears in 30 minutes, resist stretching it to six hours to find a culprit. Say plainly that no recent change explains it, and look elsewhere — most likely upstream data or a silent failure.

Triaging once and never revisiting. An incident triaged as SEV3 at minute one can become SEV1 by minute ten if it is spreading. Re-run triage as new numbers arrive.

Confusing "severe" with "hard to fix". A one-line config revert can resolve a SEV1. A confusing, low-impact bug can take a day to understand. Severity is about user impact, not difficulty.

Try it yourself

Add a previously_seen: bool field to an incident, true if this exact failure has happened before. Adjust triage so a previously-seen incident with a known fix skips straight to "apply the known fix", instead of "mitigate first" — you already know what works.

What to learn next

Researcher — Mathematics and papers.

Incident severity as a decision problem

Triage is a bounded-time classification problem: sort $n$ unresolved signals into $k$ priority classes using partial information, under a cost function where delaying a high-severity incident is far more expensive than a slightly wrong classification of a low one. This asymmetry is why the triage function above is deliberately conservative — it rounds ambiguous cases toward a higher severity ($\text{pct} \geq 0.10$ triggers SEV2 even without revenue sensitivity) rather than toward a lower one, because the expected cost of under-triaging exceeds the cost of over-triaging.

Change-point correlation

find_suspect_changes is a simple time-window filter. A more rigorous version treats the incident's start time as an unknown change point and asks which prior events most plausibly precede it, weighting by both proximity in time and known blast radius of the change type (a full model deploy outranks a documentation update). Formally, this is related to root cause analysis via event correlation in distributed systems, surveyed in Chen et al., Pinpoint: Problem Determination in Large, Dynamic Internet Services (DSN 2002) — correlating a failure's onset against a log of prior state changes, ranked by temporal proximity and historical association with past incidents.

MTTD, MTTM and MTTR as separate clocks

Incident response distinguishes three intervals, each optimised differently:

  • MTTD (mean time to detect) — from the fault occurring to it being noticed. Lowered by better monitoring, covered in monitoring and model drift.
  • MTTM (mean time to mitigate) — from detection to user impact stopping, whether or not the cause is understood. This is what triage optimises directly.
  • MTTR (mean time to resolve/repair) — from detection to the underlying cause being fixed.

Separating MTTM from MTTR is the formal justification for "mitigate first, investigate after": a rollback can drive MTTM to near zero while MTTR remains hours away, and for a user-facing incident, MTTM is very often the number that matters.

Severity taxonomies in practice

Most large operators (Google, AWS, PagerDuty's public incident-response guides) converge on a 3–5 level severity scale similar to the SEV1–SEV4 scheme used above, with SEV1 typically requiring: an incident commander assigned, a communication channel opened, and executive visibility, regardless of how quickly the technical fix arrives. The organisational machinery triggered by a severity label is often a larger design decision than the technical threshold itself — a SEV1 that pages five extra people has a real cost distinct from the incident it responds to.

What to learn next

What to learn next

These follow on from what you just read.

  • Incident Response for ML Systems

    Silent failures

    A silent failure is a model that answers every request, with no error, while being wrong — because nothing about being wrong trips a health check.

  • Incident Response for ML Systems

    Debugging a latency spike

    A latency spike is usually a small share of requests taking far longer than the rest, hidden by an average that looks fine — finding which stage they get stuck in is the whole job.

  • Incident Response for ML Systems

    When upstream data breaks

    Most ML incidents are not the model's fault — they are a change in the data arriving from somewhere else, with nobody who touched your code involved.