Incident Response for ML Systems
Triaging an ML incident
Triage is deciding how bad an incident is and what to do first, in the first few minutes, before anyone understands the full cause.
- 9 min read
- 3 reading levels
- Published
Read these first
On this page 9
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
Triage means deciding how bad an incident is, and what to do first, before you understand the full cause.
The analogy you have already lived
You have sat in a crowded hospital waiting room. Before any doctor sees you, a nurse asks a few quick questions — since when, how bad, are you bleeding. She decides who goes in first.
The nurse is not diagnosing you. There is no time for that yet, and ten more people are waiting. The nurse is deciding order and urgency. That is triage — for a model incident, the questions change, but the job is the same.
Why it exists
An incident starts with almost no information. A dashboard turned red. Someone said predictions look wrong.
You could spend twenty minutes understanding the full cause before doing anything. For a small, contained problem, that is fine. For a payments model failing for every user, twenty minutes is twenty minutes too many.
Triage buys time to think, by deciding two things fast: how many people does this hurt, and can it wait?
How it works
Incident noticed
|
v
How much of the traffic is affected?
| |
a lot a little
| |
v v
Is it revenue or Investigate on a
safety sensitive? normal timeline
| |
yes no
| |
v v
SEV1: stop SEV2/3: fix soon,
the bleeding no need to panic
first, then
investigateTwo questions do almost all the sorting: how big, and how sensitive. Together they decide "act now" or "act calmly, soon." Everything after that sort is investigation, covered by runbooks when one exists.
The fastest question in any incident
"What changed recently?"
Most incidents start within minutes of a deploy, a config change, or an upstream data update. Check the last thirty minutes of changes before anything else. It finds the cause more often than deep investigation does.
Where you have already benefited from this
An airline gate agent deciding who re-books first when a flight is cancelled. A bank's fraud team deciding which of a thousand alerts a human looks at in the next five minutes. Neither waits to understand everything before deciding what happens next.
The honest part
Triage decisions are made with incomplete information, and some will be wrong. A SEV3 that turns out to be SEV1 is not a personal failure. It is the cost of deciding fast over deciding right. Re-triage the moment new information arrives; do not defend the first guess.
Remember this
- Triage decides urgency and order, not the cause — that comes after.
- Two questions sort most incidents: how much of the traffic, and how sensitive is it.
- Check what changed recently first — it is the fastest lead you have.
What to learn next
- Silent failures — the failure type that dodges the severity check because nothing looks broken.
- Runbooks for model failures — what to do once triage has pointed at a likely cause.
- Debugging a latency spike — a deep dive into one specific SEV2-shaped symptom.
Developer — Code and libraries.
Setup
None. Both examples use only the Python standard library.
Sorting incidents by severity
def triage(incident: dict) -> tuple[str, str]:
"""Return (severity, first_action) from a handful of incident signals."""
pct = incident["affected_pct"]
revenue = incident["revenue_sensitive"]
trend = incident["trend"]
workaround = incident["workaround_exists"]
if pct >= 0.50 or (revenue and pct >= 0.10):
severity = "SEV1"
elif pct >= 0.10 or (revenue and trend == "worsening"):
severity = "SEV2"
elif pct >= 0.01:
severity = "SEV3"
else:
severity = "SEV4"
if severity in ("SEV1", "SEV2") and not workaround:
action = "mitigate first (rollback / feature flag off), investigate after"
elif severity in ("SEV1", "SEV2") and workaround:
action = "apply the workaround now, investigate on a normal timeline"
else:
action = "investigate on a normal timeline, no immediate mitigation needed"
return severity, action
INCIDENTS = [
{"name": "checkout-fraud-check down", "affected_pct": 1.00, "revenue_sensitive": True, "trend": "worsening", "workaround_exists": False},
{"name": "recsys-home scores stale", "affected_pct": 0.35, "revenue_sensitive": False, "trend": "stable", "workaround_exists": False},
{"name": "loan-scorer p99 up slightly", "affected_pct": 0.03, "revenue_sensitive": True, "trend": "stable", "workaround_exists": True},
{"name": "churn-model batch job delayed", "affected_pct": 1.00, "revenue_sensitive": False, "trend": "stable", "workaround_exists": True},
{"name": "search typo-tolerance off", "affected_pct": 0.005,"revenue_sensitive": False, "trend": "improving", "workaround_exists": False},
]
print(f"{'incident':<30}{'severity':<10}{'first action'}")
for inc in INCIDENTS:
sev, action = triage(inc)
print(f"{inc['name']:<30}{sev:<10}{action}")incident severity first action checkout-fraud-check down SEV1 mitigate first (rollback / feature flag off), investigate after recsys-home scores stale SEV2 mitigate first (rollback / feature flag off), investigate after loan-scorer p99 up slightly SEV3 investigate on a normal timeline, no immediate mitigation needed churn-model batch job delayed SEV1 apply the workaround now, investigate on a normal timeline search typo-tolerance off SEV4 investigate on a normal timeline, no immediate mitigation needed
Notice checkout-fraud-check down and churn-model batch job delayed are both SEV1 — both affect 100% of their traffic — but their first actions differ. One has no workaround and needs an immediate rollback. The other has a workaround, so the fix can wait for daylight while the workaround absorbs the impact.
What changed recently
from datetime import datetime, timedelta
LOOKBACK = timedelta(minutes=30)
def find_suspect_changes(alert_time: datetime, changes: list[dict]) -> list[dict]:
"""Changes inside the lookback window before the alert, most recent first."""
suspects = [c for c in changes if alert_time - LOOKBACK <= c["at"] <= alert_time]
return sorted(suspects, key=lambda c: c["at"], reverse=True)
ALERT_TIME = datetime(2026, 8, 31, 14, 20)
CHANGES = [
{"at": datetime(2026, 8, 31, 14, 5), "what": "deployed loan-scorer v12"},
{"at": datetime(2026, 8, 31, 11, 0), "what": "rotated an API key"},
{"at": datetime(2026, 8, 31, 13, 50), "what": "upstream 'income' feature pipeline re-ran"},
{"at": datetime(2026, 8, 30, 9, 0), "what": "deployed loan-scorer v11"},
]
suspects = find_suspect_changes(ALERT_TIME, CHANGES)
lookback_min = int(LOOKBACK.total_seconds() // 60)
print(f"Alert fired at {ALERT_TIME:%H:%M}. Changes in the {lookback_min} minutes before it:\n")
if not suspects:
print("None found — this is not a 'something recently changed' incident.")
for c in suspects:
minutes_before = int((ALERT_TIME - c["at"]).total_seconds() // 60)
print(f" {c['at']:%H:%M} ({minutes_before:>2} min before) {c['what']}")Alert fired at 14:20. Changes in the 30 minutes before it: 14:05 (15 min before) deployed loan-scorer v12 13:50 (30 min before) upstream 'income' feature pipeline re-ran
Two suspects, not one. find_suspect_changes does not decide which caused the incident — it narrows a search that would otherwise cover the entire deploy history down to two events worth checking first. The API key rotation at 11:00, well outside the window, is correctly excluded.
Common mistakes
Investigating before mitigating on a SEV1. Understanding the exact cause feels more satisfying than a blunt rollback. On a severe, revenue-sensitive incident, roll back or fail over first. Understand it after the bleeding stops.
Widening the lookback window until something matches. If nothing appears in 30 minutes, resist stretching it to six hours to find a culprit. Say plainly that no recent change explains it, and look elsewhere — most likely upstream data or a silent failure.
Triaging once and never revisiting. An incident triaged as SEV3 at minute one can become SEV1 by minute ten if it is spreading. Re-run triage as new numbers arrive.
Confusing "severe" with "hard to fix". A one-line config revert can resolve a SEV1. A confusing, low-impact bug can take a day to understand. Severity is about user impact, not difficulty.
Try it yourself
Add a previously_seen: bool field to an incident, true if this exact failure has happened before. Adjust triage so a previously-seen incident with a known fix skips straight to "apply the known fix", instead of "mitigate first" — you already know what works.
What to learn next
- Silent failures — the failure type that dodges the severity check because nothing looks broken.
- Runbooks for model failures — what to do once triage has pointed at a likely cause.
- Debugging a latency spike — a deep dive into one specific SEV2-shaped symptom.
Researcher — Mathematics and papers.
Incident severity as a decision problem
Triage is a bounded-time classification problem: sort $n$ unresolved signals into $k$ priority classes using partial information, under a cost function where delaying a high-severity incident is far more expensive than a slightly wrong classification of a low one. This asymmetry is why the triage function above is deliberately conservative — it rounds ambiguous cases toward a higher severity ($\text{pct} \geq 0.10$ triggers SEV2 even without revenue sensitivity) rather than toward a lower one, because the expected cost of under-triaging exceeds the cost of over-triaging.
Change-point correlation
find_suspect_changes is a simple time-window filter. A more rigorous version treats the incident's start time as an unknown change point and asks which prior events most plausibly precede it, weighting by both proximity in time and known blast radius of the change type (a full model deploy outranks a documentation update). Formally, this is related to root cause analysis via event correlation in distributed systems, surveyed in Chen et al., Pinpoint: Problem Determination in Large, Dynamic Internet Services (DSN 2002) — correlating a failure's onset against a log of prior state changes, ranked by temporal proximity and historical association with past incidents.
MTTD, MTTM and MTTR as separate clocks
Incident response distinguishes three intervals, each optimised differently:
- MTTD (mean time to detect) — from the fault occurring to it being noticed. Lowered by better monitoring, covered in monitoring and model drift.
- MTTM (mean time to mitigate) — from detection to user impact stopping, whether or not the cause is understood. This is what triage optimises directly.
- MTTR (mean time to resolve/repair) — from detection to the underlying cause being fixed.
Separating MTTM from MTTR is the formal justification for "mitigate first, investigate after": a rollback can drive MTTM to near zero while MTTR remains hours away, and for a user-facing incident, MTTM is very often the number that matters.
Severity taxonomies in practice
Most large operators (Google, AWS, PagerDuty's public incident-response guides) converge on a 3–5 level severity scale similar to the SEV1–SEV4 scheme used above, with SEV1 typically requiring: an incident commander assigned, a communication channel opened, and executive visibility, regardless of how quickly the technical fix arrives. The organisational machinery triggered by a severity label is often a larger design decision than the technical threshold itself — a SEV1 that pages five extra people has a real cost distinct from the incident it responds to.
What to learn next
- Silent failures — the failure type that dodges the severity check because nothing looks broken.
- Runbooks for model failures — what to do once triage has pointed at a likely cause.
- Debugging a latency spike — a deep dive into one specific SEV2-shaped symptom.