Incident Response for ML Systems

Runbooks for model failures

A runbook is a written, step-by-step set of checks for one specific kind of failure, so the person who gets paged does not have to think from zero.

On this page 9
  1. The short answer
  2. The analogy you have already lived
  3. Why it exists
  4. How it works
  5. What a runbook is not
  6. Where you have already benefited from this
  7. The honest part
  8. Remember this
  9. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

A runbook is a set of steps for one failure, so nobody starts from zero at 2 a.m.

The analogy you have already lived

Every lift you have ridden has a small card near the buttons: what to do if it stops between floors. Press the alarm. Do not force the doors. Wait for the number on the card to call you back.

Nobody reads that card for the first time while stuck. It is written calmly, in advance, by someone who is not panicking, for the benefit of someone who will be. A runbook is that card, for a broken model.

Why it exists

At 2 a.m., a paged person is not at their best. They are tired, the dashboard is unfamiliar in the dark, and every minute the model stays broken has a cost.

Without a runbook, that person starts from a guess. They open five tabs, check things in a random order, and re-derive knowledge somebody already had.

A runbook removes the guessing. It is written by someone calm, ahead of time, for exactly this failure.

How it works

Alert: "fraud-model returning odd scores"
        |
        v
Open the runbook for THIS alert (not a general one)
        |
   Step 1: is the health check green?     no -> restart, stop here
        |  yes
   Step 2: does a known-good request work? no -> check the last deploy
        |  yes
   Step 3: is the response shaped right?   no -> schema problem, see [silent failures]
        |  yes
   Step 4: escalate — this needs a human decision

Each step has one question and one action. Nobody has to decide what to check next — the runbook already decided, calmly, before the incident happened.

What a runbook is not

It is not a general troubleshooting guide, and it is not documentation of how the system works. It is tied to one specific alert, and it stops being useful the moment it tries to cover everything.

A runbook titled "when the model server is unhealthy" is too broad to trust. A runbook titled "when /healthz on fraud-model returns not_ready" can be followed at 2 a.m. without thinking.

Where you have already benefited from this

Every call centre script a support agent reads from. Every pre-flight checklist a pilot runs before takeoff, in the same order, every single time — however experienced they are.

The honest part

Runbooks go stale. The system changes, and the runbook still describes last year's version. A stale runbook is worse than none, because it wastes the exact minutes it was meant to save. Attach a "last verified" note to each one, and revisit it whenever the incident it covers happens again.

Remember this

  • A runbook covers one specific failure, with steps written calmly, in advance.
  • Each step is one question and one action — no judgment calls in the middle of a fire.
  • A runbook that is never re-checked goes stale, and a stale one can mislead worse than none.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install fastapi "uvicorn[standard]" httpx

A tiny service to have an incident on

This mirrors the server from model serving, trimmed down for this lesson.

service.py
from contextlib import asynccontextmanager

from fastapi import FastAPI
from pydantic import BaseModel

STATE = {}


@asynccontextmanager
async def lifespan(app: FastAPI):
    STATE["model_loaded"] = True
    STATE["model_version"] = "v3"
    yield
    STATE.clear()


app = FastAPI(lifespan=lifespan)


class Row(BaseModel):
    income: float
    years: float
    age: float


@app.get("/healthz")
def healthz():
    return {"status": "ok" if STATE.get("model_loaded") else "not_ready"}


@app.post("/predict")
def predict(row: Row):
    score = min(0.99, row.income / 100)  # a stand-in for a real model
    return {"probability": round(score, 4), "model_version": STATE["model_version"]}

The runbook itself, written as code

A runbook does not have to live only in a document. Writing the checks as a function means anyone on the rotation runs the exact same steps, in the exact same order, every time — the same guarantee a pilot's checklist gives.

runbook.py
from fastapi.testclient import TestClient

from service import app


def run_runbook(client: TestClient) -> list[tuple[str, bool, str]]:
    """Each entry: (step name, passed, detail). Stops adding steps after the
    first hard failure — the same point a human should stop and escalate."""
    steps = []

    r = client.get("/healthz")
    ok = r.status_code == 200 and r.json().get("status") == "ok"
    steps.append(("health check responds 'ok'", ok, f"HTTP {r.status_code}, body {r.json()}"))
    if not ok:
        steps.append(("STOP", False, "restart the service, then re-run this runbook"))
        return steps

    sample = {"income": 60.0, "years": 7.0, "age": 34}
    r = client.post("/predict", json=sample)
    ok = r.status_code == 200
    steps.append(("a known-good request returns HTTP 200", ok, f"HTTP {r.status_code}"))
    if not ok:
        steps.append(("STOP", False, "check the last deploy's logs for a startup error"))
        return steps

    body = r.json()
    has_prob = "probability" in body and isinstance(body["probability"], float)
    steps.append(("response has a numeric 'probability' field", has_prob, f"body={body}"))

    in_range = has_prob and 0.0 <= body["probability"] <= 1.0
    steps.append(("probability is between 0 and 1", in_range, f"got {body.get('probability')}"))

    has_version = "model_version" in body
    steps.append(("response reports which model version answered", has_version,
                  f"version={body.get('model_version')}"))
    return steps


with TestClient(app) as client:
    print("=== Runbook: fraud-model returning odd scores ===")
    for name, ok, detail in run_runbook(client):
        mark = "PASS" if ok else "FAIL"
        print(f"[{mark}] {name}\n       {detail}")
Output
=== Runbook: fraud-model returning odd scores ===
[PASS] health check responds 'ok'
       HTTP 200, body {'status': 'ok'}
[PASS] a known-good request returns HTTP 200
       HTTP 200
[PASS] response has a numeric 'probability' field
       body={'probability': 0.6, 'model_version': 'v3'}
[PASS] probability is between 0 and 1
       got 0.6
[PASS] response reports which model version answered
       version=v3

A clean run. Now see what the same runbook prints when the model genuinely failed to load — change STATE["model_loaded"] = True to False in service.py's lifespan, simulating a missing model file at startup, and run it again.

Output
=== Runbook: fraud-model returning odd scores ===
[FAIL] health check responds 'ok'
       HTTP 200, body {'status': 'not_ready'}
[FAIL] STOP
       restart the service, then re-run this runbook

Both output blocks above are real runs of this exact code, on one machine — a TestClient run has no network, no timing variance, and no randomness, so this pass/fail pattern will repeat identically for anyone who runs it.

Line-by-line walkthrough

run_runbook returns after the first STOP. This matches how a human should behave — there is no point checking whether the probability is in range on a server that is not even healthy.

Each steps.append carries a detail string. A runbook that only says PASS or FAIL forces the responder to go re-gather the evidence by hand. One that carries the evidence with it saves that trip.

Common mistakes

Writing runbooks as prose instead of steps. A paragraph explaining "the model might be unhealthy, in which case check the logs" makes a tired reader parse English at 2 a.m. A numbered step with one question is faster to follow under pressure.

One giant runbook for the whole system. Split by alert. The runbook for "health check red" and the runbook for "latency over 2 seconds" share almost nothing.

No owner, no review date. A runbook nobody owns does not get updated when the system changes underneath it. Put a name and a date at the top.

Steps that assume tools the responder might not have open. "Check the dashboard" is not a step if the dashboard needs a URL, a login, and a specific panel. Write the exact link.

Try it yourself

Add a fifth step to run_runbook: send a request missing the income field and check that the server returns 422, not a crash. That is the validation behaviour from model serving, and a runbook is exactly where you would confirm it still holds during an incident.

What to learn next

Researcher — Mathematics and papers.

Runbooks versus playbooks versus automation

The incident-response literature draws a three-way distinction worth keeping precise:

  • A runbook is a fixed sequence of diagnostic and remediation steps for one known failure mode, meant to be followed exactly.
  • A playbook is broader guidance for a class of incidents (for example, "a security incident"), where judgment fills gaps the document cannot.
  • Automation (sometimes "auto-remediation") is a runbook with no human step — the checks and fixes run unattended, with a human notified only on escalation.

The run_runbook function above sits at the boundary between the first and the third: fully automatic today, and a natural candidate to trigger a real restart automatically once its false-positive rate is known to be low.

Where runbooks fit in the incident lifecycle

Allspaw's writing on resilience engineering in software operations (see Allspaw, "Trade-Offs Under Pressure," 2015, and the broader Etsy/Netflix chaos-engineering lineage) frames a runbook as encoding anticipated failure — a scenario the team has seen or predicted. Novel incidents, by definition, have no runbook, which is why triaging an incident and general diagnostic skill remain necessary even in a mature on-call culture with good coverage.

Coverage as a metric

A useful operational number: the fraction of paged incidents in the last quarter that had an applicable runbook, versus those handled from first principles. A low and falling number signals either novel failure modes outpacing documentation, or runbooks that exist but were not findable during the incident — a discoverability failure, not a content one. Both are worth tracking separately; conflating them hides which problem to fix.

Executable runbooks and their limits

Encoding a runbook as code, as this lesson does, buys reproducibility — the same steps run in the same order — at the cost of coverage. Code can check what it was written to check. A human responder notices the unfamiliar log line, the coworker's Slack message about a related deploy, or the customer complaint that arrived through a different channel than the alert. Google's SRE literature on toil reduction (Beyer et al., 2016, ch. 5) recommends automating the mechanical majority of steps while keeping a human in the loop for judgment calls — the STOP, escalate branch in a well-designed runbook, not the absence of one.

What to learn next