Incident Response for ML Systems

On-call for ML systems

On-call means someone is reachable and ready to act when a model-serving system breaks, on a rotating schedule, so a 2 a.m. failure never depends on luck.

On this page 9
  1. The short answer
  2. The analogy you have already lived
  3. Why it exists
  4. How it works
  5. Not every alert should wake someone up
  6. Where you have already benefited from this
  7. The honest part
  8. Remember this
  9. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

On-call means someone is awake and reachable, ready to fix a broken model at 2 a.m.

The analogy you have already lived

Walk past a 24-hour pharmacy at night. A small card is taped to the shutter — this week's night-duty chemist, with a phone number. The name on that card changes every week. No single person carries every night forever.

That rotating card is on-call. Someone reachable, on a fixed schedule, so help never depends on who happens to be awake.

Why it exists

A fraud model, a ride-hailing price model, a delivery-time model — these run every hour of every day. Users do not stop tapping "Pay" at 10 p.m.

When one of these breaks at midnight, "we will look at it Monday morning" is not good enough. Money moves, orders go out, someone gets an unfair price, while nobody is watching.

On-call is the answer to one question: who gets told, right now, when something breaks while everyone else is asleep?

How it works

Roster, one week each
  Asha  ->  Rahul  ->  Priya  ->  (back to Asha)

Alert fires at 2:14 am
   |
   v
Page whoever is "primary" this week
   |
   no response in 5 minutes?
   v
Page "secondary"
   |
   still no response?
   v
Page the team lead

The roster decides who. The path down the arrows decides what happens if that person does not answer. Both parts matter — a roster with no escalation path is one missed phone call away from nobody knowing.

Not every alert should wake someone up

This is the part teams get wrong first. If every small wobble in a metric pages a human at night, people learn to silence their phone. Then the alert that matters gets silenced too.

A model serving wrong-but-confident answers for an hour is bad. It is rarely worse than waking a tired person into a mistake. Reserve a page for things actively hurting users right now, that cannot wait until morning.

Where you have already benefited from this

  • UPI fraud checks that keep working at midnight on a payday.
  • Food delivery ETAs that stay sane during a Friday night rush, not only during office hours.
  • Ride-hailing pricing that gets fixed within minutes of going wrong, not the next business day.

The honest part

Being on-call is tiring, and a badly run rotation burns people out. Three choices decide whether a rotation can last. A fair schedule. A real escalation path. Alerts that page only for real emergencies.

Remember this

  • On-call means a specific, reachable person is responsible right now, decided by a rotation.
  • An alert that does not need a human awake should route to a ticket, not a page.
  • Escalation exists so one missed phone call does not mean nobody responds.

What to learn next

Developer — Code and libraries.

Setup

None. Both examples below use only the Python standard library — no install needed.

Deciding what deserves a page

A page interrupts sleep. It has to earn that. This function turns raw signals from a running service into a decision.

alert_policy.py
def classify_alert(event: dict) -> str:
    """Return what should happen to a single alert, given its signals."""
    if not event["health_check_ok"]:
        return "PAGE_NOW"
    if event["error_rate"] > 0.05:
        return "PAGE_NOW"
    if event["p99_latency_ms"] > 2000:
        return "PAGE_NOW"
    if event["drift_score"] > 0.30:
        return "OPEN_TICKET"
    if event["drift_score"] > 0.10:
        return "LOG_ONLY"
    return "OK"


# Six services reporting their last minute of signals. Fixed values, not
# random ones, so this table looks the same on every machine that runs it.
EVENTS = [
    {"name": "loan-scorer",  "health_check_ok": True,  "error_rate": 0.001, "p99_latency_ms": 180,  "drift_score": 0.04},
    {"name": "fraud-model",  "health_check_ok": False, "error_rate": 0.000, "p99_latency_ms": 90,   "drift_score": 0.02},
    {"name": "search-rank",  "health_check_ok": True,  "error_rate": 0.081, "p99_latency_ms": 240,  "drift_score": 0.06},
    {"name": "delivery-eta", "health_check_ok": True,  "error_rate": 0.002, "p99_latency_ms": 3100, "drift_score": 0.05},
    {"name": "churn-model",  "health_check_ok": True,  "error_rate": 0.003, "p99_latency_ms": 210,  "drift_score": 0.34},
    {"name": "recsys-home",  "health_check_ok": True,  "error_rate": 0.004, "p99_latency_ms": 190,  "drift_score": 0.14},
]

print(f"{'service':<14}{'decision':<14}{'why'}")
for e in EVENTS:
    decision = classify_alert(e)
    if decision == "PAGE_NOW":
        if not e["health_check_ok"]:
            why = "health check is down"
        elif e["error_rate"] > 0.05:
            why = f"error rate {e['error_rate']:.1%} over 5%"
        else:
            why = f"p99 latency {e['p99_latency_ms']}ms over 2000ms"
    elif decision == "OPEN_TICKET":
        why = f"drift score {e['drift_score']:.2f} over 0.30"
    elif decision == "LOG_ONLY":
        why = f"drift score {e['drift_score']:.2f}, watching"
    else:
        why = "within normal range"
    print(f"{e['name']:<14}{decision:<14}{why}")

paged = sum(1 for e in EVENTS if classify_alert(e) == "PAGE_NOW")
print(f"\n{paged} of {len(EVENTS)} services would wake someone up right now.")
Output
service       decision      why
loan-scorer   OK            within normal range
fraud-model   PAGE_NOW      health check is down
search-rank   PAGE_NOW      error rate 8.1% over 5%
delivery-eta  PAGE_NOW      p99 latency 3100ms over 2000ms
churn-model   OPEN_TICKET   drift score 0.34 over 0.30
recsys-home   LOG_ONLY      drift score 0.14, watching

3 of 6 services would wake someone up right now.

churn-model and recsys-home both drifted, but neither pages anyone. A slowly drifting model is a daytime problem, worked in monitoring and model drift. A down health check or a 3-second response is not — that is money leaving right now.

The escalation path

A page that nobody sees is the same as no page. This checks minutes elapsed against a fixed acknowledgement window and moves to the next person.

escalate.py
ROTATION = ["Asha (primary)", "Rahul (secondary)", "Team lead"]
ACK_WINDOW_MIN = 5


def escalation_chain(minutes_since_page: int, acknowledged: bool) -> str:
    if acknowledged:
        return "no further paging needed"
    steps_missed = minutes_since_page // ACK_WINDOW_MIN
    index = min(steps_missed, len(ROTATION) - 1)
    return f"page {ROTATION[index]}"


for minute in [0, 4, 5, 9, 10, 16]:
    print(f"t={minute:>2}min  acked=False  ->  {escalation_chain(minute, False)}")
print(f"t= 3min  acked=True   ->  {escalation_chain(3, True)}")
Output
t= 0min  acked=False  ->  page Asha (primary)
t= 4min  acked=False  ->  page Asha (primary)
t= 5min  acked=False  ->  page Rahul (secondary)
t= 9min  acked=False  ->  page Rahul (secondary)
t=10min  acked=False  ->  page Team lead
t=16min  acked=False  ->  page Team lead
t= 3min  acked=True   ->  no further paging needed

minutes_since_page // ACK_WINDOW_MIN turns elapsed time into "how many 5-minute windows have passed", and min(..., len(ROTATION) - 1) stops it from running off the end of the list once everyone has been tried.

A real paging tool (PagerDuty, Opsgenie, Grafana OnCall) runs this same logic against a wall clock and an actual phone call. The function above is the decision underneath it, stripped to something you can read in ten seconds.

Common mistakes

One person on-call every week, forever. It works for a month, then that person quits. Cap it — no more than one week in four for anyone, written down, not assumed.

Paging on every metric that moves. Route anything that is not actively hurting a user to a ticket, as classify_alert does for drift. A rotation that pages ten times a night for nothing trains people to ignore their phone.

No secondary. If the only on-call person is unreachable — dead phone, no signal, asleep through the alarm — nothing happens next unless an escalation step exists.

Treating "server down" and "server slow" as the same alert. They need different first steps. Keep the signals that triggered a page (which classify_alert prints as why) attached to it, so the responder does not start blind.

Try it yourself

Add a queue_depth signal to an event, and a rule in classify_alert that pages when it climbs past a threshold you choose. A growing queue is often the earliest sign of a system in trouble, before latency or errors show it.

What to learn next

Researcher — Mathematics and papers.

Alert volume as a Poisson process

Treat incidents that reach a given on-call rotation as arriving independently at an average rate $\lambda$ per week. Under a Poisson process, the probability of exactly $k$ pages in a week is

$$P(k) = \frac{e^{-\lambda}\lambda^{k}}{k!}$$

where $\lambda$ is the mean number of pages per week and $k$ is a specific count. This is the standard model used to size a rotation: teams generally target $\lambda$ low enough that $P(k \geq 2)$ in a single night stays small, since a second incident while the first is unresolved is where response quality collapses.

MTTA and MTTR

Two numbers describe a rotation's health:

  • MTTA (mean time to acknowledge) — average time from page fired to a human confirming they saw it: $\text{MTTA} = \frac{1}{n}\sum_{i=1}^{n} t_{\text{ack},i}$, where $t_{\text{ack},i}$ is the acknowledgement delay for incident $i$ and $n$ the number of incidents.
  • MTTR (mean time to resolve) — average time from page to the incident being closed, computed the same way using resolution time in place of acknowledgement time.

MTTA rising over weeks, with MTTR flat, points at rotation fatigue or escalation misconfiguration rather than at the underlying system.

The four golden signals

Google's SRE practice (Beyer, Jones, Petoff and Murphy, Site Reliability Engineering, 2016) recommends alerting on four signals for any served system: latency, traffic, errors, saturation. classify_alert above encodes three of the four (latency, an error-rate proxy, and a health check standing in for traffic reachability); a production system adds saturation — queue depth, GPU memory, thread pool occupancy — as its own alert class, not folded into the others.

What is specific to ML on-call

Sculley et al., Hidden Technical Debt in Machine Learning Systems (NeurIPS 2015), documents that ML systems accumulate monitoring debt beyond ordinary services: a model can satisfy every infrastructure signal (up, fast, no errors) while being wrong, because "correct" is a statistical property invisible to a health check. This is why the drift-based OPEN_TICKET / LOG_ONLY split exists separately from the infrastructure PAGE_NOW path — the two failure classes need different urgency and different first responders, one usually an ML engineer and the other usually whoever owns the serving infrastructure.

Rotation design in practice

Follow-the-sun rotations, splitting a 24-hour day across time zones so nobody is paged at 3 a.m. their local time, are standard at organisations large enough to staff multiple regions. Blank-Edelman (ed.), Seeking SRE (O'Reilly, 2018), collects case studies on rotation length, compensation, and the tradeoff between rotation size and individual page frequency — a larger rotation lowers frequency per person but raises the ramp-up cost of maintaining familiarity with the system.

What to learn next

What to learn next

These follow on from what you just read.

  • Incident Response for ML Systems

    Runbooks for model failures

    A runbook is a written, step-by-step set of checks for one specific kind of failure, so the person who gets paged does not have to think from zero.

  • Incident Response for ML Systems

    Triaging an ML incident

    Triage is deciding how bad an incident is and what to do first, in the first few minutes, before anyone understands the full cause.

  • Incident Response for ML Systems

    Silent failures

    A silent failure is a model that answers every request, with no error, while being wrong — because nothing about being wrong trips a health check.