Incident Response for ML Systems

Postmortems for ML incidents

A postmortem is a written, blameless account of what happened, why, and what changes — so the same incident does not happen again for the same reason.

On this page 9
  1. The short answer
  2. The analogy you have already lived
  3. Why it exists
  4. How it works
  5. Blameless, and why that word is load-bearing
  6. Where you have already benefited from this
  7. The honest part
  8. Remember this
  9. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

A postmortem is a written, blameless account of what happened, why, and what changes next.

The analogy you have already lived

Picture a cricket team watching video after a lost match. Nobody stands up and blames the fielder who dropped a catch. The room asks a different question instead: what led to that moment, and what changes before the next match.

A postmortem is that video session, in writing, for an incident. It looks backward calmly, once the fire is out, to make the next fire smaller or rarer.

Why it exists

An incident, once fixed, is tempting to close and forget. The pressure is off, everyone is tired, and there is new work waiting.

Closing it without writing it up wastes the one useful thing a bad night produced: a concrete lesson. Six months later, a slightly different version of the same failure happens. Nobody remembers a fix was ever decided.

A postmortem turns a bad night into something that cannot happen the same way twice.

How it works

Incident resolved
       |
       v
  Write the timeline: what happened, in order, with times
       |
       v
  State the impact: how many users, how much money, how long
       |
       v
  Find the root cause — and the contributing factors around it
       |
       v
  List action items, each with an owner and a real deadline
       |
       v
  Share it. Read by more people than were in the incident.

The document is not the point. The action items getting done is the point. A postmortem with no owners and no deadlines is a story, not a fix.

Blameless, and why that word is load-bearing

"Blameless" does not mean nobody made a mistake. People make mistakes; systems should be built assuming they will.

It means the postmortem asks "what let one mistake become an incident," not "whose fault was this." The first question produces fixes. The second produces people who hide problems next time, afraid of being the one named.

Where you have already benefited from this

Every safety improvement made to a road after a bad accident. Aviation's habit of investigating every serious incident, for decades. School fire drills, refined after every real evacuation. None of them exist to assign blame — all of them exist to make the next one smaller.

The honest part

Writing a good postmortem right after a stressful incident is hard. The first version is often too kind to the real cause. Come back a day later, with a clearer head. Ask again whether the cause you wrote down is the real one, or the comfortable one.

Remember this

  • A postmortem is written, blameless, and shared — its value is the record, not the memory of the room.
  • It asks what let a mistake become an incident, not who made the mistake.
  • Action items with owners and dates are the actual output. Everything above them is context.

What to learn next

Developer — Code and libraries.

Setup

None. Pure Python standard library — dataclasses to structure the facts.

Turning incident facts into a document

Writing a postmortem from a blank page is slower and easier to skip. Structuring the facts first, then rendering them, keeps every postmortem the same shape.

postmortem.py
from dataclasses import dataclass


@dataclass
class Event:
    time: str
    what: str


@dataclass
class ActionItem:
    owner: str
    task: str
    due: str


@dataclass
class Postmortem:
    title: str
    severity: str
    incident_date: str
    summary: str
    timeline: list[Event]
    impact: str
    root_cause: str
    contributing_factors: list[str]
    action_items: list[ActionItem]


def render_postmortem(pm: Postmortem) -> str:
    lines = [
        f"# Postmortem: {pm.title}",
        f"**Severity:** {pm.severity}  **Date:** {pm.incident_date}",
        "", "## Summary", pm.summary, "", "## Timeline",
    ]
    for e in pm.timeline:
        lines.append(f"- {e.time} - {e.what}")
    lines += ["", "## Impact", pm.impact, "", "## Root cause", pm.root_cause,
              "", "## Contributing factors"]
    for f in pm.contributing_factors:
        lines.append(f"- {f}")
    lines += ["", "## Action items"]
    for a in pm.action_items:
        lines.append(f"- [ ] {a.task} - owner: {a.owner}, due: {a.due}")
    return "\n".join(lines)


pm = Postmortem(
    title="Fraud model under-scoring Pune transactions",
    severity="SEV2",
    incident_date="2026-08-31",
    summary=("For nine hours, the fraud model scored every transaction from Pune "
              "as if it originated in Mumbai, silently under-estimating risk."),
    timeline=[
        Event("09:00", "A new dark store in Pune began sending city='PUN' in live traffic"),
        Event("09:00", "City encoder silently mapped unknown cities to Mumbai's index"),
        Event("14:20", "A support ticket flagged three approved transactions that looked wrong"),
        Event("15:05", "On-call engineer confirmed via the default-rate metric, opened SEV2"),
        Event("15:40", "Deployed a fix: unknown cities route to a manual-review queue"),
        Event("18:00", "Backfilled a re-score of the affected nine-hour window"),
    ],
    impact=("Approximately 1,200 transactions scored with the wrong city risk factor; "
            "no confirmed fraud losses found in the backfill."),
    root_cause="The city encoder defaulted unrecognised values to index 0 instead of raising or flagging them.",
    contributing_factors=[
        "No schema validation gate existed for the 'city' feature at ingestion",
        "The default-rate metric existed but had no alert threshold configured",
        "Store-opening announcements from the operations team do not reach the ML team",
    ],
    action_items=[
        ActionItem("ML platform team", "Add an alert on default-rate crossing 5%", "2026-09-07"),
        ActionItem("Data platform team", "Add ingestion-time schema validation for categorical features", "2026-09-14"),
        ActionItem("ML + Ops leads", "Add ML team to the new-store-opening notification list", "2026-09-05"),
    ],
)

print(render_postmortem(pm))
Output
# Postmortem: Fraud model under-scoring Pune transactions
**Severity:** SEV2  **Date:** 2026-08-31

## Summary
For nine hours, the fraud model scored every transaction from Pune as if it originated in Mumbai, silently under-estimating risk.

## Timeline
- 09:00 - A new dark store in Pune began sending city='PUN' in live traffic
- 09:00 - City encoder silently mapped unknown cities to Mumbai's index
- 14:20 - A support ticket flagged three approved transactions that looked wrong
- 15:05 - On-call engineer confirmed via the default-rate metric, opened SEV2
- 15:40 - Deployed a fix: unknown cities route to a manual-review queue
- 18:00 - Backfilled a re-score of the affected nine-hour window

## Impact
Approximately 1,200 transactions scored with the wrong city risk factor; no confirmed fraud losses found in the backfill.

## Root cause
The city encoder defaulted unrecognised values to index 0 instead of raising or flagging them.

## Contributing factors
- No schema validation gate existed for the 'city' feature at ingestion
- The default-rate metric existed but had no alert threshold configured
- Store-opening announcements from the operations team do not reach the ML team

## Action items
- [ ] Add an alert on default-rate crossing 5% - owner: ML platform team, due: 2026-09-07
- [ ] Add ingestion-time schema validation for categorical features - owner: Data platform team, due: 2026-09-14
- [ ] Add ML team to the new-store-opening notification list - owner: ML + Ops leads, due: 2026-09-05

This incident is the same one worked through in silent failures — the postmortem is where its lesson gets written down and turned into three concrete, owned commitments, instead of staying a story a few people remember.

Finding the root cause, not the first excuse

"The encoder had a bug" is true and useless — it does not say why the bug was possible. The five whys technique keeps asking "why" past the first, comfortable answer.

five_whys.py
FIVE_WHYS = [
    ("Why were Pune transactions scored wrong?", "The city encoder mapped an unknown city to Mumbai."),
    ("Why did it map to Mumbai instead of failing?", "The encoder used .get(city, 0), a silent default."),
    ("Why was a silent default written in the first place?", "No policy required unknown categories to be rejected."),
    ("Why was there no such policy?", "Schema validation was optional, added per-feature by habit, not enforced."),
    ("Why was it not enforced?", "No CI check blocks a merge that skips it. This is the fix to make."),
]

for i, (q, a) in enumerate(FIVE_WHYS, 1):
    print(f"{i}. {q}")
    print(f"   {a}")
Output
1. Why were Pune transactions scored wrong?
   The city encoder mapped an unknown city to Mumbai.
2. Why did it map to Mumbai instead of failing?
   The encoder used .get(city, 0), a silent default.
3. Why was a silent default written in the first place?
   No policy required unknown categories to be rejected.
4. Why was there no such policy?
   Schema validation was optional, added per-feature by habit, not enforced.
5. Why was it not enforced?
   No CI check blocks a merge that skips it. This is the fix to make.

Notice where the chain stops. It stops at something a specific action item can fix — a missing CI check — not at "an engineer wrote a bug," which is not fixable in any useful sense. That is the test for whether you have gone deep enough: the final answer names a system gap, not a person.

Common mistakes

Stopping at the first cause. "The encoder had a bug" would have produced an action item to fix that one encoder. Five whys deep, the action item is a CI check that prevents the entire class of bug.

Writing action items with no owner. "We should add validation" gets done by nobody. ActionItem("Data platform team", ..., "2026-09-14") gets done by someone, checked on a date.

Naming a person in the root cause. "Priya forgot to add validation" blames an individual for a gap the system should have caught. "No CI check enforces validation" describes the same gap without a name attached, and points at the actual fix.

Skipping low-severity incidents. A SEV3 that took ten minutes to fix can still reveal a real gap. The postmortem does not need to be long — it needs to exist.

Try it yourself

Add a field to Postmortem for detection_method — how the incident was actually noticed (a support ticket, in this example, not an automated alert). Then add an action item that closes that specific gap: an alert that would have caught it sooner than a human did.

What to learn next

Researcher — Mathematics and papers.

The origin of blameless postmortems

The practice traces most directly to aviation and to Charles Perrow's Normal Accidents (1984), which argues that in sufficiently complex, tightly-coupled systems, failure is not an occasional aberration but a structural inevitability — the design question is not "how do we prevent all failure" but "how do we make failure visible, contained, and instructive." John Allspaw's writing while at Etsy (see "Blameless PostMortems and a Just Culture," 2012) brought this framing explicitly into software operations, arguing that punishing the individual closest to a failure removes the person best positioned to explain how the system actually behaves under stress — and teaches everyone else to hide near-misses rather than report them.

Root cause versus contributing factors

Modern incident-analysis practice increasingly resists the phrase "the root cause," on the grounds that complex systems rarely fail for one reason — they fail when several individually-tolerable weaknesses align. Sidney Dekker's The Field Guide to Understanding 'Human Error' (2014) argues for multiple contributing factors over a single named cause, because a single root cause invites a single fix and a false sense the system is now safe, when in practice several independent factors (in the example above: no schema gate, no alert threshold, no notification path) each independently would have prevented or shortened the incident.

The Swiss cheese model

Reason's Swiss cheese model (Reason, Human Error, 1990) visualises a system's defences as layered slices, each with holes (weaknesses) that shift over time. An incident occurs when the holes in every layer momentarily align, letting a failure pass through all of them. This is the formal version of "contributing factors": each factor above is one slice with a hole in it, and the fix for each closes that hole without claiming to be the single cause.

Postmortem quality as a measurable practice

Google's SRE book (Beyer et al., 2016, ch. 15) proposes concrete criteria for postmortem quality: was it written within a defined window after the incident, does it avoid blaming individuals, are action items tracked to completion with the same rigor as any other engineering work, and is it discoverable by people who were not in the incident. A common failure mode this catches: a well-written postmortem that never gets read past the team that wrote it, so the lesson never reaches the teams likely to hit the same gap next.

Incident review at scale

As an organisation accumulates postmortems, patterns across them become their own signal — the same contributing factor ("no schema validation gate") appearing across unrelated incidents indicates a systemic gap worth fixing once, centrally, rather than once per incident. This is the basis for periodic "meta-postmortem" reviews some organisations run quarterly, aggregating contributing factors across all incidents in the period rather than reading each in isolation.

What to learn next