Production and MLOps Section 100

Incident Response for ML Systems

What to do at 2am when the model is wrong, slow or quietly broken — and how to be ready for it.

9 of 9 lessons published Three reading levels on every lesson

Start with “On-call for ML systems”

Lessons in order

Work top to bottom. Each lesson assumes the one above it.

  1. On-call for ML systems
  2. Runbooks for model failures
  3. Triaging an ML incident
  4. Silent failures
  5. Debugging a latency spike
  6. When upstream data breaks
  7. Surviving a provider outage
  8. Postmortems for ML incidents
  9. Chaos testing an ML service