Incident Response for ML Systems

Chaos testing an ML service

Chaos testing means breaking things on purpose, on a normal afternoon, so the first time a failure happens is not during a real incident.

On this page 9
  1. The short answer
  2. The analogy you have already lived
  3. Why it exists
  4. How it works
  5. What to break, in order
  6. Where you have already benefited from this
  7. The honest part
  8. Remember this
  9. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

Chaos testing means breaking things on purpose, so a real incident is never the first time it happens.

The analogy you have already lived

Every school runs fire drills. Nobody waits for a real fire to discover the exit door is locked. Nobody waits to find out half the class does not know where to gather outside.

The alarm rings on a boring Tuesday, on purpose. So the real version, if it ever comes, is not the first time anyone tried it.

Chaos testing is a fire drill for software. You trigger the failure yourself, when nothing real is at risk, and see what actually happens.

Why it exists

Every lesson in this section so far has assumed something breaks, and asked what to do next. Runbooks. Triage. A circuit breaker for a dead provider.

All of that only works if it was actually tested. A circuit breaker nobody has watched trip, a fallback nobody has exercised — these are guesses, not verified behaviour. The gap between "should handle this" and "actually handles this" is exactly where real incidents live.

Chaos testing closes that gap on your schedule, not the failure's.

How it works

Pick one specific thing to break
   (a dependency times out, a cache is empty, a whole server dies)
              |
              v
   Break it on purpose, in a controlled way
              |
              v
   Watch what actually happens
              |
        did it degrade gracefully?  or did it fall over unhandled?
              |                              |
             yes                            no
              |                              |
     confidence, move on            you found a real gap,
                                     for free, on a Tuesday

The outcome that matters is not "nothing broke." It is finding out, safely, exactly what breaks. You want to be the one who finds it, not a real outage.

What to break, in order

Start small. A single slow dependency. Then a dependency that is fully down. Then two things wrong at once — the case teams forget to plan for, since it feels unlikely.

Where you have already benefited from this

Netflix famously runs "Chaos Monkey." It deliberately turns off servers in production, to prove the rest of the system tolerates it. Every time a video kept streaming through a real outage, a chaos test like that is very likely why.

The honest part

Chaos testing on a real production system is a serious practice with its own safety rules. A blast radius kept small. An easy way to stop the experiment. A team that expects it.

This lesson teaches the reasoning, and demonstrates it on a local simulation. Doing this against a real production system needs real infrastructure and real guardrails this lesson does not build for you.

Remember this

  • Chaos testing means breaking something on purpose, so a real failure is not the first time it happens.
  • The useful outcome is finding a real gap, not proving nothing is wrong.
  • Start with one small, controlled failure — and build up to two things failing at once.

What to learn next

Developer — Code and libraries.

Setup

None. Pure Python standard library — random, seeded, so the chaos is reproducible.

Injecting failure on purpose, at a fixed seed

This reuses the circuit breaker from surviving a provider outage, and deliberately fails a fraction of calls to see what the rest of the system does about it.

chaos_experiment.py
import random


class CircuitBreaker:
    def __init__(self, failure_threshold=3, cooldown_calls=3):
        self.failure_threshold = failure_threshold
        self.cooldown_calls = cooldown_calls
        self.state = "CLOSED"
        self.consecutive_failures = 0
        self.calls_since_opened = 0

    def call(self, provider_fn, fallback_fn, *args):
        if self.state == "OPEN":
            self.calls_since_opened += 1
            if self.calls_since_opened >= self.cooldown_calls:
                self.state = "HALF_OPEN"
            else:
                return fallback_fn(*args)
        try:
            result = provider_fn(*args)
        except RuntimeError:
            self.consecutive_failures += 1
            if self.state == "HALF_OPEN" or self.consecutive_failures >= self.failure_threshold:
                self.state = "OPEN"
                self.calls_since_opened = 0
            return fallback_fn(*args)
        self.state = "CLOSED"
        self.consecutive_failures = 0
        return result


def run_chaos_experiment(cache: dict, seed: int, n_requests: int = 50, failure_rate: float = 0.35):
    """A fixed seed makes this reproducible. Real chaos tooling injects
    failure into a live system; this injects into a local simulation of
    one, which is what a single laptop can honestly demonstrate."""
    rng = random.Random(seed)
    cities = ["MUM", "DEL", "BLR", "PUN", "HYD", "KOL"]

    def flaky_provider(city):
        if rng.random() < failure_rate:
            raise RuntimeError("provider timeout")
        return f"live({city})"

    def cached_fallback(city):
        if city not in cache:
            raise KeyError(f"no cached value for {city}")
        return f"cached({city})"

    breaker = CircuitBreaker()
    ok, unhandled = 0, []
    for _ in range(n_requests):
        city = rng.choice(cities)
        try:
            breaker.call(flaky_provider, cached_fallback, city)
            ok += 1
        except KeyError as e:
            unhandled.append(str(e))
    return ok, unhandled


# The cache only ever warmed up for the three cities the system has always
# served. Pune, Hyderabad and Kolkata are newer and have no cache entry.
warm_cache = {"MUM": 1, "DEL": 1, "BLR": 1}
ok, unhandled = run_chaos_experiment(warm_cache, seed=7)
print("Experiment 1: cold cache for new cities")
print(f"  {ok} requests survived, {len(unhandled)} were unhandled failures")
for u in unhandled[:3]:
    print(f"    unhandled: {u}")
if len(unhandled) > 3:
    print(f"    ... and {len(unhandled) - 3} more")
Output
Experiment 1: cold cache for new cities
  40 requests survived, 10 were unhandled failures
    unhandled: 'no cached value for PUN'
    unhandled: 'no cached value for HYD'
    unhandled: 'no cached value for KOL'
    ... and 7 more

Ten of fifty simulated requests crashed. Not because the provider was down — the circuit breaker handles that fine, as it did in the previous lesson — but because the fallback itself had nothing for three cities. The breaker degraded to a fallback that had its own untested edge case.

This is what chaos testing is for: warm_cache looked like a perfectly reasonable fallback until deliberately failing the provider revealed what it does not cover.

Fixing the gap the experiment found

chaos_experiment_fixed.py
from chaos_experiment import CircuitBreaker
import random

warm_cache = {"MUM": 1, "DEL": 1, "BLR": 1}


def cached_fallback_with_default(cache, city):
    if city in cache:
        return f"cached({city})"
    return "default-neutral-score"


def run_chaos_experiment_fixed(cache: dict, seed: int, n_requests: int = 50, failure_rate: float = 0.35):
    rng = random.Random(seed)
    cities = ["MUM", "DEL", "BLR", "PUN", "HYD", "KOL"]

    def flaky_provider(city):
        if rng.random() < failure_rate:
            raise RuntimeError("provider timeout")
        return f"live({city})"

    breaker = CircuitBreaker()
    ok, unhandled = 0, []
    for _ in range(n_requests):
        city = rng.choice(cities)
        try:
            breaker.call(flaky_provider, lambda c: cached_fallback_with_default(cache, c), city)
            ok += 1
        except Exception as e:
            unhandled.append(str(e))
    return ok, unhandled


ok2, unhandled2 = run_chaos_experiment_fixed(warm_cache, seed=7)
print("Experiment 2: same chaos, with a final default answer added")
print(f"  {ok2} requests survived, {len(unhandled2)} were unhandled failures")
Output
Experiment 2: same chaos, with a final default answer added
  50 requests survived, 0 were unhandled failures

Same seed, same injected failures, same fifty requests. The only change is a final, always-available default in the fallback path. Zero unhandled failures — because there is nothing left for the system to have no answer for.

Line-by-line walkthrough

random.Random(seed) creates an independent random generator, not the shared global one — the exact same sequence of "does this call fail" decisions happens every run, which is what makes an experiment like this reproducible and worth discussing with a team, not a one-off fluke.

failure_rate=0.35 is a deliberately harsh, artificial number. Real outages are rarely this severe. Chaos testing turns the dial higher than reality on purpose, to find gaps faster than waiting for reality to find them at its own pace.

Common mistakes

Only testing the happy path of the fallback. warm_cache worked fine in every normal test, because normal tests never simulate the provider being down at all. The gap only appears under the exact condition — a provider failure — that a fallback exists for.

Running chaos experiments only in your head. "I think it would handle that" and "it handled fifty simulated failures with zero unhandled errors" are different claims. Only one of them is evidence.

Injecting failure without a fixed seed. An unseeded random failure is different every run, which makes a finding hard to reproduce and hard to discuss precisely. Seed it, exactly as run_chaos_experiment does.

Testing only one failure at a time, forever. The hardest bugs live at the intersection — a provider outage during a cold cache, exactly as this experiment combines. Real systems fail in combinations more often than teams expect.

Try it yourself

Raise failure_rate to 0.8 in run_chaos_experiment_fixed and rerun. The default answer still catches everything — confirm that, then ask a harder question: is a system safely answering "default-neutral-score" for 80% of requests actually a system you would want to leave running, or is there a threshold where the honest response is to stop serving instead of guessing?

What to learn next

Researcher — Mathematics and papers.

Chaos engineering as a discipline

Basiri et al., Chaos Engineering (IEEE Software, 2016), formalise the practice pioneered at Netflix as the discipline of experimenting on a system in order to build confidence in its capability to withstand turbulent conditions in production. Their principles, adapted here:

  1. Define "steady state" as a measurable output (in this lesson, "zero unhandled failures"), not an internal detail.
  2. Hypothesize that steady state holds in both the control group and the experimental group with an injected fault.
  3. Inject real-world events — a dependency failure, a cache miss, network latency.
  4. Try to disprove the hypothesis by looking for a difference between the groups.

run_chaos_experiment above follows this shape in miniature: warm_cache is the experimental condition, and the discovery that ten requests broke is a disproven hypothesis about steady state — the useful outcome the methodology is designed to produce.

Blast radius and the safety of the practice itself

Running chaos experiments against real production traffic carries real risk, which is why mature practice bounds the blast radius — the fraction of real traffic or infrastructure an experiment can affect — starting deliberately small (a single instance, a small traffic percentage) and expanding only as confidence grows. An abort mechanism that can halt an experiment instantly, monitored throughout by the same signals covered in on-call for ML systems, is considered a prerequisite, not an optional extra, before running any experiment against a live system.

What is specific to ML services

An ML-serving system has failure modes beyond the infrastructure faults chaos engineering originally targeted (a dead server, a network partition). Worth injecting deliberately, beyond what this lesson demonstrates:

  • A stale or corrupted model artefact, testing whether a health check (from model serving) actually catches it.
  • An adversarial or malformed input distribution, testing the validation gate from when upstream data breaks under genuinely hostile input, not only accidental data drift.
  • A feature store returning stale values, testing whether serving code can tell the difference between "no value" and "an old value" — a distinction silent failures shows is easy to lose.

Game days

Beyond automated fault injection, many organisations run periodic game days — a scheduled, team-wide exercise simulating a specific incident scenario end to end, including the human response: who gets paged, which runbook gets followed, whether the postmortem process itself works. Google's SRE book (Beyer et al., 2016, ch. 28) documents this as "wheel of misfortune" exercises, explicitly designed to test the incident-response process, not only the code — closing the loop back to on-call and runbooks with the same rigor applied to the system itself.

What to learn next