Monitoring Models in Production

Alerts people do not ignore

A good ML alert fires rarely and means something every time, because an alert that cries wolf on noise trains the whole team to stop reading it.

On this page 8
  1. The short answer
  2. The analogy you have already lived
  3. Why it exists
  4. How it works
  5. A real example you have seen
  6. The honest part
  7. Remember this
  8. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

A good alert fires rarely, and means something real every time it does.

The analogy you have already lived

A car's seatbelt buzzer beeps for an unbuckled belt. It also beeps for a door left slightly open, low fuel, and a dozen other things, using the exact same sound.

Within a month, most drivers stop reacting to it at all. It has been wrong, or unimportant, too many times.

The same thing happens to a Slack channel full of monitoring alerts that fire on every small wobble in a metric. People mute it, and then miss the one alert that mattered.

Why it exists

Every metric in the earlier lessons of this section has natural, harmless day-to-day noise. A drift score wobbles a little even when nothing is wrong.

Alert on every wobble, and you get many alerts nobody needed to see. Alert too rarely or too late, and a real problem runs for days before anyone notices.

Alerting design is choosing the point between those two failures, on purpose.

How it works

   metric crosses the threshold  -->  is this ONE noisy day,
                                        or a real, sustained shift?
                                              |
                       wait for confirmation, not the first breach
                                              |
                              alert only once it holds up

The fix is rarely a smarter threshold. It is usually waiting for the signal to repeat before treating it as real.

A real example you have seen

A weather app does not tell you it will rain because one cloud passed overhead for two minutes. It waits for a sustained pattern before it changes the forecast shown to you.

An ML alert should apply that same patience to a metric before paging a human at 2 a.m.

The honest part

You cannot make an alert both instant and calm. Waiting for confirmation before alerting means a real problem takes a little longer to be noticed.

That trade is almost always worth it. A silenced alert channel notices nothing, instantly or otherwise.

Remember this

  • An alert that fires on every small wobble trains people to ignore it.
  • Requiring the signal to hold up for a few readings removes most false alarms.
  • Choosing when to alert is a genuine trade-off, not a solved problem.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install numpy

Two policies on the same noisy metric

A metric that wobbles around a normal baseline, with one real, sustained event hidden inside it. Compare alerting on the first breach against alerting only once a breach holds for three readings in a row.

alerting.py
import numpy as np

rng = np.random.RandomState(0)

days = 90
threshold = 0.10

# A metric (like the PSI score from earlier lessons) that is normally
# noisy around a low baseline, with one real, sustained drift episode.
metric = np.clip(rng.normal(0.05, 0.04, days), 0, None)
metric[60:75] = np.clip(rng.normal(0.35, 0.05, 15), 0, None)  # a real event

naive_alerts = metric > threshold


def confirmed_alerts(series, threshold, min_consecutive=3):
    breach = series > threshold
    flags = np.zeros(len(series), dtype=bool)
    run = 0
    for i, b in enumerate(breach):
        run = run + 1 if b else 0
        if run >= min_consecutive:
            flags[i] = True
    return flags


confirmed = confirmed_alerts(metric, threshold, min_consecutive=3)

print(f"days measured:                {days}")
print(f"NAIVE policy   -- alert days:  {naive_alerts.sum()}")
print(f"CONFIRMED policy -- alert days: {confirmed.sum()}")

real_event_days = set(range(60, 75))
naive_hit_real = any(naive_alerts[list(real_event_days)])
confirmed_hit_real = any(confirmed[list(real_event_days)])
print(f"\ndid NAIVE catch the real event?     {naive_hit_real}")
print(f"did CONFIRMED catch the real event? {confirmed_hit_real}")

false_naive = naive_alerts.sum() - sum(naive_alerts[d] for d in real_event_days)
false_confirmed = confirmed.sum() - sum(confirmed[d] for d in real_event_days)
print(f"\nfalse-alarm days -- naive:     {int(false_naive)}")
print(f"false-alarm days -- confirmed: {int(false_confirmed)}")
Output
days measured:                90
NAIVE policy   -- alert days:  26
CONFIRMED policy -- alert days: 13

did NAIVE catch the real event?     True
did CONFIRMED catch the real event? True

false-alarm days -- naive:     11
false-alarm days -- confirmed: 0

Exact output from this seeded script. Both policies caught the real 15-day event. The naive policy also raised 11 false alarms on ordinary noise; the confirmed policy raised zero.

Line-by-line walkthrough

metric[60:75] is the only genuinely abnormal stretch in the whole series. Everything else is noise around a calm baseline, by construction.

naive_alerts fires the instant the threshold is crossed, once, for any reason.

confirmed_alerts tracks a running streak of consecutive breaches, and only raises a flag once that streak reaches min_consecutive. A single noisy day can never trigger it alone.

Common mistakes

Picking a threshold and never revisiting it. A threshold set once, on one week of data, drifts out of date as normal behaviour itself shifts. Review thresholds on a schedule, not only when someone complains.

Sending every alert to the same channel at the same urgency. A slow, minor drift and a service returning errors are not the same emergency. Route them differently, and page a human only for the second kind.

Requiring confirmation for everything, including genuine emergencies. The three-day wait used above is right for a slow metric like drift. A latency spike or an error-rate jump needs to page immediately, with no waiting.

Never closing the loop on false alarms. When an alert turns out to be nothing, record that. A threshold that keeps producing false alarms is telling you it is set wrong.

Try it yourself

Change min_consecutive from 3 to 5. Rerun, and check whether the confirmed policy still catches the real 15-day event, and by how many fewer alert-days.

What to learn next

Researcher — Mathematics and papers.

Framing alerting as a sequential decision problem

Naive thresholding tests each observation independently, ignoring the temporal structure of the underlying process. Two established sequential-monitoring tools address this directly:

CUSUM (cumulative sum control chart) accumulates deviation from a reference value, raising an alarm when the running sum exceeds a bound:

$$S_t = \max(0,\ S_{t-1} + (x_t - \mu_0 - k))$$

Where $x_t$ is the observation at time $t$, $\mu_0$ is the in-control mean, and $k$ is a slack parameter (typically half the shift size worth detecting). $S_t$ resets to zero whenever the process is behaving, and grows only under sustained deviation — a principled version of the "require confirmation" heuristic used in the developer example.

EWMA (exponentially weighted moving average) control charts smooth the series with a decay parameter $\lambda$, trading detection speed for noise rejection continuously, rather than the developer example's hard consecutive-count cutoff.

The precision-recall trade-off, made explicit

Every alerting policy sits somewhere on a curve between false-alarm rate and detection delay. Formally, for a fixed detector, the Average Run Length (ARL) under the null (mean time between false alarms) trades directly against the ARL under a real shift (mean time to detect it). Page and Hinkley's original CUSUM analysis derives closed-form ARL expressions for Gaussian shifts, which is the standard reference point for setting $k$ and the alarm threshold jointly, rather than by intuition.

Alert routing and severity

Google's SRE practice (Beyer et al., 2016) formalises the routing decision in the developer tab as a distinction between symptom-based alerts (page a human, something a user would notice) and cause-based alerts (log or ticket, useful for investigation but not urgent on their own) — collapsing this distinction is a leading cause of alert fatigue in practice, independent of any statistical tuning.

Papers

  • Page, Continuous Inspection Schemes, Biometrika 1954 — the original CUSUM paper.
  • Roberts, Control Chart Tests Based on Geometric Moving Averages, Technometrics 1959 — the EWMA control chart.
  • Beyer, Jones, Petoff and Murphy (eds.), Site Reliability Engineering, O'Reilly 2016, Chapter 6 — the symptom-versus-cause alerting distinction.

What to learn next