Monitoring Models in Production
Alerts people do not ignore
A good ML alert fires rarely and means something every time, because an alert that cries wolf on noise trains the whole team to stop reading it.
- 7 min read
- 3 reading levels
- Published
Read these first
On this page 8
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
A good alert fires rarely, and means something real every time it does.
The analogy you have already lived
A car's seatbelt buzzer beeps for an unbuckled belt. It also beeps for a door left slightly open, low fuel, and a dozen other things, using the exact same sound.
Within a month, most drivers stop reacting to it at all. It has been wrong, or unimportant, too many times.
The same thing happens to a Slack channel full of monitoring alerts that fire on every small wobble in a metric. People mute it, and then miss the one alert that mattered.
Why it exists
Every metric in the earlier lessons of this section has natural, harmless day-to-day noise. A drift score wobbles a little even when nothing is wrong.
Alert on every wobble, and you get many alerts nobody needed to see. Alert too rarely or too late, and a real problem runs for days before anyone notices.
Alerting design is choosing the point between those two failures, on purpose.
How it works
metric crosses the threshold --> is this ONE noisy day,
or a real, sustained shift?
|
wait for confirmation, not the first breach
|
alert only once it holds upThe fix is rarely a smarter threshold. It is usually waiting for the signal to repeat before treating it as real.
A real example you have seen
A weather app does not tell you it will rain because one cloud passed overhead for two minutes. It waits for a sustained pattern before it changes the forecast shown to you.
An ML alert should apply that same patience to a metric before paging a human at 2 a.m.
The honest part
You cannot make an alert both instant and calm. Waiting for confirmation before alerting means a real problem takes a little longer to be noticed.
That trade is almost always worth it. A silenced alert channel notices nothing, instantly or otherwise.
Remember this
- An alert that fires on every small wobble trains people to ignore it.
- Requiring the signal to hold up for a few readings removes most false alarms.
- Choosing when to alert is a genuine trade-off, not a solved problem.
What to learn next
- When model metrics and business metrics disagree — deciding what is even worth alerting on.
- Monitoring by segment — the check this alerting logic most often wraps around.
- Feedback loops in production — a failure mode that alerting alone cannot catch.
Developer — Code and libraries.
Setup
pip install numpyTwo policies on the same noisy metric
A metric that wobbles around a normal baseline, with one real, sustained event hidden inside it. Compare alerting on the first breach against alerting only once a breach holds for three readings in a row.
import numpy as np
rng = np.random.RandomState(0)
days = 90
threshold = 0.10
# A metric (like the PSI score from earlier lessons) that is normally
# noisy around a low baseline, with one real, sustained drift episode.
metric = np.clip(rng.normal(0.05, 0.04, days), 0, None)
metric[60:75] = np.clip(rng.normal(0.35, 0.05, 15), 0, None) # a real event
naive_alerts = metric > threshold
def confirmed_alerts(series, threshold, min_consecutive=3):
breach = series > threshold
flags = np.zeros(len(series), dtype=bool)
run = 0
for i, b in enumerate(breach):
run = run + 1 if b else 0
if run >= min_consecutive:
flags[i] = True
return flags
confirmed = confirmed_alerts(metric, threshold, min_consecutive=3)
print(f"days measured: {days}")
print(f"NAIVE policy -- alert days: {naive_alerts.sum()}")
print(f"CONFIRMED policy -- alert days: {confirmed.sum()}")
real_event_days = set(range(60, 75))
naive_hit_real = any(naive_alerts[list(real_event_days)])
confirmed_hit_real = any(confirmed[list(real_event_days)])
print(f"\ndid NAIVE catch the real event? {naive_hit_real}")
print(f"did CONFIRMED catch the real event? {confirmed_hit_real}")
false_naive = naive_alerts.sum() - sum(naive_alerts[d] for d in real_event_days)
false_confirmed = confirmed.sum() - sum(confirmed[d] for d in real_event_days)
print(f"\nfalse-alarm days -- naive: {int(false_naive)}")
print(f"false-alarm days -- confirmed: {int(false_confirmed)}")days measured: 90 NAIVE policy -- alert days: 26 CONFIRMED policy -- alert days: 13 did NAIVE catch the real event? True did CONFIRMED catch the real event? True false-alarm days -- naive: 11 false-alarm days -- confirmed: 0
Exact output from this seeded script. Both policies caught the real 15-day event. The naive policy also raised 11 false alarms on ordinary noise; the confirmed policy raised zero.
Line-by-line walkthrough
metric[60:75] is the only genuinely abnormal stretch in the whole series. Everything else is noise around a calm baseline, by construction.
naive_alerts fires the instant the threshold is crossed, once, for any reason.
confirmed_alerts tracks a running streak of consecutive breaches, and only raises a flag once that streak reaches min_consecutive. A single noisy day can never trigger it alone.
Common mistakes
Picking a threshold and never revisiting it. A threshold set once, on one week of data, drifts out of date as normal behaviour itself shifts. Review thresholds on a schedule, not only when someone complains.
Sending every alert to the same channel at the same urgency. A slow, minor drift and a service returning errors are not the same emergency. Route them differently, and page a human only for the second kind.
Requiring confirmation for everything, including genuine emergencies. The three-day wait used above is right for a slow metric like drift. A latency spike or an error-rate jump needs to page immediately, with no waiting.
Never closing the loop on false alarms. When an alert turns out to be nothing, record that. A threshold that keeps producing false alarms is telling you it is set wrong.
Try it yourself
Change min_consecutive from 3 to 5. Rerun, and check whether the confirmed policy still catches the real 15-day event, and by how many fewer alert-days.
What to learn next
- When model metrics and business metrics disagree — deciding what is even worth alerting on.
- Monitoring by segment — the check this alerting logic most often wraps around.
- Feedback loops in production — a failure mode that alerting alone cannot catch.
Researcher — Mathematics and papers.
Framing alerting as a sequential decision problem
Naive thresholding tests each observation independently, ignoring the temporal structure of the underlying process. Two established sequential-monitoring tools address this directly:
CUSUM (cumulative sum control chart) accumulates deviation from a reference value, raising an alarm when the running sum exceeds a bound:
$$S_t = \max(0,\ S_{t-1} + (x_t - \mu_0 - k))$$
Where $x_t$ is the observation at time $t$, $\mu_0$ is the in-control mean, and $k$ is a slack parameter (typically half the shift size worth detecting). $S_t$ resets to zero whenever the process is behaving, and grows only under sustained deviation — a principled version of the "require confirmation" heuristic used in the developer example.
EWMA (exponentially weighted moving average) control charts smooth the series with a decay parameter $\lambda$, trading detection speed for noise rejection continuously, rather than the developer example's hard consecutive-count cutoff.
The precision-recall trade-off, made explicit
Every alerting policy sits somewhere on a curve between false-alarm rate and detection delay. Formally, for a fixed detector, the Average Run Length (ARL) under the null (mean time between false alarms) trades directly against the ARL under a real shift (mean time to detect it). Page and Hinkley's original CUSUM analysis derives closed-form ARL expressions for Gaussian shifts, which is the standard reference point for setting $k$ and the alarm threshold jointly, rather than by intuition.
Alert routing and severity
Google's SRE practice (Beyer et al., 2016) formalises the routing decision in the developer tab as a distinction between symptom-based alerts (page a human, something a user would notice) and cause-based alerts (log or ticket, useful for investigation but not urgent on their own) — collapsing this distinction is a leading cause of alert fatigue in practice, independent of any statistical tuning.
Papers
- Page, Continuous Inspection Schemes, Biometrika 1954 — the original CUSUM paper.
- Roberts, Control Chart Tests Based on Geometric Moving Averages, Technometrics 1959 — the EWMA control chart.
- Beyer, Jones, Petoff and Murphy (eds.), Site Reliability Engineering, O'Reilly 2016, Chapter 6 — the symptom-versus-cause alerting distinction.
What to learn next
- When model metrics and business metrics disagree — deciding what is even worth alerting on.
- Monitoring by segment — the check this alerting logic most often wraps around.
- Feedback loops in production — a failure mode that alerting alone cannot catch.