False alarms and alert fatigue
A monitor that cries wolf hundreds of times a day trains the humans around it to stop listening, even to the one alarm that matters.
- 8 min read
- 3 reading levels
- Published
Read these first
On this page 6
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Alert fatigue means an alarm fires so often that people stop reacting to it.
Think of a car alarm on a busy street. The first time it goes off, everyone looks up. By the tenth false trigger that week, nobody looks up anymore — not even during a real theft.
Hospital monitors can fall into the exact same trap. A machine that beeps at the smallest twitch trains the nurses near it to tune the beeping out. That is the same way a street tunes out a car alarm.
Why it exists
Building an alarm that never misses a real problem is not hard. Set the threshold low enough, and it catches almost everything. The cost is a flood of false alarms, because a low threshold also catches every minor, harmless blip.
The reverse is equally easy. Set the threshold high, and false alarms nearly disappear. The cost is missed real events — a high threshold also filters out the early, subtle version of a real problem.
There is no setting that avoids both costs. Most systems get tuned only to "catch everything." That choice picks the first failure mode, often without anyone deciding to.
How it works
Low threshold: catches almost every real problem + huge number of false alarms
|
v
staff learn to ignore alarms
|
v
the one real alarm gets missed tooThe danger is not the false alarms by themselves. It is what constant false alarms do to how humans respond to every alarm, including the true ones.
Where you have already seen it
- Phone notifications. An app that pings constantly gets its notifications silenced entirely — including the one that mattered.
- A smoke detector that goes off every time you cook. People start disabling it, or moving it, rather than fixing the underlying sensitivity.
An honest warning
A model with genuinely excellent accuracy numbers can still generate an unusable flood of false alarms. Rare events, multiplied across thousands of daily checks, still add up to a lot of noise. The next section works through exactly why, with real numbers.
Remember this
- The danger of a noisy alarm system is not the false alarms themselves. It is what they train humans to ignore.
- There is no threshold that eliminates both false alarms and missed events at the same time.
- A system's accuracy on paper can look fine while still generating an unworkable number of alarms per shift.
- Choosing where to set that threshold is a clinical decision, not a modelling one. It needs a hospital's own clinicians involved before deployment, not a default left in by a vendor.
What to learn next
- Predicting patient deterioration — where these alarm thresholds usually come from.
- Choosing a threshold from costs — turning "too many false alarms" into an actual number to optimise.
- Screening: one positive in a thousand scans — the same base-rate problem, in a different clinical setting.
Developer — Code and libraries.
Setup
pip install pandasMinimal runnable code
The scenario: a 20-bed ward, monitored continuously.
import pandas as pd
# A 20-bed ward, checked every 5 minutes: 288 checks per bed per day.
beds = 20
checks_per_bed_per_day = 288
total_checks = beds * checks_per_bed_per_day
# ILLUSTRATIVE prevalence, not a measured statistic: about 1 in 500 checks
# catches a patient who is genuinely about to decline in the next hour.
prevalence = 1 / 500
def alarms_per_day(sensitivity, specificity):
positives = total_checks * prevalence
negatives = total_checks - positives
true_alarms = positives * sensitivity
false_alarms = negatives * (1 - specificity)
total_alarms = true_alarms + false_alarms
ppv = true_alarms / total_alarms
return total_alarms, ppv
rows = []
for specificity in [0.90, 0.95, 0.99, 0.999]:
total_alarms, ppv = alarms_per_day(sensitivity=0.90, specificity=specificity)
rows.append({
"specificity": specificity,
"alarms_per_day": round(total_alarms, 1),
"real_alarm_chance": round(ppv, 3),
})
print(pd.DataFrame(rows).to_string(index=False)) specificity alarms_per_day real_alarm_chance
0.900 585.2 0.018
0.950 297.8 0.035
0.990 67.9 0.153
0.999 16.1 0.643What actually happened
At 90% specificity — a number that sounds respectable — this ward gets 585 alarms a day. Divide that across a 24-hour shift and it is an alarm roughly every two and a half minutes, on every bed combined. Only 1.8% of those alarms are real.
Even at 99% specificity, which most people would call an excellent model, more than 5 out of every 6 alarms are still false. Only pushing specificity to 99.9% gets the real-alarm chance above even coin-flip odds.
Line by line, the parts that are not obvious:
total_checksscales with how often a check happens, not only with how many patients there are. Frequent monitoring multiplies the false-alarm opportunity, even if the model itself never changes.real_alarm_chancehere is the positive predictive value (PPV) — of the alarms that fire, what fraction are true. It is not the same number as sensitivity, and it degrades fast as the real event becomes rarer.- Sensitivity is held fixed at 90% throughout this table. Only specificity moves — showing that specificity, not sensitivity, is usually the lever that actually controls alarm burden.
Common mistakes
Reporting sensitivity and specificity without also reporting expected alarms per day. Both numbers can look excellent in isolation and still describe an unusable system once multiplied by how often checks happen.
Assuming higher specificity is free. Pushing specificity up almost always costs some sensitivity — the real, missed-event trade-off this table hides by holding sensitivity fixed.
Designing the alarm threshold around the model's best F1 score. F1 balances precision and recall as if they cost the same. A missed deterioration and an extra beep on a monitor do not cost the same, and the threshold should reflect that.
Try it yourself
Change checks_per_bed_per_day to 12 — simulating hourly checks instead of every five minutes. Rerun the table and notice how much the alarm burden drops, even with identical sensitivity and specificity. Monitoring frequency is a design lever too, not only the model.
What to learn next
- Choosing a threshold from costs — formalising what a false alarm and a missed event should each be weighted at.
- ROC vs precision-recall curves — the general machinery behind sensitivity and specificity trade-offs.
- Validating a model before it touches patients — where alarm-burden numbers like these get tested against real staff, not only a spreadsheet.
Researcher — Mathematics and papers.
Formalising alarm burden
For a monitoring system running N checks over some period, with prevalence π, sensitivity Se, and specificity Sp, the expected number of alarms is:
E[alarms] = N*π*Se + N*(1-π)*(1-Sp)N— total number of checks in the periodπ— prevalence, the base rate of the real event per checkSe— sensitivity (true positive rate)Sp— specificity (true negative rate)
The positive predictive value follows directly from Bayes' theorem:
PPV = (π * Se) / (π * Se + (1-π) * (1-Sp))As π -> 0, PPV -> 0 regardless of how high Se and Sp are, provided Sp < 1. This is the base-rate fallacy made quantitative: no realistic classifier avoids a low PPV against a sufficiently rare event, because the (1-Sp) term is multiplied by a much larger (1-π) population of negatives.
The alarm-rate/coverage trade-off
Treating this as a decision problem, define a cost ratio C_FA / C_miss — the cost of one false alarm relative to the cost of one missed event. The Neyman-Pearson lemma implies the optimal decision rule is a likelihood-ratio threshold set by that cost ratio, connecting directly to choosing a threshold from costs. In practice, C_FA is rarely a fixed number — it grows nonlinearly as alarm frequency increases, because each additional alarm degrades the credibility of every alarm around it. This nonlinearity is what standard threshold-selection frameworks, which assume a fixed per-alarm cost, do not capture.
Documented consequences
Alarm fatigue has been named a top patient-safety technology hazard in recurring hazard reports from clinical safety organisations (e.g. the ECRI Institute's annual health technology hazard lists), and has been the subject of Joint Commission Sentinel Event Alerts in the United States addressing clinical alarm safety directly. Readers should consult those primary sources for current, verified figures rather than any restated number here — the specific alarm counts and harm statistics vary by report and by year, and are outside what this lesson can responsibly assert.
Cost
Reducing alarm burden without losing sensitivity generally requires either (a) a better-discriminating model — raising both Se and Sp simultaneously, which is the hard, slow route — or (b) smarter alarm logic: delay-and-confirm rules, trend-based alarms instead of instantaneous threshold crossings, and alarm suppression during known noisy periods (e.g. patient repositioning). These are engineering interventions on top of the model, not substitutes for a better model.
Key references
- Sendelbach, S. & Funk, M. (2013). Alarm Fatigue: A Patient Safety Concern. AACN Advanced Critical Care 24(4). A widely cited clinical review of the phenomenon.
- Cvach, M. (2012). Monitor Alarm Fatigue: An Integrative Review. Biomedical Instrumentation & Technology 46(4).
- Egan, J. (1975). Signal Detection Theory and ROC Analysis. Academic Press. The foundational Neyman-Pearson framing referenced above.
What to learn next
- Choosing a threshold from costs — the decision-theoretic machinery this lesson builds on.
- Reliability diagrams and calibration error — whether the underlying probability is even trustworthy before a threshold is applied to it.
- Screening: one positive in a thousand scans — the same mathematics, applied to imaging recalls instead of bedside alarms.