AI in Healthcare

False alarms and alert fatigue

A monitor that cries wolf hundreds of times a day trains the humans around it to stop listening, even to the one alarm that matters.

On this page 6
  1. Why it exists
  2. How it works
  3. Where you have already seen it
  4. An honest warning
  5. Remember this
  6. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Alert fatigue means an alarm fires so often that people stop reacting to it.

Think of a car alarm on a busy street. The first time it goes off, everyone looks up. By the tenth false trigger that week, nobody looks up anymore — not even during a real theft.

Hospital monitors can fall into the exact same trap. A machine that beeps at the smallest twitch trains the nurses near it to tune the beeping out. That is the same way a street tunes out a car alarm.

Why it exists

Building an alarm that never misses a real problem is not hard. Set the threshold low enough, and it catches almost everything. The cost is a flood of false alarms, because a low threshold also catches every minor, harmless blip.

The reverse is equally easy. Set the threshold high, and false alarms nearly disappear. The cost is missed real events — a high threshold also filters out the early, subtle version of a real problem.

There is no setting that avoids both costs. Most systems get tuned only to "catch everything." That choice picks the first failure mode, often without anyone deciding to.

How it works

Low threshold:    catches almost every real problem   +  huge number of false alarms
                                                              |
                                                              v
                                                     staff learn to ignore alarms
                                                              |
                                                              v
                                                   the one real alarm gets missed too

The danger is not the false alarms by themselves. It is what constant false alarms do to how humans respond to every alarm, including the true ones.

Where you have already seen it

  • Phone notifications. An app that pings constantly gets its notifications silenced entirely — including the one that mattered.
  • A smoke detector that goes off every time you cook. People start disabling it, or moving it, rather than fixing the underlying sensitivity.

An honest warning

A model with genuinely excellent accuracy numbers can still generate an unusable flood of false alarms. Rare events, multiplied across thousands of daily checks, still add up to a lot of noise. The next section works through exactly why, with real numbers.

Remember this

  • The danger of a noisy alarm system is not the false alarms themselves. It is what they train humans to ignore.
  • There is no threshold that eliminates both false alarms and missed events at the same time.
  • A system's accuracy on paper can look fine while still generating an unworkable number of alarms per shift.
  • Choosing where to set that threshold is a clinical decision, not a modelling one. It needs a hospital's own clinicians involved before deployment, not a default left in by a vendor.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install pandas

Minimal runnable code

The scenario: a 20-bed ward, monitored continuously.

alert_fatigue.py
import pandas as pd

# A 20-bed ward, checked every 5 minutes: 288 checks per bed per day.
beds = 20
checks_per_bed_per_day = 288
total_checks = beds * checks_per_bed_per_day

# ILLUSTRATIVE prevalence, not a measured statistic: about 1 in 500 checks
# catches a patient who is genuinely about to decline in the next hour.
prevalence = 1 / 500

def alarms_per_day(sensitivity, specificity):
    positives = total_checks * prevalence
    negatives = total_checks - positives
    true_alarms = positives * sensitivity
    false_alarms = negatives * (1 - specificity)
    total_alarms = true_alarms + false_alarms
    ppv = true_alarms / total_alarms
    return total_alarms, ppv

rows = []
for specificity in [0.90, 0.95, 0.99, 0.999]:
    total_alarms, ppv = alarms_per_day(sensitivity=0.90, specificity=specificity)
    rows.append({
        "specificity": specificity,
        "alarms_per_day": round(total_alarms, 1),
        "real_alarm_chance": round(ppv, 3),
    })

print(pd.DataFrame(rows).to_string(index=False))
Output
 specificity  alarms_per_day  real_alarm_chance
       0.900           585.2              0.018
       0.950           297.8              0.035
       0.990            67.9              0.153
       0.999            16.1              0.643

What actually happened

At 90% specificity — a number that sounds respectable — this ward gets 585 alarms a day. Divide that across a 24-hour shift and it is an alarm roughly every two and a half minutes, on every bed combined. Only 1.8% of those alarms are real.

Even at 99% specificity, which most people would call an excellent model, more than 5 out of every 6 alarms are still false. Only pushing specificity to 99.9% gets the real-alarm chance above even coin-flip odds.

Line by line, the parts that are not obvious:

  • total_checks scales with how often a check happens, not only with how many patients there are. Frequent monitoring multiplies the false-alarm opportunity, even if the model itself never changes.
  • real_alarm_chance here is the positive predictive value (PPV) — of the alarms that fire, what fraction are true. It is not the same number as sensitivity, and it degrades fast as the real event becomes rarer.
  • Sensitivity is held fixed at 90% throughout this table. Only specificity moves — showing that specificity, not sensitivity, is usually the lever that actually controls alarm burden.

Common mistakes

Reporting sensitivity and specificity without also reporting expected alarms per day. Both numbers can look excellent in isolation and still describe an unusable system once multiplied by how often checks happen.

Assuming higher specificity is free. Pushing specificity up almost always costs some sensitivity — the real, missed-event trade-off this table hides by holding sensitivity fixed.

Designing the alarm threshold around the model's best F1 score. F1 balances precision and recall as if they cost the same. A missed deterioration and an extra beep on a monitor do not cost the same, and the threshold should reflect that.

Try it yourself

Change checks_per_bed_per_day to 12 — simulating hourly checks instead of every five minutes. Rerun the table and notice how much the alarm burden drops, even with identical sensitivity and specificity. Monitoring frequency is a design lever too, not only the model.

What to learn next

Researcher — Mathematics and papers.

Formalising alarm burden

For a monitoring system running N checks over some period, with prevalence π, sensitivity Se, and specificity Sp, the expected number of alarms is:

E[alarms] = N*π*Se + N*(1-π)*(1-Sp)
  • N — total number of checks in the period
  • π — prevalence, the base rate of the real event per check
  • Se — sensitivity (true positive rate)
  • Sp — specificity (true negative rate)

The positive predictive value follows directly from Bayes' theorem:

PPV = (π * Se) / (π * Se + (1-π) * (1-Sp))

As π -> 0, PPV -> 0 regardless of how high Se and Sp are, provided Sp < 1. This is the base-rate fallacy made quantitative: no realistic classifier avoids a low PPV against a sufficiently rare event, because the (1-Sp) term is multiplied by a much larger (1-π) population of negatives.

The alarm-rate/coverage trade-off

Treating this as a decision problem, define a cost ratio C_FA / C_miss — the cost of one false alarm relative to the cost of one missed event. The Neyman-Pearson lemma implies the optimal decision rule is a likelihood-ratio threshold set by that cost ratio, connecting directly to choosing a threshold from costs. In practice, C_FA is rarely a fixed number — it grows nonlinearly as alarm frequency increases, because each additional alarm degrades the credibility of every alarm around it. This nonlinearity is what standard threshold-selection frameworks, which assume a fixed per-alarm cost, do not capture.

Documented consequences

Alarm fatigue has been named a top patient-safety technology hazard in recurring hazard reports from clinical safety organisations (e.g. the ECRI Institute's annual health technology hazard lists), and has been the subject of Joint Commission Sentinel Event Alerts in the United States addressing clinical alarm safety directly. Readers should consult those primary sources for current, verified figures rather than any restated number here — the specific alarm counts and harm statistics vary by report and by year, and are outside what this lesson can responsibly assert.

Cost

Reducing alarm burden without losing sensitivity generally requires either (a) a better-discriminating model — raising both Se and Sp simultaneously, which is the hard, slow route — or (b) smarter alarm logic: delay-and-confirm rules, trend-based alarms instead of instantaneous threshold crossings, and alarm suppression during known noisy periods (e.g. patient repositioning). These are engineering interventions on top of the model, not substitutes for a better model.

Key references

  • Sendelbach, S. & Funk, M. (2013). Alarm Fatigue: A Patient Safety Concern. AACN Advanced Critical Care 24(4). A widely cited clinical review of the phenomenon.
  • Cvach, M. (2012). Monitor Alarm Fatigue: An Integrative Review. Biomedical Instrumentation & Technology 46(4).
  • Egan, J. (1975). Signal Detection Theory and ROC Analysis. Academic Press. The foundational Neyman-Pearson framing referenced above.

What to learn next