Medical Imaging AI

Screening: one positive in a thousand scans

When real disease is rare, even a strong screening model sends mostly healthy people back for extra tests — and that math cannot be wished away.

On this page 6
  1. Why it exists
  2. How it works
  3. Where you have already seen it
  4. An honest warning
  5. Remember this
  6. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

In cancer screening, most flagged scans turn out fine.

Think about airport security scanning every bag for a dangerous item. Almost every bag is harmless. The scanner still has to check every single one, carefully, for the rare bag that is not.

Cancer screening scans work the same way. Almost every scan is normal. The model has to flag a suspicious few, knowing most flagged scans will turn out fine.

Why it exists

Screening programs scan huge numbers of healthy people to catch the rare early case. Catching it early can save a life. That is the entire point.

Rare, though, is the key word. If only about 1 scan in 1,000 has a real finding, the math turns brutal. Even a very good model must sort one true case from 999 healthy ones, using imperfect signals.

Recall is what happens when a scan gets flagged. The patient is called back for another look — maybe another scan, maybe a biopsy. Every recall costs something real: money, time, and real anxiety for the patient. A screening program has to accept a lot of unnecessary recalls to reliably catch the rare real ones.

How it works

1,000 scans  ->  1 real case, 999 healthy

Even a good model, catching 90% of real cases:
  - catches the 1 real case, most of the time
  - also flags some healthy scans as suspicious

Result: most people called back for a "recall" turn out to be fine

This is not a flaw specific to any one model. It is what a low base rate does to any screening test, however accurate.

Where you have already seen it

  • A home smoke detector goes off far more often for burnt toast than for a real fire. Real fires are rare, and the detector is tuned to catch them anyway.
  • A spam filter's "maybe spam" folder often holds mostly real mail. Most incoming mail is not spam in the first place.

An honest warning

A screening program is a genuinely regulated, high-stakes medical decision. Choosing thresholds, and deciding what counts as a "positive," involves clinical trials and regulatory review. It takes years of population-level evidence — not a threshold picked once and left alone. Nothing in this lesson's toy numbers should be read as real screening guidance.

Remember this

  • When a real disease is rare, most flagged results are false alarms, even from a strong model.
  • Every recall has a real cost — for the patient, and for the healthcare system.
  • Screening thresholds are a regulated clinical decision, not a number a model tunes on its own.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install pandas

Minimal runnable code

screening_math.py
import pandas as pd

# ILLUSTRATIVE numbers, not a specific published screening trial. Order of
# magnitude only: cancer screening finds roughly one true case per thousand
# scans, and every recall (a second scan, maybe a biopsy) costs real time,
# money, and anxiety.
scans_per_year = 100_000
prevalence = 1 / 1000
true_cases = scans_per_year * prevalence

def screening_outcomes(sensitivity, specificity):
    negatives = scans_per_year - true_cases
    true_recalls = true_cases * sensitivity
    false_recalls = negatives * (1 - specificity)
    total_recalls = true_recalls + false_recalls
    ppv = true_recalls / total_recalls
    scans_per_real_case = total_recalls / true_recalls
    return total_recalls, ppv, scans_per_real_case

rows = []
for specificity in [0.90, 0.95, 0.98, 0.995]:
    total_recalls, ppv, scans_per_case = screening_outcomes(sensitivity=0.90, specificity=specificity)
    rows.append({
        "specificity": specificity,
        "recalls_per_year": round(total_recalls),
        "chance_a_recall_is_real": round(ppv, 3),
        "recalls_needed_per_real_case": round(scans_per_case, 1),
    })

print(f"true cancers in {scans_per_year:,} scans at this prevalence: {true_cases:.0f}")
print(pd.DataFrame(rows).to_string(index=False))
Output
true cancers in 100,000 scans at this prevalence: 100
 specificity  recalls_per_year  chance_a_recall_is_real  recalls_needed_per_real_case
       0.900             10080                    0.009                         112.0
       0.950              5085                    0.018                          56.5
       0.980              2088                    0.043                          23.2
       0.995               590                    0.153                           6.6

What actually happened

At 90% specificity — which sounds reasonably good — only 0.9% of recalls turn out to be a real case. That means about 112 recalls happen for every single true cancer caught. Even at 99.5% specificity, still a demanding bar, most recalls remain false: only 15.3% turn out real.

This is not a failure of the model. Sensitivity stays fixed at 90% throughout the whole table — the model is catching the same share of real cases every time. The share of recalls that are real moves almost entirely because of how rare the disease is, not because of how good the model is.

Line by line, the parts that are not obvious:

  • prevalence = 1 / 1000 is the single number that drives this whole table. Change it, and every other number shifts, even with the model held perfectly constant.
  • chance_a_recall_is_real is the positive predictive value (PPV) — of the people called back, what share actually have the disease. It is a different question from sensitivity, which only asks about people who truly have the disease.
  • Specificity is the only thing varying across rows here. Small specificity improvements produce large drops in unnecessary recalls — this is the single highest-leverage number in the whole table.

Common mistakes

Reporting sensitivity and specificity without also reporting PPV at the real-world prevalence. Both numbers can look excellent in isolation, while still describing a program that recalls over a hundred people per real case found.

Assuming a screening model's accuracy transfers between populations. A model tuned for a population with 1-in-1,000 prevalence performs differently at 1-in-100, even with identical sensitivity and specificity — see why your model fails at the next hospital.

Treating "reduce false recalls" as a purely technical problem. It is also a threshold decision, weighing the cost of unnecessary biopsies against the cost of a missed cancer — a decision for clinicians and regulators, not an engineering team alone.

Try it yourself

Change prevalence to 1 / 100, simulating a higher-risk screened population. Rerun the table and compare how much the chance_a_recall_is_real column improves, with the model itself completely unchanged.

What to learn next

Researcher — Mathematics and papers.

PPV as a function of prevalence

From Bayes' theorem, positive predictive value depends on prevalence, sensitivity, and specificity jointly:

PPV = (pi * Se) / (pi * Se + (1 - pi) * (1 - Sp))
  • pi — prevalence, the base rate of true disease in the screened population
  • Se — sensitivity
  • Sp — specificity

Holding Se and Sp fixed, PPV -> 0 as pi -> 0. No realistic classifier escapes this: the (1-pi) term dominates the denominator as disease becomes rare, regardless of how high Sp is set (short of the degenerate case Sp = 1). This is the same base-rate mathematics covered in false alarms and alert fatigue, applied here to imaging recalls specifically.

Operating point selection under extreme imbalance

A screening classifier's ROC curve describes the full sensitivity/specificity trade-off, but choosing an operating point under a prevalence this low requires optimising an explicit cost function, not a symmetric metric like accuracy or F1:

Expected cost = pi*(1-Se)*C_miss + (1-pi)*(1-Sp)*C_recall
  • C_miss — the cost of a missed true case (delayed diagnosis, worse prognosis)
  • C_recall — the cost of an unnecessary recall (biopsy risk, patient anxiety, system cost)

Because C_miss for a missed cancer is generally judged far higher than C_recall for one unnecessary follow-up, the optimal operating point sits at a sensitivity well above what a symmetric-cost analysis would choose — which is why screening programs typically favour high sensitivity, high recall-rate operating points over the balanced accuracy an unweighted metric would select. The exact ratio of C_miss to C_recall is a clinical and regulatory judgement, not a number a model can determine on its own.

AI as a second reader

A well-studied deployment pattern in screening mammography is AI-assisted double reading: an AI system flags cases for a second human read, or triages cases by suspicion level, rather than making the final call unassisted. This pattern is specifically designed to add sensitivity for the rare true-positive case without granting an unsupervised model final diagnostic authority — directly connected to the human-in-the-loop principle covered in validating a model before it touches patients. Published studies of this pattern vary meaningfully by population, imaging protocol, and study design; readers evaluating a specific claim should consult the primary study rather than a general summary.

Key references

  • Lehman, C. et al. (2015). Diagnostic Accuracy of Digital Screening Mammography. JAMA Internal Medicine. A large-scale study of real-world screening mammography performance.
  • Egan, J. (1975). Signal Detection Theory and ROC Analysis. Academic Press. The foundational cost-weighted decision framework referenced above.
  • McKinney, S. et al. (2020). International evaluation of an AI system for breast cancer screening. Nature 577. A widely discussed and subsequently debated study of an AI screening system evaluated across multiple populations.

Current state

AI-assisted screening remains an active area of regulatory evaluation, with results varying substantially by population, imaging equipment, and study design. Findings from one population or health system do not automatically generalise to another — a point covered in depth in why your model fails at the next hospital. Deployment of any screening AI system requires regulatory clearance and prospective clinical validation specific to the population it will serve.

What to learn next