Medical Imaging AI

When the radiologists disagree

The label a model trains on is not an objective fact — it is one expert's read, and different experts reading the same scan often disagree.

Read these first

On this page 6
  1. Why it exists
  2. How it works
  3. Where you have already seen it
  4. An honest warning
  5. Remember this
  6. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Two radiologists can disagree about the same scan — and one of them still becomes the "correct" label.

Think about a football match, and a foul near the goal. Three fans watching the same replay can genuinely disagree about whether it was a penalty. All three watched carefully. The moment was genuinely hard to call.

Reading a medical scan works the same way. A shadow that one radiologist calls "probably nothing" another calls "worth a closer look." Both are experienced. The image itself was ambiguous.

Why it exists

Machine learning needs a label to learn from. For a photo of a cat, that label is easy — almost everyone agrees it is a cat.

A medical scan is often not that simple. Whether a nodule looks suspicious. Whether a fracture is present. Whether a border counts as one lesion or two. Each of these calls involves real, defensible expert judgement — not a single obvious fact.

That means the "ground truth" a model trains on is not truth in the way a photo caption is. It is a recorded human opinion, formed by someone skilled but fallible, looking at a genuinely hard case.

How it works

Same scan  ->  Radiologist A: "suspicious"
           ->  Radiologist B: "probably benign"
           ->  Radiologist C: "suspicious"
                       |
                       v
              majority vote becomes
              the training label

Using several readers and taking a majority vote is the standard fix. It does not remove disagreement — it decides how to handle it before training starts.

Where you have already seen it

  • A second medical opinion exists precisely because one expert's read is not treated as automatically final.
  • Sports video review by multiple officials reflects the same idea — one trained eye is not always enough.

An honest warning

A model trained against one radiologist's labels can only ever be as consistent as that radiologist was. If the label itself is shaky, matching it perfectly does not make a model right. It may have learned one person's particular habits instead.

This is a real limit, not a bug to be patched away. Every scan-labelling project has to decide how to handle it, and no single choice is free of trade-offs.

A model built on disagreement-prone labels needs domain-expert review before real deployment. A good score against one reader is not enough.

Remember this

  • A label on a medical scan is a recorded expert opinion, not an objective, obvious fact.
  • Different qualified radiologists reading the same scan can genuinely disagree.
  • A majority vote across several readers builds a sturdier label. It does not remove disagreement entirely.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install numpy scikit-learn

Minimal runnable code

reader_agreement.py
import numpy as np
from sklearn.metrics import cohen_kappa_score

rng = np.random.default_rng(1)
n_scans = 300

# A hidden true state, unknown to anyone in real life. We only get to see it
# here because this is a synthetic teaching example.
true_finding = rng.binomial(1, 0.25, n_scans)

def noisy_reader(true, agreement):
    """A radiologist who agrees with the true finding `agreement` of the
    time, and reads at random the rest of the time -- a stand-in for
    genuine diagnostic difficulty, not carelessness."""
    coin = rng.binomial(1, agreement, n_scans)
    random_read = rng.binomial(1, 0.25, n_scans)
    return np.where(coin == 1, true, random_read)

reader_a = noisy_reader(true_finding, agreement=0.85)
reader_b = noisy_reader(true_finding, agreement=0.80)
reader_c = noisy_reader(true_finding, agreement=0.75)

print("agreement between pairs of readers (Cohen's kappa, 1.0 = perfect):")
print(f"  A vs B: {cohen_kappa_score(reader_a, reader_b):.3f}")
print(f"  A vs C: {cohen_kappa_score(reader_a, reader_c):.3f}")
print(f"  B vs C: {cohen_kappa_score(reader_b, reader_c):.3f}")

# The usual fix: the training label is a majority vote, not any one reader.
votes = reader_a + reader_b + reader_c
consensus = (votes >= 2).astype(int)
print()
print(f"cases where all three readers agreed: {((reader_a == reader_b) & (reader_b == reader_c)).mean():.1%}")
print(f"consensus label matches the (synthetic) true finding: "
      f"{(consensus == true_finding).mean():.1%} of the time")
Output
agreement between pairs of readers (Cohen's kappa, 1.0 = perfect):
  A vs B: 0.688
  A vs C: 0.689
  B vs C: 0.652

cases where all three readers agreed: 80.7%
consensus label matches the (synthetic) true finding: 99.0% of the time

What actually happened

The kappa scores, all around 0.65 to 0.69, land in the range textbooks call "substantial agreement" — solidly good for real diagnostic work, and still far from perfect. Even skilled, careful readers disagree on close to a fifth of cases in this synthetic example.

Now look at the consensus label: 99.0% accurate against the true finding, despite no individual reader being anywhere near that reliable alone. Combining several imperfect opinions, even by plain majority vote, produces a far sturdier label than trusting any one of them.

Line by line, the parts that are not obvious:

  • cohen_kappa_score measures agreement beyond what random chance alone would produce — a raw percent-agreement number, by contrast, can look deceptively high purely because one finding is common.
  • noisy_reader deliberately never disagrees with the truth on purpose — every disagreement here comes from genuine ambiguity, exactly like real inter-reader variability, not from carelessness.
  • votes >= 2 implements a simple majority vote across three readers — the most common consensus method, though far from the only one used in practice.

Common mistakes

Treating one radiologist's read as ground truth without saying so. Many published imaging datasets were labelled by a single reader per case — a real limitation worth checking for, not assuming away.

Reporting only overall accuracy against a set of labels. A model can look excellent by matching easy, high-agreement cases, while quietly failing on exactly the hard, low-agreement ones that matter most clinically.

Assuming a majority vote of three readers is automatically enough. More readers generally help, but the right number depends on how ambiguous the task actually is — three is a convention, not a law.

Try it yourself

Lower every reader's agreement value by 0.1, simulating a genuinely harder diagnostic task. Rerun and watch both the pairwise kappa scores and the consensus accuracy drop — a concrete demonstration of how task difficulty flows all the way through to label quality.

What to learn next

Researcher — Mathematics and papers.

Cohen's kappa

For two raters classifying n items into categories, Cohen's kappa corrects observed agreement for the agreement expected by chance alone:

kappa = (p_o - p_e) / (1 - p_e)
  • p_o — observed proportion of agreement between the two raters
  • p_e — expected proportion of agreement under independence, computed from each rater's marginal label distribution

kappa = 1 indicates perfect agreement; kappa = 0 indicates agreement no better than chance. Landis & Koch's (1977) commonly cited interpretive scale treats 0.61-0.80 as "substantial" agreement and 0.81-1.00 as "almost perfect" — a scale used descriptively, not as a formally derived statistical threshold. For more than two raters, Fleiss' kappa generalises the same correction-for-chance logic across an arbitrary number of raters.

STAPLE: combining multiple noisy segmentations

For per-pixel or per-voxel segmentation labels rather than single categorical calls, majority voting is a special case of a more general problem: estimating the true segmentation from several imperfect ones. STAPLE (Warfield, Zou & Wells, 2004) treats each rater's segmentation as a noisy observation of an unknown true segmentation, and uses expectation-maximisation to jointly estimate the true label map and each rater's sensitivity and specificity:

E-step:  estimate  P(true label at voxel v | all raters' labels, current sensitivity/specificity estimates)
M-step:  re-estimate each rater's sensitivity and specificity, given the current label estimate

This iterates to convergence, producing both a consensus segmentation and a per-rater reliability estimate — strictly more informative than an unweighted majority vote, at the cost of requiring several complete segmentations of the same cases.

Documented inter-rater variability

Elmore et al. (1994, NEJM) is a widely cited study of inter-radiologist variability in mammography screening interpretation, documenting substantial disagreement among experienced readers evaluating the same films. It remains one of the foundational citations establishing that reader variability in imaging is a genuine, persistent phenomenon rather than a rare edge case — directly relevant to why any imaging dataset's labels deserve scrutiny before being treated as ground truth.

Soft labels and learning from disagreement

Rather than collapsing multiple readers into one hard label, some approaches train directly against the label distribution — treating disagreement itself as signal rather than noise to be voted away:

y_soft = [ fraction of readers choosing class 1, fraction choosing class 2, ... ]

Training against y_soft with a distributional loss (e.g. KL divergence to the predicted distribution) lets a model express calibrated uncertainty on genuinely ambiguous cases, rather than being forced toward false confidence by a majority-vote label — an approach with growing adoption in recent medical imaging literature, though it is not yet the default in most production pipelines.

Key references

  • Cohen, J. (1960). A Coefficient of Agreement for Nominal Scales. Educational and Psychological Measurement 20(1).
  • Landis, J. & Koch, G. (1977). The Measurement of Observer Agreement for Categorical Data. Biometrics 33(1).
  • Warfield, S., Zou, K. & Wells, W. (2004). Simultaneous Truth and Performance Level Estimation (STAPLE). IEEE Transactions on Medical Imaging 23(7).
  • Elmore, J. et al. (1994). Variability in Radiologists' Interpretations of Mammograms. New England Journal of Medicine 331.

Current state

Large public imaging benchmarks increasingly release multi-reader labels rather than a single consensus value, explicitly to let researchers model disagreement rather than treat it as noise. Standard evaluation practice, however, still mostly reports performance against a single consensus label — a known, actively discussed limitation of how imaging models are typically benchmarked.

What to learn next