When the radiologists disagree
The label a model trains on is not an objective fact — it is one expert's read, and different experts reading the same scan often disagree.
- 9 min read
- 3 reading levels
- Published
Read these first
On this page 6
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Two radiologists can disagree about the same scan — and one of them still becomes the "correct" label.
Think about a football match, and a foul near the goal. Three fans watching the same replay can genuinely disagree about whether it was a penalty. All three watched carefully. The moment was genuinely hard to call.
Reading a medical scan works the same way. A shadow that one radiologist calls "probably nothing" another calls "worth a closer look." Both are experienced. The image itself was ambiguous.
Why it exists
Machine learning needs a label to learn from. For a photo of a cat, that label is easy — almost everyone agrees it is a cat.
A medical scan is often not that simple. Whether a nodule looks suspicious. Whether a fracture is present. Whether a border counts as one lesion or two. Each of these calls involves real, defensible expert judgement — not a single obvious fact.
That means the "ground truth" a model trains on is not truth in the way a photo caption is. It is a recorded human opinion, formed by someone skilled but fallible, looking at a genuinely hard case.
How it works
Same scan -> Radiologist A: "suspicious"
-> Radiologist B: "probably benign"
-> Radiologist C: "suspicious"
|
v
majority vote becomes
the training labelUsing several readers and taking a majority vote is the standard fix. It does not remove disagreement — it decides how to handle it before training starts.
Where you have already seen it
- A second medical opinion exists precisely because one expert's read is not treated as automatically final.
- Sports video review by multiple officials reflects the same idea — one trained eye is not always enough.
An honest warning
A model trained against one radiologist's labels can only ever be as consistent as that radiologist was. If the label itself is shaky, matching it perfectly does not make a model right. It may have learned one person's particular habits instead.
This is a real limit, not a bug to be patched away. Every scan-labelling project has to decide how to handle it, and no single choice is free of trade-offs.
A model built on disagreement-prone labels needs domain-expert review before real deployment. A good score against one reader is not enough.
Remember this
- A label on a medical scan is a recorded expert opinion, not an objective, obvious fact.
- Different qualified radiologists reading the same scan can genuinely disagree.
- A majority vote across several readers builds a sturdier label. It does not remove disagreement entirely.
What to learn next
- Model evaluation — the general framework this lesson complicates with a genuinely uncertain label.
- Finding label errors in an image dataset — practical tools for catching labels worth a second look.
- Screening: one positive in a thousand scans — what happens next, once a label has been decided.
Developer — Code and libraries.
Setup
pip install numpy scikit-learnMinimal runnable code
import numpy as np
from sklearn.metrics import cohen_kappa_score
rng = np.random.default_rng(1)
n_scans = 300
# A hidden true state, unknown to anyone in real life. We only get to see it
# here because this is a synthetic teaching example.
true_finding = rng.binomial(1, 0.25, n_scans)
def noisy_reader(true, agreement):
"""A radiologist who agrees with the true finding `agreement` of the
time, and reads at random the rest of the time -- a stand-in for
genuine diagnostic difficulty, not carelessness."""
coin = rng.binomial(1, agreement, n_scans)
random_read = rng.binomial(1, 0.25, n_scans)
return np.where(coin == 1, true, random_read)
reader_a = noisy_reader(true_finding, agreement=0.85)
reader_b = noisy_reader(true_finding, agreement=0.80)
reader_c = noisy_reader(true_finding, agreement=0.75)
print("agreement between pairs of readers (Cohen's kappa, 1.0 = perfect):")
print(f" A vs B: {cohen_kappa_score(reader_a, reader_b):.3f}")
print(f" A vs C: {cohen_kappa_score(reader_a, reader_c):.3f}")
print(f" B vs C: {cohen_kappa_score(reader_b, reader_c):.3f}")
# The usual fix: the training label is a majority vote, not any one reader.
votes = reader_a + reader_b + reader_c
consensus = (votes >= 2).astype(int)
print()
print(f"cases where all three readers agreed: {((reader_a == reader_b) & (reader_b == reader_c)).mean():.1%}")
print(f"consensus label matches the (synthetic) true finding: "
f"{(consensus == true_finding).mean():.1%} of the time")agreement between pairs of readers (Cohen's kappa, 1.0 = perfect): A vs B: 0.688 A vs C: 0.689 B vs C: 0.652 cases where all three readers agreed: 80.7% consensus label matches the (synthetic) true finding: 99.0% of the time
What actually happened
The kappa scores, all around 0.65 to 0.69, land in the range textbooks call "substantial agreement" — solidly good for real diagnostic work, and still far from perfect. Even skilled, careful readers disagree on close to a fifth of cases in this synthetic example.
Now look at the consensus label: 99.0% accurate against the true finding, despite no individual reader being anywhere near that reliable alone. Combining several imperfect opinions, even by plain majority vote, produces a far sturdier label than trusting any one of them.
Line by line, the parts that are not obvious:
cohen_kappa_scoremeasures agreement beyond what random chance alone would produce — a raw percent-agreement number, by contrast, can look deceptively high purely because one finding is common.noisy_readerdeliberately never disagrees with the truth on purpose — every disagreement here comes from genuine ambiguity, exactly like real inter-reader variability, not from carelessness.votes >= 2implements a simple majority vote across three readers — the most common consensus method, though far from the only one used in practice.
Common mistakes
Treating one radiologist's read as ground truth without saying so. Many published imaging datasets were labelled by a single reader per case — a real limitation worth checking for, not assuming away.
Reporting only overall accuracy against a set of labels. A model can look excellent by matching easy, high-agreement cases, while quietly failing on exactly the hard, low-agreement ones that matter most clinically.
Assuming a majority vote of three readers is automatically enough. More readers generally help, but the right number depends on how ambiguous the task actually is — three is a convention, not a law.
Try it yourself
Lower every reader's agreement value by 0.1, simulating a genuinely harder diagnostic task. Rerun and watch both the pairwise kappa scores and the consensus accuracy drop — a concrete demonstration of how task difficulty flows all the way through to label quality.
What to learn next
- Finding label errors in an image dataset — practical detection tools that build on this idea.
- Writing annotation guidelines — reducing disagreement before it happens, not only after.
- Model evaluation — the general evaluation machinery this lesson complicates.
Researcher — Mathematics and papers.
Cohen's kappa
For two raters classifying n items into categories, Cohen's kappa corrects observed agreement for the agreement expected by chance alone:
kappa = (p_o - p_e) / (1 - p_e)p_o— observed proportion of agreement between the two ratersp_e— expected proportion of agreement under independence, computed from each rater's marginal label distribution
kappa = 1 indicates perfect agreement; kappa = 0 indicates agreement no better than chance. Landis & Koch's (1977) commonly cited interpretive scale treats 0.61-0.80 as "substantial" agreement and 0.81-1.00 as "almost perfect" — a scale used descriptively, not as a formally derived statistical threshold. For more than two raters, Fleiss' kappa generalises the same correction-for-chance logic across an arbitrary number of raters.
STAPLE: combining multiple noisy segmentations
For per-pixel or per-voxel segmentation labels rather than single categorical calls, majority voting is a special case of a more general problem: estimating the true segmentation from several imperfect ones. STAPLE (Warfield, Zou & Wells, 2004) treats each rater's segmentation as a noisy observation of an unknown true segmentation, and uses expectation-maximisation to jointly estimate the true label map and each rater's sensitivity and specificity:
E-step: estimate P(true label at voxel v | all raters' labels, current sensitivity/specificity estimates)
M-step: re-estimate each rater's sensitivity and specificity, given the current label estimateThis iterates to convergence, producing both a consensus segmentation and a per-rater reliability estimate — strictly more informative than an unweighted majority vote, at the cost of requiring several complete segmentations of the same cases.
Documented inter-rater variability
Elmore et al. (1994, NEJM) is a widely cited study of inter-radiologist variability in mammography screening interpretation, documenting substantial disagreement among experienced readers evaluating the same films. It remains one of the foundational citations establishing that reader variability in imaging is a genuine, persistent phenomenon rather than a rare edge case — directly relevant to why any imaging dataset's labels deserve scrutiny before being treated as ground truth.
Soft labels and learning from disagreement
Rather than collapsing multiple readers into one hard label, some approaches train directly against the label distribution — treating disagreement itself as signal rather than noise to be voted away:
y_soft = [ fraction of readers choosing class 1, fraction choosing class 2, ... ]Training against y_soft with a distributional loss (e.g. KL divergence to the predicted distribution) lets a model express calibrated uncertainty on genuinely ambiguous cases, rather than being forced toward false confidence by a majority-vote label — an approach with growing adoption in recent medical imaging literature, though it is not yet the default in most production pipelines.
Key references
- Cohen, J. (1960). A Coefficient of Agreement for Nominal Scales. Educational and Psychological Measurement 20(1).
- Landis, J. & Koch, G. (1977). The Measurement of Observer Agreement for Categorical Data. Biometrics 33(1).
- Warfield, S., Zou, K. & Wells, W. (2004). Simultaneous Truth and Performance Level Estimation (STAPLE). IEEE Transactions on Medical Imaging 23(7).
- Elmore, J. et al. (1994). Variability in Radiologists' Interpretations of Mammograms. New England Journal of Medicine 331.
Current state
Large public imaging benchmarks increasingly release multi-reader labels rather than a single consensus value, explicitly to let researchers model disagreement rather than treat it as noise. Standard evaluation practice, however, still mostly reports performance against a single consensus label — a known, actively discussed limitation of how imaging models are typically benchmarked.
What to learn next
- Finding label errors in an image dataset — systematic detection of exactly this kind of labelling problem.
- Reliability diagrams and calibration error — whether a model's own confidence reflects genuine uncertainty, of the kind shown here.
- Segmenting 3D scans — where this same disagreement problem appears at the pixel level, not only the case level.