Do your labels even agree?
Inter-annotator agreement checks whether two labellers actually agree beyond what chance alone would predict, which raw agreement cannot tell you.
- 8 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Inter-annotator agreement checks if two people labelling the same data actually agree, more than luck alone would explain.
Picture getting a second medical opinion. Two doctors look at the same scan. Do they actually agree on the diagnosis, or did they happen to land on the same guess by luck?
If a disease is rare, both doctors might say "healthy" almost every time, agreeing often purely by chance. Real agreement means agreeing more than that chance level would predict. That is what this metric checks.
Why it exists
Every metric earlier in this section compares a model against labels a human wrote. Those human labels are treated as ground truth, the trusted answer key.
That trust only holds if the labels are actually reliable. Give two labellers the same text. If they disagree constantly, your "ground truth" is not solid ground at all.
Checking human agreement first tells you whether your labels are worth trusting. Do this before spending any effort comparing a model against them.
How it works
Two people label 10 support tickets as spam or not spam.
Raw agreement: how often did they give the same label?
-> looks fine on its own
But: if 90% of tickets are "not spam" anyway,
two careless labellers could agree 90% of the time
by both guessing "not spam" every time, with no real thought.
Cohen's kappa asks a sharper question:
"How much better is this than guessing?"Cohen's kappa is the standard way to answer that sharper question. It compares actual agreement against the agreement two labellers would get purely by chance, given how common each label is.
Where you have already seen it
- Content moderation training data. Before training a spam or hate-speech classifier, teams check whether their human labellers agree on the training labels.
- Medical AI datasets. Diagnosis labels used to train medical models are checked for agreement between the doctors who provided them.
- Any crowdsourced labelling project. Platforms hiring multiple workers to label the same data use agreement scores to catch unreliable data before it reaches a model.
Remember this
- Raw agreement can look high purely by chance, especially when one label is far more common.
- Cohen's kappa corrects for that chance, giving a fairer number.
- Low agreement between human labellers means your "ground truth" needs fixing before anything else.
What to learn next
- Using an LLM to grade text — checking whether an AI judge agrees with humans, using this exact idea.
- Hypothesis testing — the broader statistical toolkit this metric borrows its logic from.
- Class weights — a related fix for the same "one label is far more common" problem, applied during training.
Developer — Code and libraries.
This example shows the exact scenario that makes raw agreement misleading: two datasets with identical raw agreement, but very different real reliability.
Setup
pip install scikit-learnSame raw agreement, very different kappa
from sklearn.metrics import cohen_kappa_score
# Two annotators labelling 10 support tickets as spam (1) or not spam (0)
annotator_a = [1, 0, 0, 1, 0, 1, 0, 0, 1, 0]
annotator_b = [1, 0, 1, 1, 0, 1, 0, 1, 1, 0]
raw_agreement = sum(a == b for a, b in zip(annotator_a, annotator_b)) / len(annotator_a)
kappa = cohen_kappa_score(annotator_a, annotator_b)
print(f"raw agreement = {raw_agreement:.2f} Cohen's kappa = {kappa:.2f}")
# Now a lopsided dataset: 9 of 10 tickets are genuinely "not spam"
annotator_a_lopsided = [0] * 9 + [1]
annotator_b_lopsided = [0] * 8 + [1, 0] # one disagreement
raw2 = sum(a == b for a, b in zip(annotator_a_lopsided, annotator_b_lopsided)) / 10
kappa2 = cohen_kappa_score(annotator_a_lopsided, annotator_b_lopsided)
print(f"raw agreement = {raw2:.2f} Cohen's kappa = {kappa2:.2f} (lopsided data)")raw agreement = 0.80 Cohen's kappa = 0.62 raw agreement = 0.80 Cohen's kappa = -0.11 (lopsided data)
Both datasets show 80% raw agreement. Look at the kappa scores: 0.62 versus -0.11. The second pair of labellers is doing worse than random guessing once you account for how easy it is to agree on a mostly one-sided dataset. Raw agreement completely hid this.
Line by line
cohen_kappa_score needs only the two label lists, in matching order. It works out the expected chance agreement internally, from how often each label appears in the two lists.
A negative kappa is a real, meaningful result. It means the two labellers agree less than chance would predict, a warning sign that something about the labelling task, or the labellers themselves, is systematically broken.
Common mistakes
Reporting only raw agreement. As shown above, this can hide a genuinely broken labelling process behind a comfortable-looking percentage.
Treating kappa's benchmark thresholds as universal law. Guidelines exist, commonly citing above 0.6 as "substantial", but these are rules of thumb from one specific field, not a fixed cutoff. Context and task difficulty matter.
Computing kappa on too few examples. With only 10 tickets, one different judgement swings the score by a lot. Kappa needs a reasonably sized sample to be a stable estimate, not a large-looking one alone.
Try it yourself
Change annotator_b_lopsided so it disagrees on two tickets instead of one. Rerun and watch how far kappa moves for a single extra disagreement, on a lopsided dataset.
That sensitivity is the whole point: on data where one label dominates, a small handful of genuine disagreements can be the difference between a strong-looking kappa and a negative one.
What to learn next
- Hypothesis testing — the general framework this kind of "better than chance" reasoning comes from.
- Class weights — handling the same one-label-dominates problem during model training.
- Using an LLM to grade text — where this metric gets reused to validate an AI judge.
Researcher — Mathematics and papers.
Cohen's kappa
kappa = (p_o - p_e) / (1 - p_e)p_ois observed agreement: the fraction of items where both annotators gave the same label.p_eis expected agreement by chance, computed from each annotator's marginal label frequencies:p_e = sum over classes c of P_a(c) * P_b(c), whereP_a(c)andP_b(c)are the fraction of items each annotator labelledc.kappa = 1means perfect agreement,kappa = 0means agreement matches chance exactly, and negative values mean agreement is worse than chance.
Cohen (1960) introduced this correction specifically because raw percent agreement is not comparable across tasks with different label-frequency distributions, exactly the failure mode demonstrated in the developer block.
Extensions beyond two annotators and nominal labels
Fleiss' kappa (Fleiss, 1971) extends the same chance-correction idea to more than two annotators, using average pairwise agreement rather than a single pair.
Weighted kappa (Cohen, 1968) applies to ordinal labels, where disagreements of different sizes should be penalised differently. A 1-star versus 2-star disagreement should count as less severe than 1-star versus 5-star, a distinction unweighted kappa cannot express.
Krippendorff's alpha (Krippendorff, 1980) generalises further still, handling missing data, more than two annotators, and multiple measurement scales (nominal, ordinal, interval) within one unified framework. It is increasingly preferred in NLP annotation work specifically because real annotation projects rarely have every annotator label every item.
Interpreting the value
Landis & Koch (1977) proposed the widely cited benchmark table (0.0–0.2 slight, 0.2–0.4 fair, 0.4–0.6 moderate, 0.6–0.8 substantial, 0.8–1.0 almost perfect). This table is a convention from one specific paper in one specific field, not a statistically derived boundary. Applying it blindly across unrelated tasks is a documented point of criticism in the annotation literature.
What counts as "good enough" agreement genuinely depends on task difficulty. Sentiment labelling on clear-cut movie reviews should reach high kappa. Labelling subtle legal clause categories, where trained lawyers themselves disagree, may never reach the same bar. A lower kappa there is not automatically a failure.
Complexity
Computing kappa is O(n) in the number of labelled items, dominated by tallying a confusion matrix between the two annotators' label distributions. Statistical significance and confidence intervals for kappa are typically obtained via bootstrap resampling rather than a closed-form test, since the sampling distribution of kappa is not simple even under standard assumptions.
Key references
- Cohen, J. (1960). A Coefficient of Agreement for Nominal Scales. Educational and Psychological Measurement.
- Cohen, J. (1968). Weighted Kappa: Nominal Scale Agreement with Provision for Scaled Disagreement or Partial Credit. Psychological Bulletin.
- Fleiss, J. L. (1971). Measuring Nominal Scale Agreement Among Many Raters. Psychological Bulletin.
- Krippendorff, K. (1980). Content Analysis: An Introduction to Its Methodology. Sage Publications.
- Landis, J. R. & Koch, G. G. (1977). The Measurement of Observer Agreement for Categorical Data. Biometrics.
Current state and open problems
Agreement statistics remain a required step before trusting any human-labelled dataset, in NLP and well beyond it. Reviewers at major venues now routinely expect a reported kappa or alpha for any newly introduced annotated dataset.
The open problem is what to do once agreement is measured and found low. Low agreement can mean unclear guidelines, an underspecified task, or genuine ambiguity in the underlying phenomenon, such as sarcasm or subjective offensiveness, where reasonable humans truly do read the same text differently. Distinguishing "fixable annotation problem" from "genuinely ambiguous task" from agreement statistics alone remains more art than science, and typically needs a manual review of the disagreement cases themselves.
What to learn next
- Using an LLM to grade text — validating an AI judge against humans with these same statistics.
- Hypothesis testing — the general statistical machinery kappa is one instance of.
- Evaluating a legal AI system — this same idea, applied where labelling disagreement carries legal consequences.