Evaluating Text Systems

Do your labels even agree?

Inter-annotator agreement checks whether two labellers actually agree beyond what chance alone would predict, which raw agreement cannot tell you.

Read these first

On this page 5
  1. Why it exists
  2. How it works
  3. Where you have already seen it
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Inter-annotator agreement checks if two people labelling the same data actually agree, more than luck alone would explain.

Picture getting a second medical opinion. Two doctors look at the same scan. Do they actually agree on the diagnosis, or did they happen to land on the same guess by luck?

If a disease is rare, both doctors might say "healthy" almost every time, agreeing often purely by chance. Real agreement means agreeing more than that chance level would predict. That is what this metric checks.

Why it exists

Every metric earlier in this section compares a model against labels a human wrote. Those human labels are treated as ground truth, the trusted answer key.

That trust only holds if the labels are actually reliable. Give two labellers the same text. If they disagree constantly, your "ground truth" is not solid ground at all.

Checking human agreement first tells you whether your labels are worth trusting. Do this before spending any effort comparing a model against them.

How it works

  Two people label 10 support tickets as spam or not spam.

  Raw agreement: how often did they give the same label?
  -> looks fine on its own

  But: if 90% of tickets are "not spam" anyway,
  two careless labellers could agree 90% of the time
  by both guessing "not spam" every time, with no real thought.

  Cohen's kappa asks a sharper question:
  "How much better is this than guessing?"

Cohen's kappa is the standard way to answer that sharper question. It compares actual agreement against the agreement two labellers would get purely by chance, given how common each label is.

Where you have already seen it

  • Content moderation training data. Before training a spam or hate-speech classifier, teams check whether their human labellers agree on the training labels.
  • Medical AI datasets. Diagnosis labels used to train medical models are checked for agreement between the doctors who provided them.
  • Any crowdsourced labelling project. Platforms hiring multiple workers to label the same data use agreement scores to catch unreliable data before it reaches a model.

Remember this

  • Raw agreement can look high purely by chance, especially when one label is far more common.
  • Cohen's kappa corrects for that chance, giving a fairer number.
  • Low agreement between human labellers means your "ground truth" needs fixing before anything else.

What to learn next

  • Using an LLM to grade text — checking whether an AI judge agrees with humans, using this exact idea.
  • Hypothesis testing — the broader statistical toolkit this metric borrows its logic from.
  • Class weights — a related fix for the same "one label is far more common" problem, applied during training.

Developer — Code and libraries.

This example shows the exact scenario that makes raw agreement misleading: two datasets with identical raw agreement, but very different real reliability.

Setup

bash
pip install scikit-learn

Same raw agreement, very different kappa

kappa_demo.py
from sklearn.metrics import cohen_kappa_score

# Two annotators labelling 10 support tickets as spam (1) or not spam (0)
annotator_a = [1, 0, 0, 1, 0, 1, 0, 0, 1, 0]
annotator_b = [1, 0, 1, 1, 0, 1, 0, 1, 1, 0]

raw_agreement = sum(a == b for a, b in zip(annotator_a, annotator_b)) / len(annotator_a)
kappa = cohen_kappa_score(annotator_a, annotator_b)
print(f"raw agreement = {raw_agreement:.2f}   Cohen's kappa = {kappa:.2f}")

# Now a lopsided dataset: 9 of 10 tickets are genuinely "not spam"
annotator_a_lopsided = [0] * 9 + [1]
annotator_b_lopsided = [0] * 8 + [1, 0]   # one disagreement

raw2 = sum(a == b for a, b in zip(annotator_a_lopsided, annotator_b_lopsided)) / 10
kappa2 = cohen_kappa_score(annotator_a_lopsided, annotator_b_lopsided)
print(f"raw agreement = {raw2:.2f}   Cohen's kappa = {kappa2:.2f}   (lopsided data)")
Output
raw agreement = 0.80   Cohen's kappa = 0.62
raw agreement = 0.80   Cohen's kappa = -0.11   (lopsided data)

Both datasets show 80% raw agreement. Look at the kappa scores: 0.62 versus -0.11. The second pair of labellers is doing worse than random guessing once you account for how easy it is to agree on a mostly one-sided dataset. Raw agreement completely hid this.

Line by line

cohen_kappa_score needs only the two label lists, in matching order. It works out the expected chance agreement internally, from how often each label appears in the two lists.

A negative kappa is a real, meaningful result. It means the two labellers agree less than chance would predict, a warning sign that something about the labelling task, or the labellers themselves, is systematically broken.

Common mistakes

Reporting only raw agreement. As shown above, this can hide a genuinely broken labelling process behind a comfortable-looking percentage.

Treating kappa's benchmark thresholds as universal law. Guidelines exist, commonly citing above 0.6 as "substantial", but these are rules of thumb from one specific field, not a fixed cutoff. Context and task difficulty matter.

Computing kappa on too few examples. With only 10 tickets, one different judgement swings the score by a lot. Kappa needs a reasonably sized sample to be a stable estimate, not a large-looking one alone.

Try it yourself

Change annotator_b_lopsided so it disagrees on two tickets instead of one. Rerun and watch how far kappa moves for a single extra disagreement, on a lopsided dataset.

That sensitivity is the whole point: on data where one label dominates, a small handful of genuine disagreements can be the difference between a strong-looking kappa and a negative one.

What to learn next

Researcher — Mathematics and papers.

Cohen's kappa

text
kappa = (p_o - p_e) / (1 - p_e)
  • p_o is observed agreement: the fraction of items where both annotators gave the same label.
  • p_e is expected agreement by chance, computed from each annotator's marginal label frequencies: p_e = sum over classes c of P_a(c) * P_b(c), where P_a(c) and P_b(c) are the fraction of items each annotator labelled c.
  • kappa = 1 means perfect agreement, kappa = 0 means agreement matches chance exactly, and negative values mean agreement is worse than chance.

Cohen (1960) introduced this correction specifically because raw percent agreement is not comparable across tasks with different label-frequency distributions, exactly the failure mode demonstrated in the developer block.

Extensions beyond two annotators and nominal labels

Fleiss' kappa (Fleiss, 1971) extends the same chance-correction idea to more than two annotators, using average pairwise agreement rather than a single pair.

Weighted kappa (Cohen, 1968) applies to ordinal labels, where disagreements of different sizes should be penalised differently. A 1-star versus 2-star disagreement should count as less severe than 1-star versus 5-star, a distinction unweighted kappa cannot express.

Krippendorff's alpha (Krippendorff, 1980) generalises further still, handling missing data, more than two annotators, and multiple measurement scales (nominal, ordinal, interval) within one unified framework. It is increasingly preferred in NLP annotation work specifically because real annotation projects rarely have every annotator label every item.

Interpreting the value

Landis & Koch (1977) proposed the widely cited benchmark table (0.0–0.2 slight, 0.2–0.4 fair, 0.4–0.6 moderate, 0.6–0.8 substantial, 0.8–1.0 almost perfect). This table is a convention from one specific paper in one specific field, not a statistically derived boundary. Applying it blindly across unrelated tasks is a documented point of criticism in the annotation literature.

What counts as "good enough" agreement genuinely depends on task difficulty. Sentiment labelling on clear-cut movie reviews should reach high kappa. Labelling subtle legal clause categories, where trained lawyers themselves disagree, may never reach the same bar. A lower kappa there is not automatically a failure.

Complexity

Computing kappa is O(n) in the number of labelled items, dominated by tallying a confusion matrix between the two annotators' label distributions. Statistical significance and confidence intervals for kappa are typically obtained via bootstrap resampling rather than a closed-form test, since the sampling distribution of kappa is not simple even under standard assumptions.

Key references

  • Cohen, J. (1960). A Coefficient of Agreement for Nominal Scales. Educational and Psychological Measurement.
  • Cohen, J. (1968). Weighted Kappa: Nominal Scale Agreement with Provision for Scaled Disagreement or Partial Credit. Psychological Bulletin.
  • Fleiss, J. L. (1971). Measuring Nominal Scale Agreement Among Many Raters. Psychological Bulletin.
  • Krippendorff, K. (1980). Content Analysis: An Introduction to Its Methodology. Sage Publications.
  • Landis, J. R. & Koch, G. G. (1977). The Measurement of Observer Agreement for Categorical Data. Biometrics.

Current state and open problems

Agreement statistics remain a required step before trusting any human-labelled dataset, in NLP and well beyond it. Reviewers at major venues now routinely expect a reported kappa or alpha for any newly introduced annotated dataset.

The open problem is what to do once agreement is measured and found low. Low agreement can mean unclear guidelines, an underspecified task, or genuine ambiguity in the underlying phenomenon, such as sarcasm or subjective offensiveness, where reasonable humans truly do read the same text differently. Distinguishing "fixable annotation problem" from "genuinely ambiguous task" from agreement statistics alone remains more art than science, and typically needs a manual review of the disagreement cases themselves.

What to learn next

What to learn next

These follow on from what you just read.

  • Computer Vision

    What is computer vision?

    Computer vision is teaching a computer to pull meaning out of pictures and video, when all it actually receives is a grid of brightness values.

  • Computer Vision

    How images are stored in a computer

    An image is a grid of numbers, one small set of numbers per dot, and every single thing in computer vision is arithmetic done on that grid.

  • Computer Vision

    OpenCV

    OpenCV is the free toolbox that reads, reshapes and measures images for you, so you never write resize or blur or edge detection by hand.