Vision Datasets and Annotation
Writing annotation guidelines
The guideline document decides whether two labellers looking at the same image produce the same answer, and that agreement is the ceiling on your accuracy.
- 13 min read
- 3 reading levels
- Updated
Read these first
On this page 9
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
The guideline is the written rulebook that makes two different people label the same picture the same way.
The analogy
Think about two people cutting vegetables for the same dish, in different kitchens, from the same recipe.
If the recipe says "chop the onions", one comes back with fine mince and the other with thick rings. Both followed the recipe. The dish is ruined anyway.
If the recipe says "onions, half-centimetre dice", both hands produce the same thing.
Annotation guidelines are that second recipe. Vague ones do not produce vague labels. They produce confidently inconsistent labels, which is much worse, because nothing looks wrong.
Why this decides your accuracy
Here is the part people learn too late.
Suppose two careful labellers disagree on a fifth of your images. No model can then be more than about four-fifths right against your test set. The disagreement is baked into the answer key.
Your test set is not the truth. It is one group of people's opinion, written down. The guideline is what makes that opinion consistent.
What a good guideline contains
A rule for every edge case you have actually met. Not imagined ones. Real ones, from real images, with a picture beside each rule.
Examples of correct and wrong labels, side by side. One picture beats three paragraphs.
A rule for how tight the box should be. Does it include the shadow? The part hidden behind a pole? A tiny gap of background?
A rule for objects that are hardly there. A person half out of frame. A tiny face in a crowd. Say where the line is.
A rule for uncertainty. Give labellers a way to say "I do not know", and make sure someone reads those cases.
BAD rule: "label all vehicles"
GOOD rule: "label cars, buses, trucks, autos and two-wheelers.
Box the visible pixels only, no shadow.
Skip anything under twenty pixels tall.
Skip parked vehicles behind a barrier.
Mark 'unclear' rather than guessing."How to know if it is working
Give the same fifty images to two labellers, separately. Compare their answers.
If they agree, your rulebook works. If they disagree, you have found the exact places where the rulebook is silent. Go fix those places.
Do this before labelling ten thousand images. Not after.
The trap in measuring agreement
Suppose ninety-five in every hundred images have no defect. Two labellers who both say "no defect" every single time will agree ninety-five times in a hundred.
That looks excellent. It means nothing. Neither of them did any work.
So the plain agreement percentage lies whenever one answer is much more common than the others. There is a proper measure that corrects for this, and the next section shows the difference in numbers.
Where you have seen this
- Two doctors reading the same scan and reaching different conclusions.
- Two exam graders giving different marks to one script.
- Two people describing the same accident to a police officer.
Remember this
- Annotator agreement is the ceiling on your measurable accuracy.
- Good guidelines are built from real edge cases, with pictures.
- Plain agreement percentage lies when one answer dominates.
What to learn next
- The COCO dataset format — how these decisions get written to a file.
- Finding label errors in an image dataset — catching the disagreements you did not prevent.
- Model evaluation — the metrics that sit on top of this answer key.
Developer — Code and libraries.
Setup
pip install numpy==1.26.4 scikit-learn==1.7.2Measuring agreement, and watching the naive metric fail
import numpy as np
from sklearn.metrics import cohen_kappa_score, confusion_matrix
# 100 images, mostly "no defect". Two annotators, no written guideline.
rng = np.random.default_rng(0)
truth = rng.choice([0, 1], 100, p=[0.92, 0.08]) # 0 = fine, 1 = defect
a = truth.copy(); b = truth.copy()
a[rng.choice(100, 4, replace=False)] ^= 1 # each makes a few mistakes
b[rng.choice(100, 5, replace=False)] ^= 1
print("raw agreement ", f"{(a == b).mean():.1%}")
print("Cohen's kappa ", f"{cohen_kappa_score(a, b):.3f}")
print("confusion between them:\n", confusion_matrix(a, b))
print("\nA lazy annotator who always says 'fine':")
lazy = np.zeros(100, int)
print("raw agreement with A ", f"{(a == lazy).mean():.1%}")
print("Cohen's kappa with A ", f"{cohen_kappa_score(a, lazy):.3f}")
print("-> raw agreement barely moved. Kappa exposed it. Never report raw agreement alone.")
# Box agreement: the same object, boxed by two people with different habits.
def iou(p, q):
ax, ay, aw, ah = p; bx, by, bw, bh = q
ix = max(0, min(ax + aw, bx + bw) - max(ax, bx))
iy = max(0, min(ay + ah, by + bh) - max(ay, by))
inter = ix * iy
return inter / (aw * ah + bw * bh - inter)
tight = (100, 100, 60, 120) # "box the visible pixels only"
print("\nguideline difference IoU with the tight box")
for name, box in [("2 px of slack", (98, 98, 64, 124)),
("5 px of slack", (95, 95, 70, 130)),
("includes the shadow", (100, 100, 60, 150)),
("includes an occluded leg", (100, 100, 90, 120)),
("amodal: the whole object", (85, 100, 90, 140))]:
print(f" {name:34s} {iou(tight, box):.3f}")
print("\nCOCO calls a detection correct at IoU 0.5 and reports the average from 0.5 to 0.95.")
print("Two annotators who disagree at 0.70 have already capped your headline number.")raw agreement 91.0% Cohen's kappa 0.689 confusion between them: [[78 3] [ 6 13]] A lazy annotator who always says 'fine': raw agreement with A 81.0% Cohen's kappa with A 0.000 -> raw agreement barely moved. Kappa exposed it. Never report raw agreement alone. guideline difference IoU with the tight box 2 px of slack 0.907 5 px of slack 0.791 includes the shadow 0.800 includes an occluded leg 0.667 amodal: the whole object 0.571 COCO calls a detection correct at IoU 0.5 and reports the average from 0.5 to 0.95. Two annotators who disagree at 0.70 have already capped your headline number.
Reading the output
Two competent annotators agree 91 percent of the time, and kappa says 0.689. That gap is the whole point of kappa. Most of the 91 percent came free, because both said "fine" on the many easy images. Kappa subtracts the agreement you would expect by chance and reports what is left.
The lazy annotator scores 81 percent raw agreement and kappa exactly 0.000. Someone who never looks at an image and always answers "fine" beats four out of five images. If your quality dashboard shows raw agreement, that person passes.
Kappa gives them the zero they earned. This is the single strongest argument for reporting kappa, and it takes one line of scikit-learn.
The IoU table converts guideline vagueness into a number. Two annotators who differ by five pixels of slack are at 0.791. That is already below COCO's IoU 0.8 threshold. At the stricter end of the mAP range, two humans following one instruction would score each other wrong.
Shadows and occlusion are worse. Including the shadow costs 0.800. Extending the box over an occluded leg costs 0.667. Amodal boxing, drawing the whole object including hidden parts, costs 0.571. It would be scored as a miss at any threshold above 0.5.
These are not annotator errors. They are unwritten-rule differences. A single sentence in the guideline removes each of them.
Interpreting kappa
The Landis and Koch (1977) bands are the convention, and they are conventions rather than laws:
| Kappa | Reading |
|---|---|
| below 0.20 | Slight; the task as written is not answerable |
| 0.21 to 0.40 | Fair; rewrite the guideline |
| 0.41 to 0.60 | Moderate; usable for a coarse task |
| 0.61 to 0.80 | Substantial; a normal target for a subjective task |
| above 0.81 | Almost perfect; expected for objective tasks |
Kappa is also known to behave badly with extreme class imbalance. It can be low even when both annotators are almost always correct. Report the confusion matrix alongside it, so the reader can see what is happening.
For more than two annotators use Fleiss's kappa. For ordinal or continuous labels, or when annotators cover different subsets, use Krippendorff's alpha, which handles missing data natively.
The guideline template that works
Structure the document so that each rule can be checked against an image.
- Purpose. One paragraph on what the model will do. Annotators make better judgement calls when they know the downstream use.
- Class definitions. One paragraph and at least two example images per class, including one hard example.
- Geometry rules. Tight versus loose, shadows, reflections, occlusion, amodal versus modal, minimum size.
- Counting rules. One box per object or one per group. What to do with a crowd.
- Decision tree for hard cases. An explicit ordered set of questions, not prose.
- The uncertainty option. When to use it, and a promise that someone will review those items.
- A changelog. Dated, with a note on whether earlier labels were re-done.
That last item is the one everybody skips and everybody needs. When the guideline changes at week three, the labels from weeks one and two follow a different rulebook. Either re-label them or record the boundary, because a model trained across an unmarked rule change learns the contradiction.
Common mistakes
Writing the guideline before labelling anything. You cannot imagine the edge cases. Label 100 images yourself first, list every hesitation, and write the guideline from that list.
Rules without pictures. "Box tightly" means five different things to five people.
Punishing the "unclear" option. If annotators are measured on throughput, they will guess instead of flagging. Genuinely ambiguous images are the most informative ones you have.
Never re-measuring. Agreement drifts as annotators develop personal habits. Re-run the redundant-labelling check monthly.
Treating the guideline as the model's specification. It is a specification for humans. If the model must handle a case the guideline forbids annotators from labelling, the model will never learn it.
Try it yourself
Replace cohen_kappa_score with a Krippendorff's alpha implementation and confirm the two agree closely on this binary complete-coverage case. Then delete 30 percent of annotator B's answers at random. Watch kappa refuse to compute while alpha handles it. That difference is why alpha is preferred in practice.
What to learn next
- The COCO dataset format — how these decisions get written to a file.
- Finding label errors in an image dataset — catching the disagreements you did not prevent.
- Model evaluation — the metrics that sit on top of this answer key.
Researcher — Mathematics and papers.
Agreement statistics
Cohen's kappa (1960) corrects observed agreement for chance:
$$ \kappa = \frac{p_o - p_e}{1 - p_e} $$
$p_o$ is observed agreement, $p_e$ the agreement expected if both annotators labelled independently at their own observed marginal rates. $\kappa = 1$ is perfect, $\kappa = 0$ is chance-level, and negative values indicate systematic disagreement.
Two documented pathologies matter in practice. The prevalence problem comes first. With a highly skewed marginal distribution, $p_e$ approaches $p_o$ and $\kappa$ collapses toward zero. That happens even when both annotators are correct nearly always. The bias problem: two annotators with different marginal rates can produce an inflated $\kappa$. Byrt et al. (1993) propose reporting prevalence-adjusted bias-adjusted kappa (PABAK) alongside, and at minimum the full confusion matrix should always be shown.
Fleiss's kappa (1971) generalises to a fixed number of raters per item without requiring the same raters throughout.
Krippendorff's alpha is the most general:
$$ \alpha = 1 - \frac{D_o}{D_e} $$
$D_o$ is observed disagreement and $D_e$ disagreement expected by chance, both computed from a user-supplied difference function $\delta$. Choosing $\delta$ makes it work for nominal, ordinal, interval and ratio data. It also handles missing values and any number of annotators. Krippendorff (2004) suggests $\alpha \ge 0.800$ for conclusions and $\alpha \ge 0.667$ for tentative ones.
For spatial tasks, agreement must be defined geometrically. Take mean IoU over matched instances, plus a matching step that itself needs a threshold. Then account separately for instances one annotator found and the other missed. Reporting only mean IoU over matched pairs hides disagreement about existence, which is usually the larger effect.
Label noise as a ceiling
Take a binary task where each annotator flips the true label independently with probability $\eta$. The Bayes-optimal classifier, evaluated against a single-annotator answer key, has accuracy at most $1 - \eta$. Two consequences follow.
Model comparison becomes unreliable near the ceiling. Northcutt et al. (2021) estimate label error rates averaging 3.3 percent, across ten commonly used test sets. On ImageNet and QuickDraw the model ranking reverses on corrected labels. Higher-capacity models that scored better on the original labels score worse on the corrected ones. They were fitting the errors.
Reported headline numbers should carry a noise estimate. A benchmark reporting 96 percent accuracy, on a test set with 4 percent label error, has saturated. Further improvement on it is not measurable.
Consensus and aggregation
Where redundant annotation is affordable, majority vote is the weakest usable aggregator. It weights a careless annotator equally with a careful one.
Dawid and Skene (1979) model each annotator with a confusion matrix. They estimate the true labels and the annotator matrices together, by expectation-maximisation. This remains the standard baseline and outperforms majority vote when annotator quality varies. Welinder et al. (2010) extend it with per-item difficulty, and per-annotator competence and bias. Ambiguous items can then be identified rather than averaged away.
The item-difficulty parameter is the useful output, more than the aggregated label. Items the model says are hard for humans are the ones to re-examine. Add them to the guideline, and exclude them from a clean evaluation split.
Study design
A defensible annotation study reports:
- Number of annotators, their selection and their training.
- The redundant fraction and how items for it were chosen. Random selection is required; using only easy items inflates agreement.
- Agreement statistic with its variant named, plus the confusion matrix or the full distance matrix.
- Guideline version, with a changelog and a statement of whether earlier labels were re-done after changes.
- Aggregation method, and what was done with unresolved disagreement.
Gebru et al. (2021), Datasheets for Datasets, provides the accepted structure for reporting this. The "Collection Process" and "Preprocessing/cleaning/labeling" sections exist to hold exactly these facts.
References
- Cohen, A Coefficient of Agreement for Nominal Scales, 1960
- Landis and Koch, The Measurement of Observer Agreement for Categorical Data, Biometrics, 1977
- Dawid and Skene, Maximum Likelihood Estimation of Observer Error-Rates, 1979
- Byrt et al., Bias, Prevalence and Kappa, J. Clinical Epidemiology, 1993
- Krippendorff, Content Analysis: An Introduction to Its Methodology, 2004
- Welinder et al., The Multidimensional Wisdom of Crowds, NIPS 2010
- Northcutt et al., Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks, 2021 — arxiv.org/abs/2103.14749
- Gebru et al., Datasheets for Datasets, 2021 — arxiv.org/abs/1803.09010
What to learn next
- The COCO dataset format — how these decisions get written to a file.
- Finding label errors in an image dataset — catching the disagreements you did not prevent.
- Model evaluation — the metrics that sit on top of this answer key.