Labelling data
Labels are opinions written down, and a rater who is consistently wrong in one direction damages a model far more than a rater who is randomly sloppy.
- 12 min read
- 3 reading levels
- Updated
Read these first
On this page 8
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Labelling means writing the correct answer next to each example, so a model has something to learn from.
Think about two teachers marking the same set of exam papers. Give both the same script and they hand back different marks. One is generous, the other is strict.
Neither is lying. They read the same words and disagreed about what those words were worth.
Labels are the same. They are opinions, written down and then treated as facts.
Why this matters more than it sounds
A model has no way to check a label. Whatever you wrote next to an example is, from the model's point of view, the truth about the world.
If your raters wrote something wrong, the model learns the wrong thing carefully and confidently. It will then repeat that error on every future case that looks similar.
That is why label quality is not a tidy-up task. It is the thing being learned.
The two ways raters go wrong, and they are not equal
Careless. A rater is tired and ticks the wrong box now and then, in no particular pattern.
Consistent. A rater has quietly understood the task differently. Perhaps they mark every borderline complaint as "not urgent", every time.
The second one is far more damaging, even when the two make the same number of mistakes. Careless errors point in all directions and partly cancel out. A consistent error points one way, and drags the model with it.
How the damage travels
guideline says rater reads it as model learns
"urgent if the -> "urgent only if -> ignores every
customer is they use the polite complaint
unhappy" word 'angry'"Nobody in that chain made a mistake they could see. The guideline was vague, so a reasonable person filled the gap, and the model absorbed their interpretation as law.
Somewhere you have seen this
Content moderation on any large platform. Two reviewers in two countries look at the same post and reach different verdicts, because the rulebook did not cover it.
Medical scans, too. Two radiologists disagree on borderline images at a measurable rate. That rate limits how good an automatic reader can be.
How teams keep labels honest
- Write the guideline with examples, including hard ones. A one-line definition always fails.
- Have two people label the same batch. Then measure how often they agree.
- Argue about the disagreements. That conversation is where the guideline gets fixed.
- Keep a small gold set — cases with an answer everyone accepts — and re-check raters against it.
The honest part
You will not reach perfect agreement, ever. On genuinely hard tasks, two experts disagree on a fifth of cases and both are being reasonable.
That disagreement rate is useful information, not a failure. It tells you roughly how good a model can get. Chasing accuracy above your own raters' agreement is chasing noise.
Remember this
- Labels are opinions recorded as facts. The model cannot tell the difference.
- A consistently biased rater hurts far more than a randomly careless one.
- Measure agreement between raters. That number is your realistic ceiling.
What to learn next
- Cleaning data — the mechanical errors that sit beneath the judgement calls.
- Model evaluation — why accuracy alone hid the strict rater's damage.
- Fairness metrics — measuring whether a rater's bias reached particular groups.
Developer — Code and libraries.
The claim above is testable. Here we give two raters the same number of mistakes and compare the damage.
Rater A is careless: two rows in ten flipped at random, in both directions. Rater B is strict: the same count of mistakes, all pushing borderline positives down to negative. Same budget of errors, opposite structure.
Setup
pip install numpy pandas scikit-learnSame error count, different structure
import numpy as np
import pandas as pd
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import accuracy_score, cohen_kappa_score, recall_score
from sklearn.model_selection import train_test_split
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler
rng = np.random.default_rng(11)
n = 1200
x1, x2 = rng.normal(0, 1, n), rng.normal(0, 1, n)
margin = 1.6 * x1 + 1.2 * x2 # how far a case sits from the real boundary
truth = (margin + rng.normal(0, 0.4, n) > 0).astype(int)
df = pd.DataFrame({"x1": x1, "x2": x2, "margin": margin, "label": truth})
train, test = train_test_split(df, test_size=0.3, random_state=0, stratify=df["label"])
r = np.random.default_rng(5)
# Rater A is careless: two rows in ten are flipped at random, in both directions.
careless = np.where(r.random(len(train)) < 0.20, 1 - train["label"], train["label"])
# Rater B is strict: same number of mistakes, all in one direction, all near the boundary.
strict = train["label"].to_numpy().copy()
budget, changed = int(0.20 * len(train)), 0
for i in train["margin"].abs().to_numpy().argsort(): # closest calls first
if changed >= budget:
break
if strict[i] == 1:
strict[i] = 0
changed += 1
def evaluate(labels):
model = make_pipeline(StandardScaler(), LogisticRegression(max_iter=1000))
model.fit(train[["x1", "x2"]], labels)
pred = model.predict(test[["x1", "x2"]])
return round(accuracy_score(test["label"], pred), 3), round(recall_score(test["label"], pred), 3)
print("wrong labels, careless rater:", int((careless != train["label"]).sum()))
print("wrong labels, strict rater :", int((strict != train["label"]).sum()))
print("agreement between them :", round((careless == strict).mean(), 3))
print("Cohen's kappa between them :", round(cohen_kappa_score(careless, strict), 3))
print()
print(f"{'labels used for training':28s}{'accuracy':>10}{'recall':>9}")
for name, labels in [("perfect labels", train["label"]),
("careless rater", careless),
("strict rater", strict)]:
a, rc = evaluate(labels)
print(f"{name:28s}{a:>10}{rc:>9}")wrong labels, careless rater: 181 wrong labels, strict rater : 168 agreement between them : 0.687 Cohen's kappa between them : 0.363 labels used for training accuracy recall perfect labels 0.939 0.94 careless rater 0.939 0.918 strict rater 0.814 0.636
This result surprises most people
168 errors cost more than 181 errors. The strict rater made fewer mistakes and did far more damage: accuracy 0.814 against 0.939, recall 0.636 against 0.918.
Random noise did almost nothing. 181 flipped labels, out of 840, and accuracy did not move. A linear model averaging over symmetric noise recovers close to the same boundary. This is a real theoretical result, not a fluke of this seed.
Recall is where the strict rater shows up. Overall accuracy fell 12 points; recall fell 30. The model learned to say "no" in exactly the cases the rater said "no" to. If you were only watching accuracy, you would have called this a mild problem.
Read this twice if it feels wrong. Almost everybody's instinct is that 20% wrong labels must be roughly 20% as damaging in both cases. The structure of the error matters more than the rate.
What the agreement numbers mean
agreement between them: 0.687 — the two raters gave the same answer on 69% of rows. That sounds tolerable.
Cohen's kappa: 0.363 — kappa corrects for agreement that would happen by chance alone. Two raters guessing at these class proportions would agree often. Kappa asks how much better than that they did.
A rough convention (Landis and Koch, 1977): below 0.20 is slight, 0.21 to 0.40 fair, 0.41 to 0.60 moderate, 0.61 to 0.80 substantial. 0.363 means these two raters have not been given a usable guideline.
Report kappa, never raw agreement. On an imbalanced task, raw agreement of 0.95 can correspond to a kappa near zero.
Common mistakes
One rater per item, forever. With a single rater you cannot measure anything. Double-label at least 5% of every batch, permanently, not once at the start.
Measuring agreement once and stopping. Rater interpretation drifts over weeks, especially as they see unusual cases. Track kappa per batch and watch the trend.
Fixing disagreements by majority vote and moving on. The disagreement is a signal that your guideline has a hole. Vote if you must, but read a sample of the conflicts and update the document.
Letting the guideline live in someone's head. If it is not written with worked examples of hard cases, every new rater invents their own version.
Assuming model errors are model errors. Before tuning anything, hand-check 50 items the model got "wrong". A meaningful share are usually cases where the model was right and the label was not.
Try it yourself
Set the strict rater's budget to 0.10 instead of 0.20 and the careless rater to 0.40. That gives the careless rater four times as many mistakes.
Predict which model does better before you run it. Then flip the strict rater's direction — turn borderline zeros into ones — and watch precision and recall swap places.
What to learn next
- Cleaning data — the mechanical errors that sit beneath the judgement calls.
- Model evaluation — why accuracy alone hid the strict rater's damage.
- Fairness metrics — measuring whether a rater's bias reached particular groups.
Researcher — Mathematics and papers.
Noise models, formally
Let $Y$ be the true label and $\tilde{Y}$ the observed one. The noise process is characterised by:
$$ \rho_{ij}(x) \;=\; p!\left(\tilde{Y} = j \mid Y = i, X = x\right) $$
Three cases, in increasing difficulty:
Symmetric (uniform) noise. $\rho_{ij}$ is constant in $x$ and equal for all $i \neq j$. Natarajan et al. (2013), Learning with Noisy Labels, NeurIPS, give an unbiased loss correction: for binary labels with rates $\rho_{+1}, \rho_{-1}$ and $\rho_{+1} + \rho_{-1} < 1$,
$$ \tilde{\ell}(t, y) \;=\; \frac{(1 - \rho_{-y})\,\ell(t, y) \;-\; \rho_{y}\,\ell(t, -y)}{1 - \rho_{+1} - \rho_{-1}} $$
Where $\ell$ is the clean loss, $t$ the prediction and $\rho_y$ the flip rate out of class $y$. Minimising $\tilde{\ell}$ on noisy data is equivalent in expectation to minimising $\ell$ on clean data. The Bayes-optimal classifier is unchanged by symmetric noise; only the estimation variance grows.
Asymmetric (class-conditional) noise. $\rho_{ij}$ still constant in $x$ but asymmetric across classes. The optimal decision threshold shifts. Correctable if the noise rates can be estimated, which is the premise of confident learning.
Instance-dependent noise. $\rho_{ij}(x)$ genuinely depends on $x$. This is the strict-rater case: flip probability is high near the decision boundary and near zero away from it. It is not identifiable in general without extra assumptions, and it shifts the Bayes-optimal boundary itself. No reweighting fixes it, because the corrupted conditional is a different function.
The developer experiment is a demonstration of exactly this hierarchy, and of why quoting a single "label noise rate" is close to meaningless.
Confident learning
Northcutt, Jiang and Chuang (2021), Confident Learning: Estimating Uncertainty in Dataset Labels, JAIR, estimate the joint distribution $p(\tilde{Y}, Y^{*})$ using out-of-sample predicted probabilities and per-class thresholds set to the average self-confidence:
$$ t_j \;=\; \frac{1}{|\mathbf{X}{\tilde{y}=j}|} \sum{x \in \mathbf{X}_{\tilde{y}=j}} \hat{p}(\tilde{y}=j \mid x) $$
Off-diagonal mass in the estimated confident joint identifies candidate mislabelled examples. The method is model-agnostic and requires only calibrated out-of-fold probabilities. It is implemented in cleanlab.
The companion paper, Northcutt, Athalye and Mueller (2021), Pervasive Label Errors in Test Sets, applied it to ten benchmark test sets and found an average 3.4% error rate, with model rankings reversing after correction on several.
Agreement statistics, and their traps
Cohen's kappa for two raters:
$$ \kappa \;=\; \frac{p_o - p_e}{1 - p_e} $$
Where $p_o$ is observed agreement and $p_e$ is agreement expected by chance under independent marginals. Two known pathologies (Feinstein and Cicchetti, 1990):
- The prevalence paradox. With a highly skewed class distribution, $p_e \to p_o$ and $\kappa$ collapses towards zero despite high raw agreement.
- The bias paradox. Two raters with opposite marginal biases can produce a higher $\kappa$ than two raters with matched marginals, because the imbalance lowers $p_e$.
Report $\kappa$ together with $p_o$ and the marginal distributions of each rater. Kappa alone is not interpretable.
Krippendorff's alpha generalises to any number of raters, missing data and ordinal, interval or nominal scales. It is the right default for real annotation pipelines, where not every rater sees every item. Artstein and Poesio (2008), Inter-Coder Agreement for Computational Linguistics, Computational Linguistics, is the definitive practical treatment.
Fleiss' kappa handles multiple raters but assumes a fixed number of ratings per item and different raters per item. It answers a different question from Cohen's and the two are not comparable.
Aggregating multiple raters
Majority vote discards information about rater reliability. Dawid and Skene (1979), Maximum Likelihood Estimation of Observer Error-Rates Using the EM Algorithm, Applied Statistics, model each rater with a confusion matrix and infer true labels and rater reliabilities jointly by expectation-maximisation. It remains a strong baseline four decades on, and typically beats majority vote when rater quality is uneven.
Extensions: Whitehill et al. (2009) add per-item difficulty (GLAD); Welinder et al. (2010) add rater bias and competence in a multidimensional model.
Programmatic labelling
Ratner et al. (2017), Snorkel: Rapid Training Data Creation with Weak Supervision, VLDB, replace hand labels with labelling functions: noisy, correlated heuristics whose accuracies are estimated without ground truth by fitting a generative model to their agreement structure. Output is a probabilistic label used to train a downstream discriminative model.
The honest limitation: the generative model recovers accuracies only when the labelling functions are not all wrong in the same way. Correlated systematic errors across functions are unidentifiable, which is the programmatic version of the strict-rater failure.
Practical protocol
- Write the guideline with 20 or more worked examples, weighted towards boundary cases.
- Double-label 5 to 10% of every batch, indefinitely, and track $\kappa$ per batch.
- Maintain a gold set with adjudicated answers; measure each rater against it, and estimate a per-rater confusion matrix.
- Adjudicate disagreements in a recorded session; every adjudication should either resolve an item or amend the guideline.
- Run confident learning over the finished set before training, and re-inspect the top-ranked suspects by hand.
- Estimate the model ceiling as the adjudicated agreement rate, and stop tuning once you are near it.
Reading
- Natarajan et al., Learning with Noisy Labels, NeurIPS 2013.
- Northcutt, Jiang and Chuang, Confident Learning, JAIR 2021 — arxiv.org/abs/1911.00068
- Dawid and Skene, Maximum Likelihood Estimation of Observer Error-Rates Using the EM Algorithm, Applied Statistics 1979.
- Artstein and Poesio, Inter-Coder Agreement for Computational Linguistics, Computational Linguistics 2008.
- Ratner et al., Snorkel, VLDB 2017 — arxiv.org/abs/1711.10160
What to learn next
- Cleaning data — the mechanical errors that sit beneath the judgement calls.
- Model evaluation — why accuracy alone hid the strict rater's damage.
- Fairness metrics — measuring whether a rater's bias reached particular groups.