Vision Datasets and Annotation

Finding label errors in an image dataset

Rank every label by how much the model disbelieves it, check the worst few hundred by hand, and you will find nearly all your wrong labels in an afternoon.

On this page 9
  1. The short answer
  2. The analogy
  3. Why there are errors at all
  4. How the ranking works
  5. The one rule that makes or breaks it
  6. What you find
  7. Where you have seen this
  8. Remember this
  9. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

Let the model tell you which labels it disbelieves, then check those by hand.

The analogy

Think about a teacher marking two hundred answer sheets, knowing a few have been marked wrongly by an assistant.

Re-marking all two hundred takes a day. So instead the teacher looks at the ones where the mark and the paper disagree strangely. A confident, well-written answer marked zero. A blank page marked nine.

That short pile catches nearly every mistake, in twenty minutes.

Finding label errors works the same way. You do not check everything. You check where the model and the label disagree most.

Why there are errors at all

Every dataset has wrong labels. Yours does too.

People get tired. Two labellers apply different rules. A file gets renamed. Two exports get merged and the class numbers shift by one.

Even famous public datasets carry errors. Researchers checked ten widely used test sets. They found wrong labels in all ten, a few in every hundred.

That matters more than it sounds. A test set is your answer key. Errors in the answer key mean your accuracy number is partly measuring the errors.

How the ranking works

Train a model. For every image, ask how much it believes the label that image was given.

Sort by that belief, lowest first. The top of that list is where the wrong labels are.

   image     label given     model's belief in that label
   -------   -----------     ----------------------------
   img_41    "dog"           0.02   <- check this first
   img_98    "cat"           0.05   <- and this
   img_12    "dog"           0.11
   ...
   img_07    "cat"           0.99   <- no need to look

The one rule that makes or breaks it

The model must not have been trained on the image it is judging.

If it was, it has memorised that image and its wrong label. It will then agree with the mistake, confidently. The whole method collapses.

The standard fix is to split the data into five parts. Train on four, score the fifth. Rotate. Every image gets scored by a model that never saw it.

What you find

Three different things, and they need different responses.

Genuinely wrong labels. Fix them.

Genuinely ambiguous images. Neither label is right. These belong in your guidelines as examples, and often should be removed from the test set.

Hard but correct images. The model is wrong, not the label. Keep them. These are the most valuable images you have.

Only a human can tell these three apart. The ranking finds candidates; it does not make decisions.

Where you have seen this

  • Photo apps asking "is this the same person?" about a face they are unsure of.
  • Spam filters asking whether something was really spam.
  • Maps asking you to confirm a business is closed.

Remember this

  • Rank every label by how much the model disbelieves it.
  • The scores must come from a model that never saw that image.
  • The list is a list of candidates. A human decides.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install numpy==1.26.4 scikit-learn==1.7.2

Confident learning, built from scratch, measured honestly

sklearn.datasets.load_digits ships with scikit-learn, so this needs no download. Eight percent of the training labels are corrupted, and the goal is to find them without being told which.

label_errors.py
import numpy as np
from sklearn.datasets import load_digits
from sklearn.linear_model import LogisticRegression
from sklearn.preprocessing import StandardScaler
from sklearn.pipeline import make_pipeline
from sklearn.model_selection import cross_val_predict, train_test_split

X, Y = load_digits(return_X_y=True)               # 1797 8x8 handwritten digits, ships with sklearn
Xtr, Xte, ytr_true, yte = train_test_split(X, Y, test_size=0.3, random_state=0, stratify=Y)
rng = np.random.default_rng(0)

# Corrupt 8% of the TRAINING labels, the way a tired annotator or a bad merge does.
ytr = ytr_true.copy()
flip = rng.choice(len(ytr), int(0.08 * len(ytr)), replace=False)
ytr[flip] = (ytr[flip] + rng.integers(1, 10, len(flip))) % 10
print(f"{len(flip)} of {len(ytr)} training labels are wrong ({len(flip)/len(ytr):.1%}). "
      f"Nobody tells us which.")

# The rule that decides whether this works: the probabilities must be OUT OF SAMPLE.
# A model scoring images it was trained on is confident about its own memorised mistakes.
clf = make_pipeline(StandardScaler(), LogisticRegression(max_iter=1000))
proba = cross_val_predict(clf, Xtr, ytr, cv=5, method="predict_proba")

self_conf = proba[np.arange(len(ytr)), ytr]       # how sure the model is of the GIVEN label
order = np.argsort(self_conf)                     # least believable first

print("\nhand-check the k least believable labels:")
print("   k   really wrong   precision   share of all errors found   random baseline")
for k in (25, 50, 100, 200, 400):
    p = order[:k]
    hit = int((ytr[p] != ytr_true[p]).sum())
    r = rng.choice(len(ytr), k, replace=False)
    print(f"{k:4d} {hit:14d} {hit/k:12.1%} {hit/len(flip):27.1%} "
          f"{(ytr[r] != ytr_true[r]).mean():17.1%}")

# Confident learning: flag a label only when the model is confident about a DIFFERENT
# class, judged against that class's own average confidence, not a fixed cutoff.
pred = proba.argmax(1)
bar = np.array([proba[ytr == c, c].mean() for c in range(10)])
issue = (pred != ytr) & (proba[np.arange(len(ytr)), pred] > bar[pred])
tp = int((ytr[issue] != ytr_true[issue]).sum())
print(f"\nconfident-learning flags {issue.sum()} images: {tp} genuinely wrong "
      f"({tp/issue.sum():.1%} precision, {tp/len(flip):.1%} recall)")

fit = lambda x, y: make_pipeline(StandardScaler(), LogisticRegression(max_iter=1000)).fit(x, y)
print("\nsame held-out test set, three training sets")
print(f"  trained on the dirty labels      {fit(Xtr, ytr).score(Xte, yte):.1%}")
print(f"  flagged rows dropped             {fit(Xtr[~issue], ytr[~issue]).score(Xte, yte):.1%}")
print(f"  an oracle fixes every bad label  {fit(Xtr, ytr_true).score(Xte, yte):.1%}")
Output
100 of 1257 training labels are wrong (8.0%). Nobody tells us which.

hand-check the k least believable labels:
   k   really wrong   precision   share of all errors found   random baseline
  25             22        88.0%                       22.0%              4.0%
  50             47        94.0%                       47.0%              2.0%
 100             90        90.0%                       90.0%              8.0%
 200            100        50.0%                      100.0%              8.0%
 400            100        25.0%                      100.0%              7.8%

confident-learning flags 104 images: 81 genuinely wrong (77.9% precision, 81.0% recall)

same held-out test set, three training sets
  trained on the dirty labels      92.4%
  flagged rows dropped             95.7%
  an oracle fixes every bad label  97.2%

Reading the output

The ranking works extraordinarily well at the top. The first 50 images contain 47 genuine errors, a precision of 94%. Random checking of the same 50 images finds one. That is roughly a fiftyfold improvement in reviewer efficiency, from six lines of code.

All 100 errors are inside the top 200. Reviewing 200 of 1257 images, sixteen percent of the data, finds one hundred percent of the errors. That is the number to take to whoever is paying for the review.

Precision drops as you go down the list, exactly as it should. At k = 400 it is 25%, meaning three in four of those images are correctly labelled and only hard. Stop reviewing when precision falls below your reviewer's patience, and note where that happened.

The random baseline sits at the corruption rate, around 8 percent, because that is what random sampling gives you. It wanders between 2 and 8 percent across the rows purely from sampling noise on small draws.

Confident learning is more conservative and needs no threshold from you. It flags 104 images at 77.9% precision, catching 81% of the errors. It does so without being told how many to look at. That self-calibration is its main advantage over ranking by confidence. The per-class threshold adapts to classes the model finds hard.

The three-way accuracy comparison is the honest bottom line. 92.4% dirty, 95.7% after dropping flagged rows, 97.2% with a perfect oracle. Dropping the flagged rows recovered about two-thirds of the available gain. All three numbers come from the same held-out test set.

That last detail matters. A common way to overstate this method is to evaluate the cleaned model on the cleaned subset. That removes the hard examples from the test set as well. Always evaluate on a fixed, untouched test set.

The choice: drop, relabel, or reweight

Relabelling is best and costs the most. A human looks and fixes. This is what the oracle row measures, and it is the ceiling.

Dropping is cheap and loses information. As above, it recovered most of the benefit here. It biases the dataset toward easy examples, which matters more when your flagged set is large.

Never auto-correct to the model's prediction. That collapses the dataset toward whatever the model already believed and destroys exactly the hard examples you needed. The result looks better on every internal metric and is worse in deployment.

Doing this for detection and segmentation

Classification is the easy case. The same idea extends, with more bookkeeping.

Detection. Score each ground-truth box by the best-matching prediction's confidence. Separately, look for high-confidence predictions with no matching ground truth. The second set finds missing annotations, which are usually more common than wrong ones. The cleanlab library provides cleanlab.object_detection.filter.find_label_issues for this. It flags an image when any box appears mis-annotated, missing, wrongly classed, or badly placed.

Segmentation. Compare per-pixel predicted probability against the given mask and aggregate per region. Boundary pixels always disagree, so exclude a few pixels around each edge before aggregating or every mask is flagged.

Common mistakes

Using in-sample probabilities. The single fatal error. A model scoring its own training data agrees with the labels it memorised, including the wrong ones. Use cross_val_predict, or a held-out model, every time.

Trusting the flags without looking. The list contains genuinely hard correct images. Deleting them silently makes the dataset easier and your evaluation dishonest.

Only cleaning the training set. Errors in the test set corrupt the measurement itself, which is worse. Clean the test set first, and note that this changes every historical number you have.

Using a model too weak for the task. A model that cannot learn the task disbelieves everything, and the ranking is noise. Sanity-check that overall accuracy is reasonable before trusting the ordering.

Reporting the cleaned accuracy against the cleaned test set. See above.

Try it yourself

Change the corruption from uniform-random to systematic, always mapping class 3 to class 8. Rerun. Systematic errors are much harder to find, because the model learns the confusion and starts believing it. That experiment is why a single careless annotator is more dangerous than many random mistakes.

What to learn next

Researcher — Mathematics and papers.

Confident learning

Northcutt, Jiang and Chuang (2021), Confident Learning: Estimating Uncertainty in Dataset Labels (JAIR), formalise this. Let $\tilde{y}$ be the observed (noisy) label and $y^$ the latent true label. Assume class-conditional noise, so $p(\tilde{y} = i \mid y^ = j, \mathbf{x}) = p(\tilde{y} = i \mid y^* = j)$: the corruption depends on the true class, not on the image.

Define a per-class threshold as the average self-confidence of the examples carrying that label:

$$ t_j = \frac{1}{|\mathbf{X}{\tilde{y}=j}|} \sum{\mathbf{x} \in \mathbf{X}_{\tilde{y}=j}} \hat{p}(\tilde{y} = j; \mathbf{x}, \boldsymbol{\theta}) $$

Then build the confident joint $C_{\tilde{y}, y^*}$. It counts examples labelled $i$ whose predicted probability for class $j$ exceeds $t_j$:

$$ C_{\tilde{y}=i,\, y^*=j} = \left| \left{ \mathbf{x} \in \mathbf{X}_{\tilde{y}=i} : \hat{p}(\tilde{y}=j; \mathbf{x}, \boldsymbol{\theta}) \ge t_j,\; j = \arg\max_{k:\, \hat{p}(\tilde{y}=k) \ge t_k} \hat{p}(\tilde{y}=k; \mathbf{x}, \boldsymbol{\theta}) \right} \right| $$

Normalising gives an estimate of the joint distribution $\hat{Q}_{\tilde{y}, y^*}$. The off-diagonal mass estimates the number of errors in each direction.

Two properties make this useful in practice. The per-class thresholds handle class-conditional calibration error. A model underconfident on a hard class does not have that whole class flagged. And the paper proves exact estimation of the joint under perfect model calibration. It adds consistency results under bounded calibration error. The method degrades gracefully rather than failing.

The prevalence of the problem

Northcutt, Athalye and Mueller (2021) applied this to the test sets of ten widely used benchmarks. They validated the flagged items with Mechanical Turk. Their paper is Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks. They estimate an average of 3.3 percent errors. ImageNet validation sits at about 5.8 percent, and QuickDraw at about 10.1 percent.

The consequential finding is not the error rate but its effect on ranking. On ImageNet and QuickDraw, correcting the test labels reverses the ordering of model families. Higher-capacity models that scored better on the original labels score worse on the corrected ones. Their extra capacity fit the systematic noise. Their conclusion is uncomfortable and well supported. Practitioners may be better served by lower-capacity models, on real-world datasets with noisy labels.

The implication for benchmark reading is direct. Take a reported gap of one to two percentage points. On a test set with three percent label error, it measures nothing.

Alternative signals

Confident learning uses predicted probabilities. Several other signals identify mislabelled data, with different cost and different failure modes.

  • Loss-based filtering. Small-loss examples are likely clean, since networks fit clean data first. Co-teaching (Han et al., 2018) trains two networks that select small-loss batches for each other. That prevents the self-confirmation single-network filtering suffers.
  • Training dynamics. Swayamdipta et al. (2020), Dataset Cartography, plot per-example mean confidence against its variability across epochs. Low-confidence low-variability items are the mislabelled candidates; high-variability items are the genuinely ambiguous ones. This separates two categories that confident learning merges.
  • Influence functions. Koh and Liang (2017) estimate each training example's effect on a specific test loss. Principled and expensive, and unstable in deep non-convex models (Basu et al., 2021).
  • Area under the margin. Pleiss et al. (2020) rank by the accumulated margin between the assigned logit and the largest other logit. They calibrate the threshold using deliberately mislabelled inserted examples.
  • Nearest-neighbour disagreement in embedding space. Cheap, model-agnostic, and it also surfaces near-duplicates carrying different labels.

Dataset Cartography deserves particular attention, because it is the cheapest to add. The data is already being produced during training. And it distinguishes hard-but-correct from mislabelled, exactly the distinction a reviewer needs.

Learning with noisy labels, as an alternative to cleaning

Where relabelling is impossible, robust training methods exist:

  • Robust losses. Mean absolute error is noise-robust but optimises poorly; generalised cross entropy (Zhang and Sabuncu, 2018) interpolates between the two.
  • Loss correction. Estimate the noise transition matrix and adjust the loss (Patrini et al., 2017). The confident joint provides such an estimate directly.
  • Sample reweighting. Learn per-example weights from a small clean validation set (Ren et al., 2018).
  • Early stopping. Deep networks fit clean structure before memorising noise (Arpit et al., 2017). Stopping early is therefore a genuine, if crude, noise-robustness mechanism.

Cleaning and robust training are complementary. Cleaning improves the dataset permanently and for every future model; robust training only helps the current run.

Reporting standard

Any claim that a dataset was cleaned should state five things. The model and cross-validation scheme used to produce out-of-sample probabilities. The flagging rule and its parameters. How many flagged items were reviewed by a human, and by how many reviewers. The outcome breakdown into corrected, deleted and confirmed-correct. And the evaluation protocol: whether the test set was cleaned, and whether historical numbers were recomputed.

Without the last item, before-and-after accuracy comparisons are not interpretable.

References

What to learn next