Vision Datasets and Annotation

Model-assisted labelling

A model draws the labels first and a human corrects them, which saves most of the work and quietly hands the model's blind spots to your dataset.

On this page 10
  1. The short answer
  2. The analogy
  3. How it works
  4. What it saves
  5. The problem nobody mentions in the sales page
  6. Why that is worse than it sounds
  7. The fix, which is not complicated
  8. Where you have seen this
  9. Remember this
  10. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

A model draws the labels first, and a person fixes what is wrong.

The analogy

Think about a typist working from a rough draft instead of from dictation.

Correcting a draft is far faster than typing from nothing. The words are mostly there.

But there is a cost, and everyone who has proofread knows it. You stop reading properly. Your eye slides over sentences that look about right. Errors that were in the draft survive into the final copy, because you were checking rather than writing.

Model-assisted labelling has exactly this shape. Much faster, and a different kind of mistake survives.

How it works

Train a model on whatever you have already labelled. Run it on the unlabelled images. Put its guesses into the annotation tool as draft marks.

Now a person opens the image and sees boxes already drawn. They nudge, delete, add, and move on.

   unlabelled image
          |
     [ model ]  ->  draft boxes
          |
     [ person ]  ->  accept, adjust, delete, add
          |
     final labels

What it saves

A great deal. In the measured example below, a reviewer touched about one image in ten. The final accuracy was the same as touching every image.

That is roughly a ninefold reduction in effort. Real projects see savings in this range, which is why every serious annotation tool now offers it.

The problem nobody mentions in the sales page

The errors that remain are not spread evenly. They are concentrated exactly where the model is weakest.

Think about why. The reviewer catches obvious mistakes easily. Subtle mistakes, the ones the model made confidently, look plausible and get waved through.

So your finished dataset agrees with the model precisely where the model is wrong.

Why that is worse than it sounds

Your test set gets labelled the same way. So the test set also carries the model's blind spots.

Now you measure your new model on that test set. It looks good. It looks good partly because the answer key was written by its ancestor, and shares its weaknesses.

You have built a mirror and called it a measurement.

The fix, which is not complicated

Label a portion of your data from scratch, with no draft shown. Use people who never see the model's output.

Use that portion as your test set, and compare it against the assisted labels now and then. If the two start to diverge, you have caught the drift.

It costs a bit. It is the only thing that keeps the measurement honest.

Where you have seen this

  • Photo apps that group faces and ask you to confirm.
  • Document scanners that read a form and ask you to check the fields.
  • Speech tools that produce a transcript for a human to correct.

Remember this

  • Pre-labelling saves most of the work, and the savings are real.
  • The errors that survive are the model's own errors, not random ones.
  • Keep a portion labelled from scratch, and use it as the test set.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install numpy==1.26.4 scikit-learn==1.7.2

Simulating the review process, and measuring what it inherits

The two things worth measuring are the effort saved and the correlation between residual errors and the pre-labeller's errors. This simulates a reviewer who catches obvious mistakes reliably and confident mistakes rarely. That is what the automation-bias literature describes.

model_assisted.py
import numpy as np
from sklearn.datasets import load_digits
from sklearn.linear_model import LogisticRegression
from sklearn.preprocessing import StandardScaler
from sklearn.pipeline import make_pipeline
from sklearn.model_selection import train_test_split

X, y = load_digits(return_X_y=True)
Xseed, Xnew, yseed, ynew = train_test_split(X, y, test_size=0.6, random_state=0, stratify=y)
rng = np.random.default_rng(0)

# The pre-labelling model is trained on 120 images, which is what you actually have
# on day three of a project. It is decent, and it is systematically wrong somewhere.
few = rng.choice(len(Xseed), 120, replace=False)
pre = make_pipeline(StandardScaler(), LogisticRegression(max_iter=1000)).fit(Xseed[few], yseed[few])
suggest = pre.predict(Xnew)
conf = pre.predict_proba(Xnew).max(1)
print(f"pre-labeller accuracy on the new images: {(suggest == ynew).mean():.1%}")

P_CATCH_LOW, P_CATCH_HIGH = 0.90, 0.35   # a reviewer catches an obvious error, rarely a confident one
P_SCRATCH_ERR = 0.03                     # labelling from scratch is not perfect either

def review():
    """Reviewer accepts a suggestion unless something looks off. Confidence drives attention."""
    out, edits = suggest.copy(), 0
    for i in range(len(ynew)):
        wrong = suggest[i] != ynew[i]
        p_catch = P_CATCH_HIGH if conf[i] > 0.9 else P_CATCH_LOW
        if wrong and rng.random() < p_catch:
            out[i] = ynew[i]; edits += 1
        elif not wrong and rng.random() < 0.02:      # occasionally "corrects" a right answer
            out[i] = (out[i] + 1) % 10; edits += 1
    return out, edits

def from_scratch():
    out = ynew.copy()
    bad = rng.random(len(ynew)) < P_SCRATCH_ERR
    out[bad] = (out[bad] + rng.integers(1, 10, bad.sum())) % 10
    return out, len(ynew)

for name, fn in [("review the suggestions", review), ("label from scratch", from_scratch)]:
    lab, edits = fn()
    acc = (lab == ynew).mean()
    inherited = ((lab != ynew) & (lab == suggest)).sum()
    print(f"\n{name}")
    print(f"  images the annotator had to touch : {edits} of {len(ynew)} ({edits/len(ynew):.0%})")
    print(f"  final label accuracy              : {acc:.1%}")
    print(f"  errors that are the MODEL's error : {inherited} of {(lab != ynew).sum()}")

# Where the inherited errors land: not spread out, but piled on specific classes.
lab, _ = review()
print("\nreviewed labels, error count by true class")
for c in range(10):
    m = ynew == c
    print(f"  digit {c}: {(lab[m] != ynew[m]).sum():3d} wrong of {m.sum():3d} "
          f"(pre-labeller alone got {(suggest[m] != ynew[m]).sum():3d} wrong)")
Output
pre-labeller accuracy on the new images: 90.6%

review the suggestions
  images the annotator had to touch : 112 of 1079 (10%)
  final label accuracy              : 96.4%
  errors that are the MODEL's error : 14 of 39

label from scratch
  images the annotator had to touch : 1079 of 1079 (100%)
  final label accuracy              : 96.4%
  errors that are the MODEL's error : 0 of 39

reviewed labels, error count by true class
  digit 0:   3 wrong of 107 (pre-labeller alone got   1 wrong)
  digit 1:  11 wrong of 109 (pre-labeller alone got  25 wrong)
  digit 2:   2 wrong of 106 (pre-labeller alone got  12 wrong)
  digit 3:   5 wrong of 110 (pre-labeller alone got  16 wrong)
  digit 4:   3 wrong of 109 (pre-labeller alone got   6 wrong)
  digit 5:   1 wrong of 109 (pre-labeller alone got   6 wrong)
  digit 6:   4 wrong of 109 (pre-labeller alone got   4 wrong)
  digit 7:   3 wrong of 108 (pre-labeller alone got   5 wrong)
  digit 8:   4 wrong of 104 (pre-labeller alone got  19 wrong)
  digit 9:   2 wrong of 108 (pre-labeller alone got   7 wrong)

This is a simulation with an assumed reviewer model, driven by a random generator with a fixed seed. The reviewer catch rates are stated in the code, and they are the assumption. Measure your own team's rates rather than adopting these. The two accuracy figures landing on the same value is a coincidence of this seed.

Reading the output

The effort saving is real and large. 112 touched images against 1079, roughly a tenth of the work.

The accuracy is the same. Both routes landed at 96.4% here. That is the finding that makes model-assisted labelling worth doing, and it holds broadly. With a decent pre-labeller and an attentive reviewer, you get the same quality for a fraction of the effort.

14 of the 39 remaining errors are the model's own errors, retained verbatim. In the from-scratch route that number is zero, by construction. The errors have a different shape, even when they have the same count.

That difference is the entire point. Errors correlated with the model are not equivalent to random errors of the same rate. A model trained on them reproduces the same bias. A test set carrying them will not detect it.

The per-class table shows where the errors settle. Digit 1 has 11 remaining errors, the worst class. The pre-labeller alone got 25 wrong there, also its worst class. Digit 8, the pre-labeller's second-worst at 19, retained 4. The reviewer knocked the counts down everywhere but did not change the ranking.

Your finished dataset is weakest where your pre-labeller was weakest. If you now train on it and evaluate on it, that weakness is invisible in every number you produce.

The safeguards that work

A blind holdout. Label 5 to 10 percent from scratch, by annotators who never see a suggestion. Use it as the test set. This is not optional if the test set decides anything.

Randomise whether a suggestion is shown. For a small fraction of images, show a blank frame. Comparing the two populations gives you a direct estimate of the acceptance bias.

Record the provenance of every annotation. Created by a human, accepted from a model, or edited from a model. Most tools store this; most teams never query it. It lets you compute accuracy separately for accepted and edited annotations after the fact.

Track the acceptance rate over time. If it rises steadily, either the model is improving or the reviewers are becoming less attentive. Only the blind holdout distinguishes those.

Hide confidence scores from reviewers. Showing a high confidence score measurably increases acceptance. The reviewer's job is to look at the image, not at the model's opinion of itself.

What to pre-label with

Modern practice for vision uses a general segmentation model for the first round, not a task-specific one. It needs no training data at all.

Meta's SAM 3, released 20 November 2025, performs promptable concept segmentation. It detects, segments and tracks every instance of a concept, across images and video. The prompt can be a noun phrase such as "yellow school bus", an example image, or an interactive click. SAM 3.1 Object Multiplex followed on 27 March 2026 with a shared-memory approach for faster joint multi-object tracking. CVAT ships auto-annotation integrations including SAM-family and YOLO models.

The workflow follows from this. Use a general model with text prompts for round one. Train a task-specific model on the corrected results, and switch to that for round two. The general model gets you off zero; the specific model gets you accuracy.

Common mistakes

Pre-labelling with a model trained on the same images. It produces suspiciously good suggestions and no useful correction signal. Use out-of-sample predictions, as in the label-errors lesson.

Pre-labelling the test set. The most damaging version of everything above.

Showing low-quality suggestions. Below roughly 70 percent accuracy, correcting is slower than starting fresh, and reviewers begin deleting everything. Measure the pre-labeller before deploying it into the workflow.

No provenance field. Without it you cannot audit the accepted labels afterwards, and the problem becomes undetectable rather than only present.

Paying reviewers by images per hour. That is a direct incentive to accept everything. Pay for time, and audit quality separately.

Try it yourself

Set P_CATCH_HIGH = 0.90, making the reviewer equally attentive whether or not the model is confident. Watch the inherited-error count fall. The gap between the two runs is the exact cost of automation bias in this simulation. It tells you what reviewer training is worth.

What to learn next

Researcher — Mathematics and papers.

Automation bias is a measured phenomenon

The effect has a substantial literature outside machine learning. Skitka, Mosier and Burdick (1999), Does automation bias decision-making?, distinguish two failure types. Omission errors: the human misses an event because the automation did not flag it. Commission errors: the human acts on an incorrect automated recommendation despite contrary evidence being available.

Both appear in annotation. Omission is a missing box the reviewer did not notice was missing. It is the more common and less visible failure, because an absent annotation draws no attention. Commission is accepting a wrong class the model asserted confidently.

The medical imaging literature on computer-aided detection is the closest analogue, and the most carefully measured. Studies report sensitivity gains from the aid, and reduced detection of findings the aid missed. The general finding is consistent. Aided performance exceeds unaided performance on average. The error distribution shifts to align with the aid's own weaknesses.

Why the correlation matters more than the rate

Model an annotation process producing labels $\tilde{y}$ from true labels $y^*$. Two processes with identical error rate $\eta$ are not equivalent.

Independent noise. $p(\tilde{y} \ne y^*) = \eta$ with errors independent of any model. Under class-conditional independent noise, many training procedures are provably robust. A test set with independent noise gives an unbiased, if attenuated, ranking of models.

Model-correlated noise. $\tilde{y}$ agrees with a pre-labeller $f_0$ wherever $f_0$ is confidently wrong. Now the errors sit in a structured region of input space, and three things follow.

First, a model $f_1$ trained on $\tilde{y}$ inherits $f_0$'s decision boundary in exactly the wrong region. Gradient descent has no signal to correct it.

Second, a test set labelled the same way scores $f_1$ highly in that region. The error is undetectable by the evaluation.

Third, the effect compounds across rounds. Each generation's pre-labeller is trained on the previous generation's assisted labels. The shared blind spot is reinforced rather than diluted. This is the dataset analogue of model collapse. The blind holdout is the only measurement that breaks the loop.

Foundation models as pre-labellers

Kirillov et al. (2023), Segment Anything, established the pattern. It is a promptable segmentation model trained on SA-1B: 1.1 billion masks over 11 million images. A three-stage data engine built that corpus, and it is the canonical description of model-assisted labelling at scale. Stage one was assisted-manual, stage two semi-automatic with the model proposing confident objects, stage three fully automatic.

SAM 3 (November 2025) extends this to promptable concept segmentation. It detects and segments every instance of a concept, across images and video. The prompt is a noun phrase, an image exemplar, or both. Its data engine automatically annotated over four million unique concepts. SAM 3.1 Object Multiplex (March 2026) adds shared-memory joint multi-object tracking.

Two consequences for dataset construction.

The cold start problem is largely solved for segmentation. A general model with text prompts produces usable first-round annotations with no task-specific training data.

The bias source moved upstream. Your dataset now inherits the foundation model's blind spots rather than your own model's. Those blind spots are shared across every team using the same model. That makes them systematic across the field, and correspondingly harder to notice. A failure mode common to every dataset in a benchmark suite is invisible to every comparison within it.

Study design for validating an assisted pipeline

To claim assisted labelling did not degrade your dataset, you need a comparison the pipeline cannot influence.

  • Randomised assignment. For a random subset, assign images to assisted or unassisted labelling, with annotators blind to the condition where possible.
  • Report accuracy conditioned on model correctness. Specifically the acceptance rate on suggestions that were wrong, split by the model's confidence. That is the automation-bias estimate.
  • Report time per image in both arms, so the saving is quantified alongside the cost.
  • Retain provenance per annotation and recompute the analysis after the project, not only during the pilot. Reviewer attention drifts.
  • Report the residual error correlation: the fraction of remaining errors that match the pre-labeller's prediction. Independent noise gives roughly the chance rate; anything higher quantifies the inheritance.

That last statistic is the one this lesson's simulation makes concrete. It is almost never reported in dataset papers that used model assistance.

References

What to learn next