Vision Datasets and Annotation

Active learning for labelling

Instead of labelling images in the order they arrived, let a model pick the ones it is least sure about, and reach the same accuracy with far fewer labels.

On this page 10
  1. The short answer
  2. The analogy
  3. Why the order matters so much
  4. How the model picks
  5. The measure that works best
  6. The trap at the beginning
  7. The other trap
  8. Where you have seen this
  9. Remember this
  10. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

Let the model choose which images to label next, picking the ones it finds hardest.

The analogy

Think about a student revising for an exam with a thousand practice questions and one evening left.

Working through them in order is the worst plan. Most are easy, and time spent on those teaches nothing.

A better plan. Try each one quickly. Spend the evening on the ones you got wrong or guessed at. Same evening, far more learning.

Active learning is that plan, with the model as the student.

Why the order matters so much

Labelling costs money and time. Ten thousand images might take weeks.

Most of those images are easy. A well-lit picture of a plainly visible object teaches the model almost nothing new. That stays true after the first hundred like it.

The hard ones are where the learning is. Odd angles, poor light, rare classes, things half hidden.

If you could label only the hard ones, you would need far fewer labels for the same result.

How the model picks

Train a model on whatever labels you have. Let it look at everything unlabelled and say how sure it is.

Then pick the images it is least sure about, and send those for labelling.

   label a small starter set
            |
            v
      train a model
            |
            v
   score everything unlabelled
            |
            v
   pick the least certain ones  ->  send to a person
            |                              |
            +<-----------------------------+
                     repeat

Each round the model gets better, so its choices get better too.

The measure that works best

The obvious measure is "how confident is the top answer". Low confidence means hard.

There is a better one. Look at the top two answers and how close they are.

If the model says seventy percent cat and twenty-five percent dog, it is genuinely torn between two things. That image sits on a boundary the model drew in the wrong place. Labelling it moves the boundary.

That measure is called margin, and in the results below it beats every other option tested.

The trap at the beginning

There is a failure that catches almost everyone.

At the very start the model knows nothing, so its uncertainty is meaningless. Picking by uncertainty then is barely better than picking at random, and can be worse.

You will see this measured next. One popular method loses to random picking in the early rounds.

So start with a decent random batch, then switch. Do not start with active learning on day one.

The other trap

Uncertain images cluster. Ask for the twenty-five most uncertain and you may get twenty-five near-identical pictures of the same confusing thing.

You paid for twenty-five labels and learned roughly one thing.

The fix is to also insist the batch is varied. Pick the uncertain ones, then spread the selection out across them.

Where you have seen this

  • Apps that ask you to confirm exactly the faces they are unsure about.
  • Voice assistants asking you to repeat one word.
  • Spam filters asking about borderline messages, never about obvious ones.

Remember this

  • Label the images the model finds hardest, not the ones that arrived first.
  • Margin, the gap between the top two answers, beats plain confidence.
  • Start with a random batch. Uncertainty means nothing until the model knows something.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install numpy==1.26.4 scikit-learn==1.7.2

Four strategies, five seeds, one honest table

active_learning.py
import numpy as np, time
from sklearn.datasets import load_digits
from sklearn.linear_model import LogisticRegression
from sklearn.preprocessing import StandardScaler
from sklearn.pipeline import make_pipeline
from sklearn.model_selection import train_test_split
from sklearn.cluster import KMeans

X, y = load_digits(return_X_y=True)
Xpool, Xte, ypool, yte = train_test_split(X, y, test_size=0.3, random_state=0, stratify=y)
model = lambda: make_pipeline(StandardScaler(), LogisticRegression(max_iter=1000))

def entropy(p):  return -(p * np.log(p + 1e-12)).sum(1)
def margin(p):   s = np.sort(p, 1); return -(s[:, -1] - s[:, -2])   # smaller gap = more useful
def least_conf(p): return 1 - p.max(1)

def acquire(strategy, seed, rounds=8, batch=25, start=30):
    rng = np.random.default_rng(seed)
    labelled = list(rng.choice(len(Xpool), start, replace=False))   # the seed set nobody can avoid
    scores = []
    for r in range(rounds):
        m = model().fit(Xpool[labelled], ypool[labelled])
        scores.append(m.score(Xte, yte))
        rest = np.setdiff1d(np.arange(len(Xpool)), labelled)
        if strategy == "random":
            pick = rng.choice(rest, batch, replace=False)
        else:
            p = m.predict_proba(Xpool[rest])
            u = {"entropy": entropy, "margin": margin, "least-confident": least_conf}[strategy](p)
            if strategy == "margin":                       # add a diversity step on top
                top = rest[np.argsort(-u)[:batch * 4]]
                km = KMeans(batch, n_init=4, random_state=seed).fit(Xpool[top])
                pick = [top[np.argmin(((Xpool[top] - c) ** 2).sum(1))] for c in km.cluster_centers_]
            else:
                pick = rest[np.argsort(-u)[:batch]]
        labelled += list(dict.fromkeys(pick))
    m = model().fit(Xpool[labelled], ypool[labelled])
    scores.append(m.score(Xte, yte))
    return scores, len(labelled)

t0 = time.time()
print("labels used:      30    55    80   105   130   155   180   205   230")
for s in ("random", "least-confident", "entropy", "margin"):
    runs = np.array([acquire(s, seed)[0] for seed in range(5)])     # 5 seeds, then average
    print(f"{s:16s}" + "".join(f"{v:6.1%}" for v in runs.mean(0)))
full = model().fit(Xpool, ypool).score(Xte, yte)
print(f"\nall {len(Xpool)} labels: {full:.1%}   (seconds: {time.time()-t0:.0f})")
Output
labels used:      30    55    80   105   130   155   180   205   230
random           71.7% 81.5% 85.7% 89.0% 89.9% 90.6% 91.0% 91.5% 92.6%
least-confident  71.7% 80.7% 84.9% 88.9% 91.4% 92.6% 93.8% 94.9% 95.5%
entropy          71.7% 79.5% 82.6% 86.1% 88.5% 90.6% 92.3% 93.5% 94.8%
margin           71.7% 86.6% 90.2% 92.7% 93.4% 94.4% 94.5% 95.4% 96.1%

all 1257 labels: 97.2%   (seconds: 4)

Exact numbers depend on your scikit-learn build. The ordering and the cold-start effect reproduce.

Reading the output

All four start identically at 71.7% because the first 30 labels are random for everyone. That is the correct experimental design, and papers that skip it produce misleading curves.

Margin with diversity wins throughout. At 230 labels it reaches 96.1% against random's 92.6%. Read that horizontally instead: random needs well past 230 labels to reach what margin achieves at 105. That is the saving, and it is real.

Entropy loses to random for the first three rounds. 79.5% against 81.5% at 55 labels, 82.6% against 85.7% at 80. This is the cold-start problem, made visible. With ten classes and 30 labels, entropy is maximised by images the model has no purchase on. Those are often outliers rather than informative examples. It only becomes competitive once the model knows something.

Entropy never catches margin here. Entropy sums uncertainty over all ten classes. An image spread thinly over many classes scores high, even with no single decision boundary at stake. Margin looks at the two classes actually in contention, which is where a label changes the model.

230 of 1257 labels, eighteen percent, reaches 96.1% against 97.2% for everything. That one percentage point is what the remaining 82 percent of the labelling budget would buy.

What the diversity step is doing

The margin strategy here does not take the top 25 by margin. It takes the top 100, clusters them into 25 groups, and picks one image nearest each cluster centre.

Without that step, the 25 most uncertain images are frequently 25 near-copies of one confusing case. You paid for 25 labels and learned one thing.

This matters only in batch mode. If you could retrain after every single label, pure uncertainty would be fine. The model would immediately stop being uncertain about that region. Nobody retrains after every label, so diversity is not optional.

The honest caveats from the literature

Two findings that published comparisons agree on, and that vendor material tends to omit.

Augmentation shrinks the advantage. Empirical evaluations report that gradient-embedding methods such as BADGE beat entropy sampling without data augmentation. They lose that advantage once strong augmentation is used. Augmentation is standard practice in vision, so the gains measured in unaugmented studies overstate what you will see.

Diversity-only methods are weak on their own. Feature-diversity approaches such as Core-Set are reported as performing barely better than random on imbalanced problems. Diversity is a corrective on top of uncertainty, not a substitute for it.

Measure it on your own data before committing. The measurement costs one afternoon. Run random against your chosen strategy for five rounds. If the curves overlap, use random and spend the effort elsewhere.

Common mistakes

Starting with active learning. The first batch must be random, and should be large enough for the model to be meaningfully trained. Otherwise you are selecting on noise.

Ignoring class balance. Uncertainty sampling can starve a rare class it has learned to ignore. Enforce a minimum per class in each batch, or use a class-balanced acquisition.

Evaluating on a test set built from your own acquisitions. The test set must be random and fixed from the start. Actively acquired images are deliberately unrepresentative.

Retraining from scratch every round on a large model. Warm-starting is much cheaper. It carries a risk: the model gets stuck in whatever it believed at round one. On small pools, retrain; on large ones, warm-start and check the two agree occasionally.

Reporting a single run. The table above averages five seeds. Single-run differences of two or three percentage points are noise at these sizes.

Try it yourself

Add a strategy that picks the images the model is most confident about. Watch the curve flatten below random. That inverted control is the best sanity check that your uncertainty measure is wired up correctly. It takes one line.

What to learn next

Researcher — Mathematics and papers.

The framing

Pool-based active learning starts with three things. A small labelled set $\mathcal{L}$, a large unlabelled pool $\mathcal{U}$, and a budget $B$. Choose a subset $\mathcal{S} \subset \mathcal{U}$ with $|\mathcal{S}| = B$ to be labelled. The goal is minimising the risk of the model trained on $\mathcal{L} \cup \mathcal{S}$. The problem is combinatorial, and every practical method is a greedy heuristic on a surrogate objective.

Acquisition functions

Uncertainty family. For predicted distribution $p(y \mid \mathbf{x})$:

$$ \text{least confident:} \quad \arg\max_{\mathbf{x}} \; 1 - \max_y p(y \mid \mathbf{x}) $$ $$ \text{margin:} \quad \arg\min_{\mathbf{x}} \; p(\hat{y}_1 \mid \mathbf{x}) - p(\hat{y}2 \mid \mathbf{x}) $$ $$ \text{entropy:} \quad \arg\max{\mathbf{x}} \; -\sum_y p(y \mid \mathbf{x}) \log p(y \mid \mathbf{x}) $$

$\hat{y}_1, \hat{y}_2$ are the two highest-probability classes. The three coincide for binary problems and diverge as the number of classes grows. Margin targets the decision boundary between the two contending classes. Entropy rewards mass spread over many classes. For large label spaces that selects images the model has no purchase on.

Disagreement family. BALD (Houlsby et al., 2011) maximises the mutual information between the label and the model parameters:

$$ \mathbb{I}[y, \boldsymbol{\theta} \mid \mathbf{x}, \mathcal{D}] = \mathbb{H}\big[\mathbb{E}{p(\boldsymbol{\theta}|\mathcal{D})}[p(y \mid \mathbf{x}, \boldsymbol{\theta})]\big] - \mathbb{E}{p(\boldsymbol{\theta}|\mathcal{D})}\big[\mathbb{H}[p(y \mid \mathbf{x}, \boldsymbol{\theta})]\big] $$

This isolates epistemic uncertainty, reducible by more data, from aleatoric uncertainty, inherent noise that no label will fix. That distinction is the theoretical case for BALD over entropy. A genuinely ambiguous image has high entropy and low mutual information. Labelling it teaches nothing. Gal et al. (2017) implement it with MC dropout; Kirsch et al. (2019), BatchBALD, extend it to batches by accounting for the joint mutual information. That is expensive, and it is the correct treatment of redundancy.

Diversity family. Core-Set (Sener and Savarese, 2018) casts selection as a $k$-centre problem in feature space. The covering radius bounds the generalisation gap. Purely geometric, with no reference to the model's predictions.

Hybrid. BADGE (Ash et al., 2020) computes the gradient of the loss with respect to the final layer. The predicted label acts as a pseudo-label. Its magnitude encodes uncertainty and its direction encodes the class involved. Running k-means++ seeding on those embeddings selects points that are uncertain and mutually diverse. No hyperparameter balances the two.

What the careful comparisons found

The literature contains a large gap between reported gains and reproducible ones.

Munjal et al., Towards Robust and Reproducible Active Learning Using Neural Networks, use a matched, well-tuned training pipeline. The differences between published acquisition functions then shrink dramatically. Several fail to beat random sampling with statistical significance. Variance across random seeds exceeds the reported differences between methods.

Effective evaluations of deep active learning on image classification report the same pattern, with more nuance. BADGE outperforms entropy sampling without data augmentation. It loses that advantage once augmentation is applied. Feature-diversity methods such as Core-Set perform barely better than random on class-imbalanced problems. Gradient-based methods with k-centre are generally the strongest, with entropy lagging early.

The methodological requirements that follow are not optional:

  • Identical training recipe, hyperparameters and augmentation across all strategies.
  • Multiple seeds with variance reported, since single-run curves are uninterpretable.
  • The same random initial labelled set for every strategy.
  • A random baseline on the same plot, always.
  • A fixed, randomly drawn test set that is never part of the pool.

Failure modes with named causes

Cold start. With a very small $\mathcal{L}$ the model is uncalibrated and its uncertainty is uninformative. Visible in the developer table as entropy losing to random for three rounds.

Sampling bias. Actively acquired data is not i.i.d. from the input distribution, so the labelled set becomes systematically unrepresentative. This breaks any downstream use that assumes representativeness, including estimating class priors and calibration.

Batch redundancy. Top-$B$ by any pointwise score selects correlated points. BatchBALD and BADGE address it in principle; clustering the top candidates addresses it cheaply.

Model-strategy coupling. A dataset acquired under model $A$ is optimised for model $A$'s uncertainty landscape. Changing architecture later can lose the advantage entirely, and this is rarely reported.

Outlier attraction. Uncertainty scores are maximised by corrupt, blank or out-of-distribution images. In production pools this is not a corner case; filter the pool first.

When it is worth it

Active learning helps most when labelling is expensive relative to compute. It also wants a large, highly redundant pool and a cheaply retrained model. It helps least when labels are cheap, the pool is small, or classes are balanced. Strong augmentation and self-supervised pretraining also hurt it. Both reduce the marginal value of an extra label.

Given the reproducibility findings, the defensible engineering position is a small experiment. Run a random-versus-strategy comparison on your own data for a few rounds. Adopt the strategy only if it wins by more than the seed variance.

References

What to learn next