Evaluating Vision Models

Detection error analysis

A low mAP is a symptom, not a diagnosis. Sorting every wrong box into wrong-label, wrong-place, duplicate, invented or missed tells you what to fix.

On this page 9
  1. The short answer
  2. The analogy you have lived
  3. Why one number is not enough
  4. The five piles
  5. Why the piles are not equal
  6. Where you have seen this
  7. The honest part
  8. Remember this
  9. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

Error analysis sorts a detector's mistakes into named piles, so you know which one to work on.

The analogy you have lived

You get your exam paper back. Forty out of a hundred. That number tells you nothing you can act on.

Then you read the paper. Twenty marks lost by misreading what the question asked. Fifteen lost to careless arithmetic. Five lost on a page you never turned to.

Now you know what to do. Three different mistakes, three different fixes, and one of them is worth four times the others.

A detector's score is the forty. Error analysis is reading the paper.

Why one number is not enough

Two detectors can score the same and be broken in completely different ways.

One of them draws boxes in exactly the right places but calls every motorbike a bicycle. The other knows what a motorbike is but draws boxes around the rider's shoulder.

The first has a labelling problem. More classes of training data, or a better classification head, will help. The second has a placement problem. Better box regression or higher input resolution will help.

Give them the same score and you will try the same fix on both. One of those attempts is wasted weeks.

The five piles

For every box your model produced, ask two questions. Is the label right? Is the box in the right place? The answers give you five named piles.

   model drew a box
        |
        +-- right label, right place, first one there  ->  correct
        +-- right label, right place, someone beat it  ->  duplicate
        +-- wrong label, right place                   ->  classification error
        +-- right label, box too loose                 ->  localisation error
        +-- nothing is there at all                    ->  invented (background)

   model drew nothing where something was              ->  missed

Those last two are the dangerous ones. Invented means the model reported an object in empty space. Missed means a real object got no box at all, and no confidence score to lower.

Why the piles are not equal

Here is the part people skip. Fixing a pile does not always buy you what you expect.

Turning a classification error into a correct box gains you recall. A real object that was being called the wrong thing is now found.

Removing an invented box gains you precision, because you stop pointing at nothing.

Removing a duplicate gains you almost nothing on some metrics, and a lot on others. It depends on where the duplicate sat in the confidence ranking.

So the honest question is not "which pile is biggest". It is "which pile, if I emptied it, would move the number I care about".

Where you have seen this

  • A parking camera that reads the same car twice and bills you twice.
  • A shelf-scanning robot that reports biscuits where there is a shadow.
  • A crowd-counting screen at a station that misses people wearing dark clothes.
  • A helmet-detection system that boxes the rider's head but calls it a bag.

Each of those is a different pile, and each got fixed differently.

The honest part

The piles are defined by thresholds you choose. Move the "close enough" line and a localisation error becomes a correct box, or an invented one.

There is no natural, objective boundary. The standard choices are conventions that make results comparable, nothing more. Report the thresholds alongside the counts, always.

Remember this

  • A single score cannot tell you what is wrong, only that something is.
  • Every wrong box lands in one of five piles: duplicate, wrong label, wrong place, invented, or missed.
  • Fix the pile that would move the number you actually care about.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install numpy

The classification below follows TIDE (Bolya et al., ECCV 2020), which is the standard framing. TIDE uses a foreground threshold of 0.5 IoU and a background threshold of 0.1: below 0.1 overlap with anything, the box is counted as invented rather than badly placed.

Sorting every box into its pile

error_piles.py
import numpy as np
from collections import Counter

# (image, class, box)
GT = [
    (1, "car",    [10, 10, 60, 50]),
    (1, "person", [70, 20, 90, 80]),
    (2, "car",    [20, 20, 70, 60]),
    (2, "person", [80, 30, 95, 85]),
    (3, "car",    [30, 30, 80, 70]),
]

# (image, class, box, score)
PRED = [
    (1, "car",    [11, 12, 59, 51], 0.95),   # good
    (1, "car",    [70, 21, 90, 79], 0.88),   # right box, wrong class
    (2, "car",    [40, 40, 90, 80], 0.82),   # right class, sloppy box
    (2, "person", [79, 31, 96, 84], 0.77),   # good
    (2, "person", [80, 29, 94, 86], 0.71),   # duplicate of the same person
    (3, "car",    [10, 10, 25, 25], 0.66),   # nothing is there
    # image 3's car is detected by nothing -> a miss
]

TF, TB = 0.5, 0.1          # TIDE's foreground and background IoU thresholds


def iou(a, b):
    ix1, iy1 = max(a[0], b[0]), max(a[1], b[1])
    ix2, iy2 = min(a[2], b[2]), min(a[3], b[3])
    inter = max(0, ix2 - ix1) * max(0, iy2 - iy1)
    ua = (a[2]-a[0])*(a[3]-a[1]) + (b[2]-b[0])*(b[3]-b[1]) - inter
    return inter / ua


def classify_errors(preds, gts):
    used = set()
    rows = []
    for img, cls, box, score in sorted(preds, key=lambda p: -p[3]):
        same, other = 0.0, 0.0
        same_g = None
        for g, (gimg, gcls, gbox) in enumerate(gts):
            if gimg != img:
                continue
            v = iou(box, gbox)
            if gcls == cls:
                if v > same:
                    same, same_g = v, g
            else:
                other = max(other, v)

        if same >= TF and same_g not in used:
            used.add(same_g)
            label = "correct"
        elif same >= TF:
            label = "duplicate"          # right object, someone got there first
        elif other >= TF:
            label = "classification"     # box is fine, label is wrong
        elif same >= TB:
            label = "localisation"       # right label, box too loose
        elif max(same, other) < TB:
            label = "background"         # invented an object
        else:
            label = "other"
        rows.append((img, cls, score, round(same, 2), round(other, 2), label))

    missed = [(g[0], g[1]) for i, g in enumerate(gts) if i not in used]
    return rows, missed


rows, missed = classify_errors(PRED, GT)

print(f"{'img':>3} {'class':<7} {'score':>5} {'IoU same':>8} {'IoU other':>9}  verdict")
for img, cls, score, s, o, label in rows:
    print(f"{img:>3} {cls:<7} {score:>5.2f} {s:>8.2f} {o:>9.2f}  {label}")

print("\nerror counts:")
for name, n in Counter(r[5] for r in rows).most_common():
    print(f"  {name:<15} {n}")
print(f"  {'missed (FN)':<15} {len(missed)}   {missed}")

# What would mAP-style recall look like if we fixed one error type at a time?
n_gt = len(GT)
correct = sum(1 for r in rows if r[5] == "correct")
print(f"\nrecall now:                      {correct}/{n_gt} = {correct/n_gt:.2f}")
for fix in ("classification", "localisation"):
    gain = sum(1 for r in rows if r[5] == fix)
    print(f"recall if {fix:<14} fixed: {(correct+gain)}/{n_gt} = {(correct+gain)/n_gt:.2f}")
Output
img class   score IoU same IoU other  verdict
  1 car      0.95     0.89      0.00  correct
  1 car      0.88     0.00      0.97  classification
  2 car      0.82     0.18      0.16  localisation
  2 person   0.77     0.85      0.00  correct
  2 person   0.71     0.90      0.00  duplicate
  3 car      0.66     0.00      0.00  background

error counts:
  correct         2
  classification  1
  localisation    1
  duplicate       1
  background      1
  missed (FN)     3   [(1, 'person'), (2, 'car'), (3, 'car')]

recall now:                      2/5 = 0.40
recall if classification fixed: 3/5 = 0.60
recall if localisation   fixed: 3/5 = 0.60

What that table is telling you

Two IoU columns, not one. IoU same is the best overlap with a ground-truth box of the predicted class. IoU other is the best overlap with a box of any different class. The pair is what separates "wrong label" from "wrong place", and no single IoU number can do it.

Row two is the clearest case. IoU same is 0.00 and IoU other is 0.97. The model put a near-perfect box around the person in image 1 and called it a car. Nothing is wrong with the detector's eyes. Its vocabulary is wrong.

Row three is the opposite failure. IoU same is 0.18 — it found the car, and drew a box that is a third too big and shifted down. The label is right and the geometry is bad.

Row six invented something. Both overlaps are 0.00. There is nothing near that box. In production these are the boxes that wake people up at night, because a downstream system acts on them.

The three missed objects never appear in the prediction table at all. That is why you have to enumerate ground truth separately. A detector that outputs nothing has zero false positives and looks flawless on any metric that only inspects its output.

The last three lines are the decision, not the diagnosis. Fixing all classification errors and fixing all localisation errors each move recall from 0.40 to 0.60 here. They are worth the same. In a real dataset they will not be, and that ratio is what tells you where to spend the next two weeks.

Turning this into a habit

Three views are worth building once and reusing on every project.

Per-class counts. Break the piles down by ground-truth class. A single class contributing most of the missed objects is a data problem, not a model problem.

Per-size counts. Split ground truth by area, at 32 by 32 and 96 by 96 pixels as COCO does. Missed objects concentrated in the small bucket point at input resolution or feature-pyramid levels, not at your loss function.

The invented boxes, sorted by confidence. Look at the top twenty by eye. They are almost never random. They are one reflection, one poster, one shadow, repeated hundreds of times, and they are usually fixable with a few hundred targeted negatives added to your training data.

Common mistakes

Analysing errors at one confidence threshold and reporting a metric computed over all of them. mAP integrates over every threshold; your error table used a single cut. Say which cut you used, and check that the piles look similar at a couple of others.

Treating duplicates as harmless because non-maximum suppression will remove them. NMS removes overlapping boxes of the same class. A duplicate labelled differently survives it, and shows up as a classification error you cannot suppress.

Forgetting that ground truth is wrong too. Before rebuilding the model, look at fifty missed objects. On most real datasets a meaningful share of them are annotation errors: unlabelled objects, boxes on the wrong thing, or genuinely ambiguous cases. See data labelling.

Comparing error counts across datasets. The piles are counts, not rates. A dataset with more objects per image produces more of everything. Normalise by the number of ground-truth instances before comparing.

Try it yourself

Add a second person box in image 1 that overlaps the first at IoU 0.55 with score 0.90. Predict which pile it lands in before running. Then change TB from 0.1 to 0.3 and watch a localisation error turn into an invented box.

What to learn next

Researcher — Mathematics and papers.

The TIDE decomposition

Bolya et al. (2020), TIDE: A General Toolbox for Identifying Object Detection Errors, define six error types using two IoU thresholds, a foreground $t_f = 0.5$ and a background $t_b = 0.1$. For a detection with predicted class $\hat{c}$, let

$$ \mathrm{IoU}{\max}^{=} = \max{g \,:\, c_g = \hat{c}} \mathrm{IoU}(b, b_g), \qquad \mathrm{IoU}{\max}^{\ne} = \max{g \,:\, c_g \ne \hat{c}} \mathrm{IoU}(b, b_g) $$

TypeCondition
Classification$\mathrm{IoU}_{\max}^{=} < t_b$ and $\mathrm{IoU}_{\max}^{\ne} \ge t_f$
Localisation$t_b \le \mathrm{IoU}_{\max}^{=} < t_f$
Both$t_b \le \mathrm{IoU}_{\max}^{\ne} < t_f$ and $\mathrm{IoU}_{\max}^{=} < t_b$
Duplicate$\mathrm{IoU}_{\max}^{=} \ge t_f$ but the matched ground truth is already claimed
Background$\mathrm{IoU}_{\max} < t_b$ for every ground-truth box
Missedground truth unmatched and not already explained by a classification or localisation error

The Both category — right neither in label nor in place — is the one the runnable example above folds into localisation and background for brevity. It is usually the smallest pile.

Why TIDE reports $\Delta\mathrm{AP}$, not counts

The contribution that matters is not how many errors of a type exist. It is how much AP would rise if that type were repaired and nothing else changed:

$$ \Delta\mathrm{AP}{t} = \mathrm{AP}{\text{oracle}(t)} - \mathrm{AP} $$

where $\mathrm{AP}_{\text{oracle}(t)}$ recomputes AP after applying a per-type oracle: relabelling for classification, snapping the box to the ground truth for localisation, deleting for duplicate and background, and adding a perfect detection for missed.

These oracles are not additive. Repairing localisation errors can create duplicates, and repairing classification errors changes the ranking that AP integrates over. TIDE's own paper reports that the sum of individual $\Delta\mathrm{AP}$ values typically exceeds the gap to a perfect score. Treat them as a ranking of opportunities, not a budget.

The predecessor, and what changed

Hoiem et al. (2012), Diagnosing Error in Object Detectors, established the approach: false positives split into localisation, confusion with similar categories, confusion with dissimilar categories, and background. The TIDE contribution was making the decomposition (a) independent of the number of detections evaluated, and (b) computable without per-image manual inspection, which is what made it practical on COCO-scale data.

Where the decomposition is weakest

Threshold sensitivity. The boundary between localisation and background at $t_b = 0.1$ is arbitrary. Reporting the pile sizes at two or three values of $t_b$ costs nothing and immediately shows whether a conclusion is real.

Crowd and ignore regions. COCO's iscrowd boxes and Open Images' group-of boxes are matched under different rules. A detector firing inside a crowd region should be neither correct nor a false positive. Decomposition code that ignores these flags systematically overcounts background errors on pedestrian datasets.

Annotation noise is not a category. Every missed object is attributed to the model. Ma et al. (2022) and the LVIS federated-annotation design both address the fact that a substantial fraction of "missed" detections on large datasets are correct detections of unlabelled objects. On any dataset with sparse annotation, missed is an upper bound on real misses, not a measurement.

No calibration view. The decomposition says nothing about whether confidence scores are meaningful. A detector whose background errors all score above 0.9 is a different engineering problem from one whose background errors score 0.2, and the piles are identical. Pair this analysis with calibration work.

Practical protocol

  1. Compute the piles at the operating threshold you plan to ship at, not only over the full ranking.
  2. Break each pile down by class and by object area.
  3. Sample fifty from the largest pile and look at the images. Every serious detection project has found a systematic annotation or data-collection defect this way.
  4. Only then decide between more data, higher resolution, a different head, or a different loss.

Papers

What to learn next