Evaluating Vision Models

Mean average precision

mAP squeezes a detector's whole ranked list of guesses into one number, by asking how much of it is right and how much it invented.

On this page 10
  1. The short answer
  2. The analogy you have lived
  3. The two words you need
  4. Why "average" precision
  5. The "close enough" rule
  6. And finally "mean"
  7. What you have already used
  8. The honest part
  9. Remember this
  10. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

Mean average precision is one number for how well a detector finds things without inventing things.

The analogy you have lived

You are on a crowded railway platform, looking for a relative. You ask the people around you to point them out. Each person points somewhere and tells you how sure they are.

You follow the most confident finger first, then the next. Some fingers land on your relative. Some land on a stranger. And your relative may be standing behind a pillar, where nobody points at all.

Two things decide whether the crowd helped you. How many of the fingers were right. And how many of your relatives got found at all.

The two words you need

Precision means: of everything the model pointed at, how much was real? Point at forty things when three exist, and precision is dreadful.

Recall means: of everything that was really there, how much did the model find? Find one relative out of four, and recall is dreadful.

You can win either one alone by cheating. Point at one thing you are certain about, and precision is perfect while recall is near zero. Draw a box around every square inch, and recall is perfect while precision collapses.

A useful score has to hold both at once.

Why "average" precision

A detector does not give a yes or no. It gives a list, sorted from most confident to least.

So you walk down the list, one guess at a time, keeping score.

   rank  guess          real?   precision so far   recall so far
   ----  -------------  -----   ----------------   -------------
    1    "car, 95%"      yes           1.00            0.25
    2    "car, 90%"      yes           1.00            0.50
    3    "car, 80%"      yes           1.00            0.75
    4    "car, 70%"      no            0.75            0.75
    5    "car, 60%"      no            0.60            0.75

Average precision is the summary of that whole walk. It rewards a model whose good guesses sit at the top of the list. It punishes a model that mixes rubbish in among them.

Notice what the table already tells you. The wrong guesses at ranks four and five arrived after every real car was found. They cost nothing at all.

The "close enough" rule

There is one more question. If the true box is here and the guess is a few pixels off, is that a hit?

The referee is intersection over union, or IoU. You look at how much the two boxes overlap, compared to how much space they cover together. Full overlap scores one. No overlap scores zero.

   truth  ┌──────────┐            overlap ÷ combined area
          │     ┌────┼─────┐      = how much they agree
          │     │////│     │
          └─────┼────┘     │      lots of ///  →  score near one
                └──────────┘      barely any   →  score near zero
                guess

Pick a cut-off, say a bit over half. Above it, the guess counts as correct. Below it, the guess counts as invented.

And finally "mean"

Two averages get taken on top.

The first runs over the cut-off. A generous cut-off flatters sloppy boxes. So the COCO benchmark scores the detector ten times, from a lenient rule to a strict one. Then it averages.

The second runs over the classes. Cars, people, dogs and traffic lights each get their own score, and the mean is reported. A model that nails cars and ignores cyclists cannot hide behind the car number.

That is the whole of mAP. Rank the guesses, walk the list, decide what counts as close enough, then average over strictness and over classes.

What you have already used

  • A phone gallery grouping photos by what is in them.
  • Toll-gate cameras reading number plates.
  • Shop checkouts that recognise the fruit you placed on the scale.
  • Reversing cameras that draw a box around the child behind your car.

Every one of those was shipped only after somebody stared at an mAP table.

The honest part

mAP is a leaderboard number, not a safety number. It says nothing about which mistakes matter.

A model that misses one pedestrian in a hundred can score the same as one missing a traffic cone. Your users will not feel the same way about those two.

Treat mAP as a first filter. Then look at the mistakes themselves.

Remember this

  • Precision is how much of what you pointed at was real. Recall is how much of what was real you found.
  • Average precision summarises the whole ranked list, not one confidence setting.
  • Mean average precision averages that over strictness cut-offs and over classes.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install numpy

That is all this needs. The point of writing it out is that mAP stops being a mystery once you have seen the fifty lines it takes.

The whole metric, from scratch

map_from_scratch.py
import numpy as np

# Ground truth: (image_id, box). Boxes are [x1, y1, x2, y2].
GT = [
    (1, [10, 10, 50, 50]),
    (1, [60, 60, 90, 90]),
    (2, [20, 20, 60, 60]),
    (2, [70, 20, 95, 45]),      # nothing will ever detect this one
]

# Predictions: (image_id, box, confidence score).
PRED = [
    (1, [12, 12, 52, 52], 0.95),
    (2, [22, 18, 58, 62], 0.90),
    (1, [58, 58, 88, 88], 0.80),
    (1, [70, 10, 95, 40], 0.70),   # nothing is there: a false positive
    (2, [21, 19, 59, 61], 0.60),   # second box on an already-detected object
    (1, [11, 11, 49, 49], 0.50),   # ditto
]


def iou(a, b):
    ix1, iy1 = max(a[0], b[0]), max(a[1], b[1])
    ix2, iy2 = min(a[2], b[2]), min(a[3], b[3])
    inter = max(0, ix2 - ix1) * max(0, iy2 - iy1)
    area_a = (a[2] - a[0]) * (a[3] - a[1])
    area_b = (b[2] - b[0]) * (b[3] - b[1])
    return inter / (area_a + area_b - inter)


def match(preds, gts, thr):
    """Greedy matching, highest confidence first. Returns 1 for TP, 0 for FP."""
    order = sorted(range(len(preds)), key=lambda i: -preds[i][2])
    taken = set()                       # each ground-truth box can be used once
    flags = []
    for i in order:
        img, box, _ = preds[i]
        best, best_g = 0.0, None
        for g, (gimg, gbox) in enumerate(gts):
            if gimg != img or g in taken:
                continue
            v = iou(box, gbox)
            if v > best:
                best, best_g = v, g
        if best >= thr:
            taken.add(best_g)
            flags.append(1)
        else:
            flags.append(0)
    return np.array(flags), [preds[i][2] for i in order]


def pr_curve(flags, n_gt):
    tp = np.cumsum(flags)
    fp = np.cumsum(1 - flags)
    recall = tp / n_gt
    precision = tp / (tp + fp)
    return precision, recall


def ap_101(precision, recall):
    """COCO's 101-point interpolated average precision."""
    rec_thrs = np.linspace(0.0, 1.0, 101)
    # for each recall level, the best precision achieved at that recall or beyond
    p_interp = [precision[recall >= r].max() if (recall >= r).any() else 0.0
                for r in rec_thrs]
    return float(np.mean(p_interp))


flags, scores = match(PRED, GT, 0.5)
prec, rec = pr_curve(flags, len(GT))

print("detections sorted by confidence, at IoU 0.50")
print("score  TP?  precision  recall")
for s, f, p, r in zip(scores, flags, prec, rec):
    print(f"{s:.2f}   {f}    {p:9.3f}  {r:6.3f}")

print(f"\nAP@0.50 = {ap_101(prec, rec):.4f}")

print("\nAP at each COCO IoU threshold:")
aps = []
for thr in np.arange(0.5, 1.0, 0.05):
    f, _ = match(PRED, GT, thr)
    p, r = pr_curve(f, len(GT))
    a = ap_101(p, r)
    aps.append(a)
    print(f"  IoU {thr:.2f} -> AP {a:.4f}")
print(f"\nmAP@[.50:.95] = {np.mean(aps):.4f}")

# Now move the false positive to the top of the ranking and nothing else.
LOUD = [(i, b, 0.99 if s == 0.70 else s) for i, b, s in PRED]
f, _ = match(LOUD, GT, 0.5)
p, r = pr_curve(f, len(GT))
print(f"\nsame boxes, the false positive now scores 0.99: AP@0.50 = {ap_101(p, r):.4f}")
Output
detections sorted by confidence, at IoU 0.50
score  TP?  precision  recall
0.95   1        1.000   0.250
0.90   1        1.000   0.500
0.80   1        1.000   0.750
0.70   0        0.750   0.750
0.60   0        0.600   0.750
0.50   0        0.500   0.750

AP@0.50 = 0.7525

AP at each COCO IoU threshold:
  IoU 0.50 -> AP 0.7525
  IoU 0.55 -> AP 0.7525
  IoU 0.60 -> AP 0.7525
  IoU 0.65 -> AP 0.7525
  IoU 0.70 -> AP 0.7525
  IoU 0.75 -> AP 0.7525
  IoU 0.80 -> AP 0.5050
  IoU 0.85 -> AP 0.1683
  IoU 0.90 -> AP 0.1683
  IoU 0.95 -> AP 0.0000

mAP@[.50:.95] = 0.5356

same boxes, the false positive now scores 0.99: AP@0.50 = 0.5644

Reading that output properly

The ceiling is set by what you missed. One of the four ground-truth boxes is never detected. Recall stops at 0.750 and stays there. Since 76 of the 101 recall points sit at or below 0.75, and precision is a perfect 1.000 at each of them, AP lands at 76/101 = 0.7525. Missing a quarter of the objects caps you near three quarters. No confidence tuning recovers it.

Low-confidence false positives were free. Three of the six predictions were wrong. They still cost nothing, because every real object had already been found by the time they appeared. This is the single most misread property of AP. Lowering your confidence threshold to dump more boxes into the output can raise AP even while making the model useless to a human reading its output.

Move one false positive to the top and AP drops by 0.19. Same boxes. Same count of errors. Only the ranking changed, from 0.7525 to 0.5644. AP is a measure of ordering, not of counts.

The IoU sweep is where sloppy boxes get caught. AP is flat at 0.7525 up to IoU 0.75, then falls off a cliff. Our boxes are two pixels out. At the strict end that stops being close enough, and the model that looked fine scores zero. mAP@[.50:.95] of 0.5356 is the honest summary.

The parts that trip people up

taken = set() is the reason a duplicate box is a false positive. Each ground-truth object may be claimed once. Remove that line and a detector spraying ten boxes on one car scores ten true positives.

Matching walks the predictions in confidence order, not in file order. A greedy match by confidence is what COCO does, and it is why a confident-but-slightly-worse box can steal the ground truth from a less confident, better-fitting one.

precision[recall >= r].max() is the interpolation step. Raw precision-recall curves are saw-toothed; each new true positive bumps precision up. COCO smooths this by taking, at each recall level, the best precision available at that recall or higher. The 101 points come straight from cocoeval.py, where recThrs = np.linspace(0, 1, 101).

Real COCO evaluation adds two things this code leaves out. It caps detections per image at 100, and it reports AP separately for small, medium and large objects, split at 32 by 32 and 96 by 96 pixels. Small-object AP is usually where the pain is.

Use the real thing in production

bash
pip install pycocotools
python
from pycocotools.coco import COCO
from pycocotools.cocoeval import COCOeval

coco_gt = COCO("instances_val.json")
coco_dt = coco_gt.loadRes("detections.json")
e = COCOeval(coco_gt, coco_dt, iouType="bbox")
e.evaluate(); e.accumulate(); e.summarize()

No output block for this one. It depends entirely on your annotation file and your model, and an invented table would teach you to expect numbers that will not appear.

Two practical notes. detections.json must be a flat list of {"image_id", "category_id", "bbox", "score"}, with bbox in [x, y, width, height] — not the corner format used above. And category_id must match the ground-truth file's ids, not your model's class indices. Silently mismatched ids are the most common cause of a suspiciously low mAP.

Common mistakes

Running non-maximum suppression too aggressively before evaluating. NMS removes overlapping boxes. Set the threshold too tight and you delete real detections of two people standing close together, which shows up as poor recall you will blame on the model.

Comparing an mAP@0.5 number to a mAP@[.50:.95] number. They are different metrics. Pascal VOC style mAP@0.5 numbers look far higher. Papers quote both, and press releases quote whichever is bigger.

Filtering low-confidence detections before submitting. Since low-ranked false positives are nearly free, cutting them off can only lose you recall. Submit everything and let the metric sort the ranking out. Filter for humans, not for the metric.

Averaging per-image mAP. mAP is computed once over the whole dataset, pooling every detection. Averaging a per-image score gives a different, wrong number, and images containing no objects of a class break it entirely.

Try it yourself

Add a fifth ground-truth box in image 1 and a matching high-confidence prediction. Predict what happens to AP before you run it. Then shrink every predicted box by four pixels on each side and watch the IoU sweep collapse while AP@0.50 barely moves.

What to learn next

Researcher — Mathematics and papers.

Definition

Let the detections for one class be ranked by confidence $s_1 \ge s_2 \ge \dots \ge s_N$. Matching each in turn against unclaimed ground truth at IoU threshold $\tau$ gives an indicator $z_i \in {0,1}$. Then

$$ P(k) = \frac{1}{k}\sum_{i=1}^{k} z_i, \qquad R(k) = \frac{1}{G}\sum_{i=1}^{k} z_i $$

where $G$ is the number of ground-truth instances of that class. $P(k)$ is precision after $k$ detections, $R(k)$ is recall.

The raw curve is non-monotonic, so it is replaced by its upper envelope:

$$ P_{\text{interp}}(r) = \max_{k \,:\, R(k) \ge r} P(k) $$

COCO then samples that envelope at 101 evenly spaced recall points:

$$ \mathrm{AP}{\tau} = \frac{1}{101}\sum{r \in {0, 0.01, \dots, 1}} P_{\text{interp}}(r) $$

and averages over ten IoU thresholds and over classes $c$:

$$ \mathrm{mAP} = \frac{1}{|C|}\sum_{c \in C} \frac{1}{10}\sum_{\tau \in {0.50, 0.55, \dots, 0.95}} \mathrm{AP}_{\tau}^{(c)} $$

The 101-point rule and the threshold grid are literal transcriptions of Params.setDetParams in pycocotools/cocoeval.py, which defines recThrs = np.linspace(0, 1, 101) and iouThrs = np.linspace(0.5, 0.95, 10).

Intersection over union

$$ \mathrm{IoU}(A, B) = \frac{|A \cap B|}{|A \cup B|} = \frac{|A \cap B|}{|A| + |B| - |A \cap B|} $$

IoU is scale-invariant, which is desirable, but its gradient vanishes for non-overlapping boxes, which is why it is a poor training loss on its own. GIoU (Rezatofighi et al., 2019) adds a penalty proportional to the empty area of the smallest enclosing box; DIoU and CIoU (Zheng et al., 2020) add centre-distance and aspect-ratio terms. All three are training losses, not evaluation changes: the metric stays plain IoU.

Three lineages of the metric

BenchmarkInterpolationIoUNotes
Pascal VOC 200711 recall points0.50Coarse; deprecated after 2010
Pascal VOC 2010–12all recall points0.50Exact area under the envelope
COCO (2014–)101 points0.50:0.05:0.95Also reports AP_S, AP_M, AP_L
Open Imagesall points, group-of boxes0.50Handles crowd regions and label hierarchy
LVIS101 points0.50:0.05:0.95Federated annotation; long-tail categories

Numbers are comparable only within a row. Everingham et al. (2010), The Pascal Visual Object Classes Challenge, is still the clearest writeup of why interpolation is needed at all.

Known pathologies

AP is dominated by ranking, not by counting. Two detectors with identical sets of true and false positives score differently if their confidence orderings differ. This is a feature for retrieval and a liability for deployment, where a fixed operating threshold is what actually ships.

Low-scoring false positives are close to free. The precision at high recall is what AP integrates over; a false positive ranked below every true positive changes nothing. Oksuz et al. (2021), Localization Recall Precision (LRP) Error, propose an alternative that penalises them, and show AP-optimal thresholds are frequently far from the deployment-optimal ones.

The IoU grid conflates two abilities. A model can lose AP by mislabelling or by localising loosely, and mAP@[.50:.95] cannot distinguish them. That is the motivation for the error decomposition covered in detection error analysis.

Greedy matching is not optimal matching. COCO assigns each detection to its best available ground truth in confidence order. A globally optimal bipartite matching would sometimes score higher. Hoiem et al. (2012), Diagnosing Error in Object Detectors, showed how sensitive the resulting error attribution is to this choice.

Class averaging hides the tail. Under LVIS-style long-tail distributions, rare classes contribute equally to mAP while contributing a handful of instances, so their AP estimates carry enormous variance. Gupta et al. (2019), LVIS, discuss the federated-annotation workaround.

Cost

Evaluation is $O(N \log N)$ for the sort plus $O(N \cdot G)$ for naive matching per class per threshold. pycocotools computes the IoU matrix once per image and reuses it across all ten thresholds, which is the only optimisation that matters in practice. On COCO val2017 with 100 detections per image, evaluation takes seconds, not minutes; if yours takes minutes, you are recomputing IoU.

Papers

What to learn next