Mean average precision
mAP squeezes a detector's whole ranked list of guesses into one number, by asking how much of it is right and how much it invented.
- 16 min read
- 3 reading levels
- Updated
Read these first
On this page 10
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
Mean average precision is one number for how well a detector finds things without inventing things.
The analogy you have lived
You are on a crowded railway platform, looking for a relative. You ask the people around you to point them out. Each person points somewhere and tells you how sure they are.
You follow the most confident finger first, then the next. Some fingers land on your relative. Some land on a stranger. And your relative may be standing behind a pillar, where nobody points at all.
Two things decide whether the crowd helped you. How many of the fingers were right. And how many of your relatives got found at all.
The two words you need
Precision means: of everything the model pointed at, how much was real? Point at forty things when three exist, and precision is dreadful.
Recall means: of everything that was really there, how much did the model find? Find one relative out of four, and recall is dreadful.
You can win either one alone by cheating. Point at one thing you are certain about, and precision is perfect while recall is near zero. Draw a box around every square inch, and recall is perfect while precision collapses.
A useful score has to hold both at once.
Why "average" precision
A detector does not give a yes or no. It gives a list, sorted from most confident to least.
So you walk down the list, one guess at a time, keeping score.
rank guess real? precision so far recall so far
---- ------------- ----- ---------------- -------------
1 "car, 95%" yes 1.00 0.25
2 "car, 90%" yes 1.00 0.50
3 "car, 80%" yes 1.00 0.75
4 "car, 70%" no 0.75 0.75
5 "car, 60%" no 0.60 0.75Average precision is the summary of that whole walk. It rewards a model whose good guesses sit at the top of the list. It punishes a model that mixes rubbish in among them.
Notice what the table already tells you. The wrong guesses at ranks four and five arrived after every real car was found. They cost nothing at all.
The "close enough" rule
There is one more question. If the true box is here and the guess is a few pixels off, is that a hit?
The referee is intersection over union, or IoU. You look at how much the two boxes overlap, compared to how much space they cover together. Full overlap scores one. No overlap scores zero.
truth ┌──────────┐ overlap ÷ combined area
│ ┌────┼─────┐ = how much they agree
│ │////│ │
└─────┼────┘ │ lots of /// → score near one
└──────────┘ barely any → score near zero
guessPick a cut-off, say a bit over half. Above it, the guess counts as correct. Below it, the guess counts as invented.
And finally "mean"
Two averages get taken on top.
The first runs over the cut-off. A generous cut-off flatters sloppy boxes. So the COCO benchmark scores the detector ten times, from a lenient rule to a strict one. Then it averages.
The second runs over the classes. Cars, people, dogs and traffic lights each get their own score, and the mean is reported. A model that nails cars and ignores cyclists cannot hide behind the car number.
That is the whole of mAP. Rank the guesses, walk the list, decide what counts as close enough, then average over strictness and over classes.
What you have already used
- A phone gallery grouping photos by what is in them.
- Toll-gate cameras reading number plates.
- Shop checkouts that recognise the fruit you placed on the scale.
- Reversing cameras that draw a box around the child behind your car.
Every one of those was shipped only after somebody stared at an mAP table.
The honest part
mAP is a leaderboard number, not a safety number. It says nothing about which mistakes matter.
A model that misses one pedestrian in a hundred can score the same as one missing a traffic cone. Your users will not feel the same way about those two.
Treat mAP as a first filter. Then look at the mistakes themselves.
Remember this
- Precision is how much of what you pointed at was real. Recall is how much of what was real you found.
- Average precision summarises the whole ranked list, not one confidence setting.
- Mean average precision averages that over strictness cut-offs and over classes.
What to learn next
- Detection error analysis — splitting a low mAP into the mistakes that caused it.
- Object detection — the task this metric scores.
- Model evaluation — precision and recall in the simpler classification setting.
Developer — Code and libraries.
Setup
pip install numpyThat is all this needs. The point of writing it out is that mAP stops being a mystery once you have seen the fifty lines it takes.
The whole metric, from scratch
import numpy as np
# Ground truth: (image_id, box). Boxes are [x1, y1, x2, y2].
GT = [
(1, [10, 10, 50, 50]),
(1, [60, 60, 90, 90]),
(2, [20, 20, 60, 60]),
(2, [70, 20, 95, 45]), # nothing will ever detect this one
]
# Predictions: (image_id, box, confidence score).
PRED = [
(1, [12, 12, 52, 52], 0.95),
(2, [22, 18, 58, 62], 0.90),
(1, [58, 58, 88, 88], 0.80),
(1, [70, 10, 95, 40], 0.70), # nothing is there: a false positive
(2, [21, 19, 59, 61], 0.60), # second box on an already-detected object
(1, [11, 11, 49, 49], 0.50), # ditto
]
def iou(a, b):
ix1, iy1 = max(a[0], b[0]), max(a[1], b[1])
ix2, iy2 = min(a[2], b[2]), min(a[3], b[3])
inter = max(0, ix2 - ix1) * max(0, iy2 - iy1)
area_a = (a[2] - a[0]) * (a[3] - a[1])
area_b = (b[2] - b[0]) * (b[3] - b[1])
return inter / (area_a + area_b - inter)
def match(preds, gts, thr):
"""Greedy matching, highest confidence first. Returns 1 for TP, 0 for FP."""
order = sorted(range(len(preds)), key=lambda i: -preds[i][2])
taken = set() # each ground-truth box can be used once
flags = []
for i in order:
img, box, _ = preds[i]
best, best_g = 0.0, None
for g, (gimg, gbox) in enumerate(gts):
if gimg != img or g in taken:
continue
v = iou(box, gbox)
if v > best:
best, best_g = v, g
if best >= thr:
taken.add(best_g)
flags.append(1)
else:
flags.append(0)
return np.array(flags), [preds[i][2] for i in order]
def pr_curve(flags, n_gt):
tp = np.cumsum(flags)
fp = np.cumsum(1 - flags)
recall = tp / n_gt
precision = tp / (tp + fp)
return precision, recall
def ap_101(precision, recall):
"""COCO's 101-point interpolated average precision."""
rec_thrs = np.linspace(0.0, 1.0, 101)
# for each recall level, the best precision achieved at that recall or beyond
p_interp = [precision[recall >= r].max() if (recall >= r).any() else 0.0
for r in rec_thrs]
return float(np.mean(p_interp))
flags, scores = match(PRED, GT, 0.5)
prec, rec = pr_curve(flags, len(GT))
print("detections sorted by confidence, at IoU 0.50")
print("score TP? precision recall")
for s, f, p, r in zip(scores, flags, prec, rec):
print(f"{s:.2f} {f} {p:9.3f} {r:6.3f}")
print(f"\nAP@0.50 = {ap_101(prec, rec):.4f}")
print("\nAP at each COCO IoU threshold:")
aps = []
for thr in np.arange(0.5, 1.0, 0.05):
f, _ = match(PRED, GT, thr)
p, r = pr_curve(f, len(GT))
a = ap_101(p, r)
aps.append(a)
print(f" IoU {thr:.2f} -> AP {a:.4f}")
print(f"\nmAP@[.50:.95] = {np.mean(aps):.4f}")
# Now move the false positive to the top of the ranking and nothing else.
LOUD = [(i, b, 0.99 if s == 0.70 else s) for i, b, s in PRED]
f, _ = match(LOUD, GT, 0.5)
p, r = pr_curve(f, len(GT))
print(f"\nsame boxes, the false positive now scores 0.99: AP@0.50 = {ap_101(p, r):.4f}")detections sorted by confidence, at IoU 0.50 score TP? precision recall 0.95 1 1.000 0.250 0.90 1 1.000 0.500 0.80 1 1.000 0.750 0.70 0 0.750 0.750 0.60 0 0.600 0.750 0.50 0 0.500 0.750 AP@0.50 = 0.7525 AP at each COCO IoU threshold: IoU 0.50 -> AP 0.7525 IoU 0.55 -> AP 0.7525 IoU 0.60 -> AP 0.7525 IoU 0.65 -> AP 0.7525 IoU 0.70 -> AP 0.7525 IoU 0.75 -> AP 0.7525 IoU 0.80 -> AP 0.5050 IoU 0.85 -> AP 0.1683 IoU 0.90 -> AP 0.1683 IoU 0.95 -> AP 0.0000 mAP@[.50:.95] = 0.5356 same boxes, the false positive now scores 0.99: AP@0.50 = 0.5644
Reading that output properly
The ceiling is set by what you missed. One of the four ground-truth boxes is never detected. Recall stops at 0.750 and stays there. Since 76 of the 101 recall points sit at or below 0.75, and precision is a perfect 1.000 at each of them, AP lands at 76/101 = 0.7525. Missing a quarter of the objects caps you near three quarters. No confidence tuning recovers it.
Low-confidence false positives were free. Three of the six predictions were wrong. They still cost nothing, because every real object had already been found by the time they appeared. This is the single most misread property of AP. Lowering your confidence threshold to dump more boxes into the output can raise AP even while making the model useless to a human reading its output.
Move one false positive to the top and AP drops by 0.19. Same boxes. Same count of errors. Only the ranking changed, from 0.7525 to 0.5644. AP is a measure of ordering, not of counts.
The IoU sweep is where sloppy boxes get caught. AP is flat at 0.7525 up to IoU 0.75, then falls off a cliff. Our boxes are two pixels out. At the strict end that stops being close enough, and the model that looked fine scores zero. mAP@[.50:.95] of 0.5356 is the honest summary.
The parts that trip people up
taken = set() is the reason a duplicate box is a false positive. Each ground-truth object may be claimed once. Remove that line and a detector spraying ten boxes on one car scores ten true positives.
Matching walks the predictions in confidence order, not in file order. A greedy match by confidence is what COCO does, and it is why a confident-but-slightly-worse box can steal the ground truth from a less confident, better-fitting one.
precision[recall >= r].max() is the interpolation step. Raw precision-recall curves are saw-toothed; each new true positive bumps precision up. COCO smooths this by taking, at each recall level, the best precision available at that recall or higher. The 101 points come straight from cocoeval.py, where recThrs = np.linspace(0, 1, 101).
Real COCO evaluation adds two things this code leaves out. It caps detections per image at 100, and it reports AP separately for small, medium and large objects, split at 32 by 32 and 96 by 96 pixels. Small-object AP is usually where the pain is.
Use the real thing in production
pip install pycocotoolsfrom pycocotools.coco import COCO
from pycocotools.cocoeval import COCOeval
coco_gt = COCO("instances_val.json")
coco_dt = coco_gt.loadRes("detections.json")
e = COCOeval(coco_gt, coco_dt, iouType="bbox")
e.evaluate(); e.accumulate(); e.summarize()No output block for this one. It depends entirely on your annotation file and your model, and an invented table would teach you to expect numbers that will not appear.
Two practical notes. detections.json must be a flat list of {"image_id", "category_id", "bbox", "score"}, with bbox in [x, y, width, height] — not the corner format used above. And category_id must match the ground-truth file's ids, not your model's class indices. Silently mismatched ids are the most common cause of a suspiciously low mAP.
Common mistakes
Running non-maximum suppression too aggressively before evaluating. NMS removes overlapping boxes. Set the threshold too tight and you delete real detections of two people standing close together, which shows up as poor recall you will blame on the model.
Comparing an mAP@0.5 number to a mAP@[.50:.95] number. They are different metrics. Pascal VOC style mAP@0.5 numbers look far higher. Papers quote both, and press releases quote whichever is bigger.
Filtering low-confidence detections before submitting. Since low-ranked false positives are nearly free, cutting them off can only lose you recall. Submit everything and let the metric sort the ranking out. Filter for humans, not for the metric.
Averaging per-image mAP. mAP is computed once over the whole dataset, pooling every detection. Averaging a per-image score gives a different, wrong number, and images containing no objects of a class break it entirely.
Try it yourself
Add a fifth ground-truth box in image 1 and a matching high-confidence prediction. Predict what happens to AP before you run it. Then shrink every predicted box by four pixels on each side and watch the IoU sweep collapse while AP@0.50 barely moves.
What to learn next
- Detection error analysis — splitting a low mAP into the mistakes that caused it.
- Object detection — the task this metric scores.
- Model evaluation — precision and recall in the simpler classification setting.
Researcher — Mathematics and papers.
Definition
Let the detections for one class be ranked by confidence $s_1 \ge s_2 \ge \dots \ge s_N$. Matching each in turn against unclaimed ground truth at IoU threshold $\tau$ gives an indicator $z_i \in {0,1}$. Then
$$ P(k) = \frac{1}{k}\sum_{i=1}^{k} z_i, \qquad R(k) = \frac{1}{G}\sum_{i=1}^{k} z_i $$
where $G$ is the number of ground-truth instances of that class. $P(k)$ is precision after $k$ detections, $R(k)$ is recall.
The raw curve is non-monotonic, so it is replaced by its upper envelope:
$$ P_{\text{interp}}(r) = \max_{k \,:\, R(k) \ge r} P(k) $$
COCO then samples that envelope at 101 evenly spaced recall points:
$$ \mathrm{AP}{\tau} = \frac{1}{101}\sum{r \in {0, 0.01, \dots, 1}} P_{\text{interp}}(r) $$
and averages over ten IoU thresholds and over classes $c$:
$$ \mathrm{mAP} = \frac{1}{|C|}\sum_{c \in C} \frac{1}{10}\sum_{\tau \in {0.50, 0.55, \dots, 0.95}} \mathrm{AP}_{\tau}^{(c)} $$
The 101-point rule and the threshold grid are literal transcriptions of Params.setDetParams in pycocotools/cocoeval.py, which defines recThrs = np.linspace(0, 1, 101) and iouThrs = np.linspace(0.5, 0.95, 10).
Intersection over union
$$ \mathrm{IoU}(A, B) = \frac{|A \cap B|}{|A \cup B|} = \frac{|A \cap B|}{|A| + |B| - |A \cap B|} $$
IoU is scale-invariant, which is desirable, but its gradient vanishes for non-overlapping boxes, which is why it is a poor training loss on its own. GIoU (Rezatofighi et al., 2019) adds a penalty proportional to the empty area of the smallest enclosing box; DIoU and CIoU (Zheng et al., 2020) add centre-distance and aspect-ratio terms. All three are training losses, not evaluation changes: the metric stays plain IoU.
Three lineages of the metric
| Benchmark | Interpolation | IoU | Notes |
|---|---|---|---|
| Pascal VOC 2007 | 11 recall points | 0.50 | Coarse; deprecated after 2010 |
| Pascal VOC 2010–12 | all recall points | 0.50 | Exact area under the envelope |
| COCO (2014–) | 101 points | 0.50:0.05:0.95 | Also reports AP_S, AP_M, AP_L |
| Open Images | all points, group-of boxes | 0.50 | Handles crowd regions and label hierarchy |
| LVIS | 101 points | 0.50:0.05:0.95 | Federated annotation; long-tail categories |
Numbers are comparable only within a row. Everingham et al. (2010), The Pascal Visual Object Classes Challenge, is still the clearest writeup of why interpolation is needed at all.
Known pathologies
AP is dominated by ranking, not by counting. Two detectors with identical sets of true and false positives score differently if their confidence orderings differ. This is a feature for retrieval and a liability for deployment, where a fixed operating threshold is what actually ships.
Low-scoring false positives are close to free. The precision at high recall is what AP integrates over; a false positive ranked below every true positive changes nothing. Oksuz et al. (2021), Localization Recall Precision (LRP) Error, propose an alternative that penalises them, and show AP-optimal thresholds are frequently far from the deployment-optimal ones.
The IoU grid conflates two abilities. A model can lose AP by mislabelling or by localising loosely, and mAP@[.50:.95] cannot distinguish them. That is the motivation for the error decomposition covered in detection error analysis.
Greedy matching is not optimal matching. COCO assigns each detection to its best available ground truth in confidence order. A globally optimal bipartite matching would sometimes score higher. Hoiem et al. (2012), Diagnosing Error in Object Detectors, showed how sensitive the resulting error attribution is to this choice.
Class averaging hides the tail. Under LVIS-style long-tail distributions, rare classes contribute equally to mAP while contributing a handful of instances, so their AP estimates carry enormous variance. Gupta et al. (2019), LVIS, discuss the federated-annotation workaround.
Cost
Evaluation is $O(N \log N)$ for the sort plus $O(N \cdot G)$ for naive matching per class per threshold. pycocotools computes the IoU matrix once per image and reuses it across all ten thresholds, which is the only optimisation that matters in practice. On COCO val2017 with 100 detections per image, evaluation takes seconds, not minutes; if yours takes minutes, you are recomputing IoU.
Papers
- Everingham et al., The Pascal VOC Challenge, IJCV 2010 — doi.org/10.1007/s11263-009-0275-4
- Lin et al., Microsoft COCO: Common Objects in Context, ECCV 2014 — arxiv.org/abs/1405.0312
- Hoiem et al., Diagnosing Error in Object Detectors, ECCV 2012
- Rezatofighi et al., Generalized Intersection over Union, CVPR 2019 — arxiv.org/abs/1902.09630
- Gupta et al., LVIS: A Dataset for Large Vocabulary Instance Segmentation, CVPR 2019 — arxiv.org/abs/1908.03195
- Oksuz et al., One Metric to Measure Them All: Localisation Recall Precision, TPAMI 2021 — arxiv.org/abs/2011.10772
What to learn next
- Detection error analysis — splitting a low mAP into the mistakes that caused it.
- Object detection — the task this metric scores.
- Model evaluation — precision and recall in the simpler classification setting.