Segmentation metrics
IoU and Dice score how well two shapes agree, and both exist because pixel accuracy lies whenever the thing you care about is small.
- 14 min read
- 3 reading levels
- Updated
Read these first
On this page 10
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
IoU and Dice both measure how much two shapes overlap, as a number between zero and one.
The analogy you have lived
Someone spills tea on a white tablecloth. You lay tracing paper over it and draw around the stain with a pencil. Your friend does the same on a second sheet.
Now hold the two sheets up to the light. The part where both pencil lines enclose the same cloth is where you agree. The rest is where one of you was wrong.
Overlap divided by everything either of you covered is the score. Trace it perfectly and you get one. Trace a completely different corner of the cloth and you get zero.
That fraction has a name: intersection over union, usually shortened to IoU.
Why pixel accuracy is a trap
The obvious-looking measure is to count how many pixels you labelled correctly. It fails, and it fails in a way that will embarrass you.
Take a road-camera image. Nearly the whole frame is road and sky. A signboard occupies a tiny corner.
Label every single pixel "road" and you score enormously well on that count, while finding no signboard at all. The metric says you are excellent. The car drives past the sign.
truth a lazy prediction
..........SS.. ..............
..........SS.. ..............
RRRRRRRRRRRRRR RRRRRRRRRRRRRR
RRRRRRRRRRRRRR RRRRRRRRRRRRRR
S = signboard, R = road, . = background
most pixels are right, and the sign is goneIoU refuses to be fooled, because it is computed per class. The signboard gets its own score, computed only over signboard pixels. Losing the sign gives a signboard IoU near zero, and that number goes into the average with the same weight as road.
Dice, the other one
Dice is the second name you will meet. It measures the same agreement, weighted a little differently: it counts the shared part twice, once for each tracing.
Dice is always at least as large as IoU, and usually a bit larger. They rise and fall together, so no model ever wins on one and loses on the other.
Which one you see depends on the field you are in. Medical imaging papers report Dice almost always. Self-driving and satellite papers report IoU almost always. They are the same idea wearing different clothes.
The other trap: averaging
Once every class has its own score, you have to combine them.
Mean IoU treats every class equally. Rare classes count as much as common ones. This is the honest default, and it is what benchmarks report.
Frequency-weighted IoU weights each class by how many pixels it occupies. Big classes dominate. It looks kinder, and it hides exactly the failure you were trying to catch.
If somebody shows you a segmentation score without saying which averaging they used, ask.
Edges are where models really fail
Two predictions can share an IoU and look nothing alike. One is the right shape with a fuzzy border. The other is a blob in roughly the right place.
For a large object, the border is a thin ring of pixels. Getting all of it wrong barely dents the score, because the huge interior carries the number.
That is why boundary IoU exists. It throws away the interior and scores only the pixels near the edge. It is a much harsher, much more informative number for anything that will be measured, cut along, or driven beside.
Where you have seen this
- A phone's portrait mode deciding which pixels are you and which are the wall.
- A video call replacing your background, and eating part of your ear.
- Satellite maps colouring in fields, water and buildings.
- A radiologist's tool outlining an organ for a doctor to check.
The ear is a boundary error. The metric that catches it is boundary IoU.
The honest part
Every one of these numbers assumes the ground truth is right. For segmentation, that assumption is shakier than anywhere else in vision.
Ask two trained people to trace the same organ and their outlines will differ, sometimes a lot. That disagreement is a ceiling on any model score. A model reported at a Dice of 0.90 where humans agree at 0.88 has not beaten humans. It has run out of measurable room.
Remember this
- IoU is overlap divided by combined area. Dice is the same agreement counted slightly differently.
- Pixel accuracy hides small-class failures. Per-class IoU does not.
- Look at boundary scores as well; the interior of a big object flatters everything.
What to learn next
- Image segmentation — the models these numbers score.
- Keypoint metrics: OKS and PCK — the same problem for points instead of regions.
- Mean average precision — how instance segmentation borrows the detection metric.
Developer — Code and libraries.
Setup
pip install numpyBoth metrics come out of one confusion matrix, so build that first and read everything off it.
Every metric from one confusion matrix
import numpy as np
# 0 = background, 1 = road, 2 = signboard. Rows are the top of the image first.
truth = np.array([
[0,0,0,0,0,0,0,0,0,0,0,0],
[0,0,0,0,0,0,0,0,0,2,2,0],
[0,0,0,0,0,0,0,0,0,2,2,0],
[1,1,1,1,1,1,1,1,1,1,1,1],
[1,1,1,1,1,1,1,1,1,1,1,1],
[1,1,1,1,1,1,1,1,1,1,1,1],
])
pred = np.array([
[0,0,0,0,0,0,0,0,0,0,0,0],
[0,0,0,0,0,0,0,0,0,0,2,0], # found half the signboard
[0,0,0,0,0,0,0,0,0,0,0,0], # lost the rest of it
[0,0,0,0,0,0,0,0,0,0,0,0], # road starts one row late
[1,1,1,1,1,1,1,1,1,1,1,1],
[1,1,1,1,1,1,1,1,1,1,1,1],
])
NAMES = {0: "background", 1: "road", 2: "signboard"}
def confusion(t, p, k):
m = np.zeros((k, k), dtype=int)
for a, b in zip(t.ravel(), p.ravel()):
m[a, b] += 1
return m
cm = confusion(truth, pred, 3)
print("confusion matrix (rows = truth, cols = prediction)")
print(cm)
inter = np.diag(cm).astype(float)
union = cm.sum(1) + cm.sum(0) - inter
pred_n = cm.sum(0).astype(float)
true_n = cm.sum(1).astype(float)
iou = inter / union
dice = 2 * inter / (true_n + pred_n)
print(f"\n{'class':<11} {'pixels':>6} {'IoU':>6} {'Dice':>6} {'2I/(1+I)':>9}")
for c in range(3):
print(f"{NAMES[c]:<11} {int(true_n[c]):>6} {iou[c]:>6.3f} {dice[c]:>6.3f} "
f"{2*iou[c]/(1+iou[c]):>9.3f}")
acc = np.diag(cm).sum() / cm.sum()
miou = iou.mean()
fw = (true_n / true_n.sum() * iou).sum()
print(f"\npixel accuracy {acc:.3f}")
print(f"mean IoU {miou:.3f}")
print(f"frequency-weighted {fw:.3f}")
# Boundary IoU: keep only pixels within 1 cell of the mask edge, then re-score.
def boundary(mask):
pad = np.pad(mask, 1, constant_values=0)
eroded = np.ones_like(mask, dtype=bool)
for dy in (-1, 0, 1):
for dx in (-1, 0, 1):
eroded &= pad[1+dy:1+dy+mask.shape[0], 1+dx:1+dx+mask.shape[1]].astype(bool)
return mask.astype(bool) & ~eroded
for c in (1, 2):
t, p = truth == c, pred == c
bt, bp = boundary(t), boundary(p)
b_iou = (bt & bp).sum() / max((bt | bp).sum(), 1)
print(f"\n{NAMES[c]}: mask IoU {iou[c]:.3f} boundary IoU {b_iou:.3f}")confusion matrix (rows = truth, cols = prediction) [[32 0 0] [12 24 0] [ 3 0 1]] class pixels IoU Dice 2I/(1+I) background 32 0.681 0.810 0.810 road 36 0.667 0.800 0.800 signboard 4 0.250 0.400 0.400 pixel accuracy 0.792 mean IoU 0.533 frequency-weighted 0.650 road: mask IoU 0.667 boundary IoU 0.389 signboard: mask IoU 0.250 boundary IoU 0.250
Reading the output carefully
Pixel accuracy 0.792, mean IoU 0.533. Same prediction, two very different impressions. The gap is entirely the signboard, which is four pixels out of seventy-two and scores 0.250. Pixel accuracy gave those four pixels four votes out of seventy-two. Mean IoU gave them a third of the final number.
Frequency-weighted IoU sits at 0.650, right between them. It is the flattering option. Every time you see a segmentation result quoted without the averaging rule, assume it is this one until proved otherwise.
The fourth column proves Dice and IoU are the same information. 2I/(1+I) reproduces the Dice column exactly, for all three classes. There is an exact algebraic relationship between them. Reporting both is not extra evidence; it is the same evidence written twice.
Boundary IoU on road is 0.389 against a mask IoU of 0.667. The road prediction is one row short at the top. That single missing row is a small fraction of the road's area, and almost all of the road's boundary. This is the number that would have told you the shape was wrong.
Signboard boundary IoU equals its mask IoU, at 0.250. The signboard is only two pixels across. Every pixel in it is boundary, so the two measures coincide. Boundary metrics stop being informative for objects smaller than a few times the boundary width, and that is worth knowing before you report them on small objects.
Doing this at scale
Never accumulate per-image IoU and average it. Accumulate the confusion matrix and compute IoU once at the end.
totals = np.zeros((n_classes, n_classes), dtype=np.int64)
for t, p in dataset: # one image at a time
totals += np.bincount(
(t.ravel() * n_classes + p.ravel()), minlength=n_classes ** 2
).reshape(n_classes, n_classes)The bincount trick is the fast version of the double loop above and is what every serious implementation uses. Per-image averaging breaks in a specific, unpleasant way: an image containing none of class c gives that class a zero-over-zero IoU, and whatever you substitute for it — zero, one, or a skip — silently changes your headline number.
Common mistakes
Ignoring the void label. Cityscapes, ADE20K and most real datasets have pixels marked "unlabelled" or "ignore". They must be dropped before the confusion matrix, not counted as a class. Forgetting this is the most common cause of an mIoU that will not match the published baseline.
Using Dice loss and then reporting IoU as if it were independent. Optimising Dice directly is fine and often helps with class imbalance. It also means your validation IoU is measuring the thing you trained on, so it stops being an independent check.
Comparing mIoU across datasets with different class counts. A nineteen-class benchmark and a hundred-and-fifty-class benchmark produce numbers that have nothing to say to each other.
Resizing predictions with bilinear interpolation. Class indices are labels, not quantities. Interpolating between class 3 and class 7 produces class 5, which may be a completely unrelated object. Resize logits, or resize masks with nearest-neighbour.
Thresholding a probability map at 0.5 by default. For imbalanced classes the threshold that maximises Dice is often well below 0.5. Tune it on validation data and report the value you used.
Try it yourself
Shift the predicted road up by one row so it matches the truth, and watch mask IoU and boundary IoU move by very different amounts. Then make the signboard prediction two pixels wide but one row too low, so its mask IoU is zero while three of its four pixels are "nearly" right. That is the case that motivates tolerance-based boundary metrics.
What to learn next
- Image segmentation — the models these numbers score.
- Keypoint metrics: OKS and PCK — the same problem for points instead of regions.
- Mean average precision — how instance segmentation borrows the detection metric.
Researcher — Mathematics and papers.
Definitions
For a class $c$ with prediction set $P$ and ground-truth set $G$ over pixels:
$$ \mathrm{IoU}_c = \frac{|P \cap G|}{|P \cup G|} = \frac{\mathrm{TP}}{\mathrm{TP} + \mathrm{FP} + \mathrm{FN}} $$
$$ \mathrm{Dice}_c = \frac{2|P \cap G|}{|P| + |G|} = \frac{2\,\mathrm{TP}}{2\,\mathrm{TP} + \mathrm{FP} + \mathrm{FN}} $$
Dice is the $F_1$ score computed over pixels. The two are related by a strictly increasing bijection on $[0,1]$:
$$ \mathrm{Dice} = \frac{2\,\mathrm{IoU}}{1 + \mathrm{IoU}}, \qquad \mathrm{IoU} = \frac{\mathrm{Dice}}{2 - \mathrm{Dice}} $$
So they induce identical rankings over models. They differ in aggregation: averaging IoU over images and averaging Dice over images give different orderings, because the mean of a nonlinear function is not that function of the mean.
Aggregation choices, and why they matter more than the metric
$$ \mathrm{mIoU} = \frac{1}{|C|}\sum_{c} \mathrm{IoU}_c, \qquad \mathrm{fwIoU} = \sum_c \frac{|G_c|}{\sum_{c'} |G_{c'}|}\,\mathrm{IoU}_c $$
A third option, sometimes called dataset-level or aggregate IoU, pools TP, FP and FN across all images before dividing. Cityscapes reports this. Per-image averaging is a fourth. On long-tailed data these four can differ by ten points or more on the same predictions, which is why cross-paper comparison requires checking the evaluation script, not the metric name.
Panoptic quality
For panoptic segmentation, Kirillov et al. (2019) define
$$ \mathrm{PQ} = \underbrace{\frac{\sum_{(p,g) \in \mathrm{TP}} \mathrm{IoU}(p,g)}{|\mathrm{TP}|}}{\text{segmentation quality}} \times \underbrace{\frac{|\mathrm{TP}|}{|\mathrm{TP}| + \frac{1}{2}|\mathrm{FP}| + \frac{1}{2}|\mathrm{FN}|}}{\text{recognition quality}} $$
Segments match when IoU exceeds 0.5, which makes the matching unique. The factorisation is the useful part: it separates did you find the object from did you outline it well, which mIoU cannot.
Boundary-aware metrics
Boundary IoU (Cheng et al., 2021) computes IoU restricted to a band of width $d$ around each mask's contour:
$$ \mathrm{BIoU}_d = \frac{|(G_d \cap G) \cap (P_d \cap P)|}{|(G_d \cap G) \cup (P_d \cap P)|} $$
where $X_d$ denotes the set of pixels within distance $d$ of $\partial X$. Unlike Trimap IoU it is sensitive to boundary errors on large objects, and unlike the earlier F-measure of Perazzi et al. (2016) it does not saturate.
The related Hausdorff distance and its 95th-percentile variant HD95 are standard in medical imaging:
$$ d_H(P, G) = \max\left{ \sup_{p \in P} \inf_{g \in G} |p - g|,\; \sup_{g \in G} \inf_{p \in P} |p - g| \right} $$
HD95 exists because the plain supremum is decided by a single outlier pixel and is unusable in practice.
Losses versus metrics
Soft Dice loss (Milletari et al., 2016, V-Net) relaxes the sets to probabilities:
$$ \mathcal{L}_{\text{Dice}} = 1 - \frac{2\sum_i p_i g_i + \epsilon}{\sum_i p_i + \sum_i g_i + \epsilon} $$
It is differentiable and handles foreground-background imbalance far better than pixel-wise cross-entropy. Three cautions carry into evaluation. It is unstable when the ground truth for a class is empty in a batch, and the $\epsilon$ chosen then determines the loss value. Its gradients are non-local, so a single confident wrong pixel moves the whole map. And Bertels et al. (2019) show that optimising soft Dice yields systematically miscalibrated probabilities, which matters if anything downstream consumes the confidence rather than the argmax.
Tversky loss (Salehi et al., 2017) generalises Dice with asymmetric FP and FN weights $\alpha, \beta$, and is the standard tool when a missed lesion costs far more than a false alarm.
The annotation ceiling
Inter-annotator Dice is the honest upper bound. Published figures for organ segmentation commonly sit in the 0.85 to 0.95 range depending on structure and modality, and for lesions considerably lower. Joskowicz et al. (2019) and the various MICCAI challenge reports document this directly. Any model result within the inter-observer band is a measurement of the annotation protocol as much as of the model, and the correct response is to report the variability alongside the score rather than to claim superiority.
Papers
- Milletari et al., V-Net, 3DV 2016 — arxiv.org/abs/1606.04797
- Salehi et al., Tversky Loss Function, MLMI 2017 — arxiv.org/abs/1706.05721
- Kirillov et al., Panoptic Segmentation, CVPR 2019 — arxiv.org/abs/1801.00868
- Bertels et al., Optimizing the Dice Score and Jaccard Index for Medical Image Segmentation, MICCAI 2019 — arxiv.org/abs/1911.01685
- Cheng et al., Boundary IoU, CVPR 2021 — arxiv.org/abs/2103.16562
- Perazzi et al., A Benchmark Dataset and Evaluation Methodology for Video Object Segmentation, CVPR 2016
What to learn next
- Image segmentation — the models these numbers score.
- Keypoint metrics: OKS and PCK — the same problem for points instead of regions.
- Mean average precision — how instance segmentation borrows the detection metric.