Evaluating Vision Models

Segmentation metrics

IoU and Dice score how well two shapes agree, and both exist because pixel accuracy lies whenever the thing you care about is small.

On this page 10
  1. The short answer
  2. The analogy you have lived
  3. Why pixel accuracy is a trap
  4. Dice, the other one
  5. The other trap: averaging
  6. Edges are where models really fail
  7. Where you have seen this
  8. The honest part
  9. Remember this
  10. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

IoU and Dice both measure how much two shapes overlap, as a number between zero and one.

The analogy you have lived

Someone spills tea on a white tablecloth. You lay tracing paper over it and draw around the stain with a pencil. Your friend does the same on a second sheet.

Now hold the two sheets up to the light. The part where both pencil lines enclose the same cloth is where you agree. The rest is where one of you was wrong.

Overlap divided by everything either of you covered is the score. Trace it perfectly and you get one. Trace a completely different corner of the cloth and you get zero.

That fraction has a name: intersection over union, usually shortened to IoU.

Why pixel accuracy is a trap

The obvious-looking measure is to count how many pixels you labelled correctly. It fails, and it fails in a way that will embarrass you.

Take a road-camera image. Nearly the whole frame is road and sky. A signboard occupies a tiny corner.

Label every single pixel "road" and you score enormously well on that count, while finding no signboard at all. The metric says you are excellent. The car drives past the sign.

   truth                     a lazy prediction
   ..........SS..            ..............
   ..........SS..            ..............
   RRRRRRRRRRRRRR            RRRRRRRRRRRRRR
   RRRRRRRRRRRRRR            RRRRRRRRRRRRRR

   S = signboard, R = road, . = background
   most pixels are right, and the sign is gone

IoU refuses to be fooled, because it is computed per class. The signboard gets its own score, computed only over signboard pixels. Losing the sign gives a signboard IoU near zero, and that number goes into the average with the same weight as road.

Dice, the other one

Dice is the second name you will meet. It measures the same agreement, weighted a little differently: it counts the shared part twice, once for each tracing.

Dice is always at least as large as IoU, and usually a bit larger. They rise and fall together, so no model ever wins on one and loses on the other.

Which one you see depends on the field you are in. Medical imaging papers report Dice almost always. Self-driving and satellite papers report IoU almost always. They are the same idea wearing different clothes.

The other trap: averaging

Once every class has its own score, you have to combine them.

Mean IoU treats every class equally. Rare classes count as much as common ones. This is the honest default, and it is what benchmarks report.

Frequency-weighted IoU weights each class by how many pixels it occupies. Big classes dominate. It looks kinder, and it hides exactly the failure you were trying to catch.

If somebody shows you a segmentation score without saying which averaging they used, ask.

Edges are where models really fail

Two predictions can share an IoU and look nothing alike. One is the right shape with a fuzzy border. The other is a blob in roughly the right place.

For a large object, the border is a thin ring of pixels. Getting all of it wrong barely dents the score, because the huge interior carries the number.

That is why boundary IoU exists. It throws away the interior and scores only the pixels near the edge. It is a much harsher, much more informative number for anything that will be measured, cut along, or driven beside.

Where you have seen this

  • A phone's portrait mode deciding which pixels are you and which are the wall.
  • A video call replacing your background, and eating part of your ear.
  • Satellite maps colouring in fields, water and buildings.
  • A radiologist's tool outlining an organ for a doctor to check.

The ear is a boundary error. The metric that catches it is boundary IoU.

The honest part

Every one of these numbers assumes the ground truth is right. For segmentation, that assumption is shakier than anywhere else in vision.

Ask two trained people to trace the same organ and their outlines will differ, sometimes a lot. That disagreement is a ceiling on any model score. A model reported at a Dice of 0.90 where humans agree at 0.88 has not beaten humans. It has run out of measurable room.

Remember this

  • IoU is overlap divided by combined area. Dice is the same agreement counted slightly differently.
  • Pixel accuracy hides small-class failures. Per-class IoU does not.
  • Look at boundary scores as well; the interior of a big object flatters everything.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install numpy

Both metrics come out of one confusion matrix, so build that first and read everything off it.

Every metric from one confusion matrix

seg_metrics.py
import numpy as np

# 0 = background, 1 = road, 2 = signboard.  Rows are the top of the image first.
truth = np.array([
    [0,0,0,0,0,0,0,0,0,0,0,0],
    [0,0,0,0,0,0,0,0,0,2,2,0],
    [0,0,0,0,0,0,0,0,0,2,2,0],
    [1,1,1,1,1,1,1,1,1,1,1,1],
    [1,1,1,1,1,1,1,1,1,1,1,1],
    [1,1,1,1,1,1,1,1,1,1,1,1],
])

pred = np.array([
    [0,0,0,0,0,0,0,0,0,0,0,0],
    [0,0,0,0,0,0,0,0,0,0,2,0],   # found half the signboard
    [0,0,0,0,0,0,0,0,0,0,0,0],   # lost the rest of it
    [0,0,0,0,0,0,0,0,0,0,0,0],   # road starts one row late
    [1,1,1,1,1,1,1,1,1,1,1,1],
    [1,1,1,1,1,1,1,1,1,1,1,1],
])

NAMES = {0: "background", 1: "road", 2: "signboard"}


def confusion(t, p, k):
    m = np.zeros((k, k), dtype=int)
    for a, b in zip(t.ravel(), p.ravel()):
        m[a, b] += 1
    return m


cm = confusion(truth, pred, 3)
print("confusion matrix (rows = truth, cols = prediction)")
print(cm)

inter = np.diag(cm).astype(float)
union = cm.sum(1) + cm.sum(0) - inter
pred_n = cm.sum(0).astype(float)
true_n = cm.sum(1).astype(float)

iou = inter / union
dice = 2 * inter / (true_n + pred_n)

print(f"\n{'class':<11} {'pixels':>6} {'IoU':>6} {'Dice':>6} {'2I/(1+I)':>9}")
for c in range(3):
    print(f"{NAMES[c]:<11} {int(true_n[c]):>6} {iou[c]:>6.3f} {dice[c]:>6.3f} "
          f"{2*iou[c]/(1+iou[c]):>9.3f}")

acc = np.diag(cm).sum() / cm.sum()
miou = iou.mean()
fw = (true_n / true_n.sum() * iou).sum()
print(f"\npixel accuracy      {acc:.3f}")
print(f"mean IoU            {miou:.3f}")
print(f"frequency-weighted  {fw:.3f}")

# Boundary IoU: keep only pixels within 1 cell of the mask edge, then re-score.
def boundary(mask):
    pad = np.pad(mask, 1, constant_values=0)
    eroded = np.ones_like(mask, dtype=bool)
    for dy in (-1, 0, 1):
        for dx in (-1, 0, 1):
            eroded &= pad[1+dy:1+dy+mask.shape[0], 1+dx:1+dx+mask.shape[1]].astype(bool)
    return mask.astype(bool) & ~eroded


for c in (1, 2):
    t, p = truth == c, pred == c
    bt, bp = boundary(t), boundary(p)
    b_iou = (bt & bp).sum() / max((bt | bp).sum(), 1)
    print(f"\n{NAMES[c]}: mask IoU {iou[c]:.3f}   boundary IoU {b_iou:.3f}")
Output
confusion matrix (rows = truth, cols = prediction)
[[32  0  0]
 [12 24  0]
 [ 3  0  1]]

class       pixels    IoU   Dice  2I/(1+I)
background      32  0.681  0.810     0.810
road            36  0.667  0.800     0.800
signboard        4  0.250  0.400     0.400

pixel accuracy      0.792
mean IoU            0.533
frequency-weighted  0.650

road: mask IoU 0.667   boundary IoU 0.389

signboard: mask IoU 0.250   boundary IoU 0.250

Reading the output carefully

Pixel accuracy 0.792, mean IoU 0.533. Same prediction, two very different impressions. The gap is entirely the signboard, which is four pixels out of seventy-two and scores 0.250. Pixel accuracy gave those four pixels four votes out of seventy-two. Mean IoU gave them a third of the final number.

Frequency-weighted IoU sits at 0.650, right between them. It is the flattering option. Every time you see a segmentation result quoted without the averaging rule, assume it is this one until proved otherwise.

The fourth column proves Dice and IoU are the same information. 2I/(1+I) reproduces the Dice column exactly, for all three classes. There is an exact algebraic relationship between them. Reporting both is not extra evidence; it is the same evidence written twice.

Boundary IoU on road is 0.389 against a mask IoU of 0.667. The road prediction is one row short at the top. That single missing row is a small fraction of the road's area, and almost all of the road's boundary. This is the number that would have told you the shape was wrong.

Signboard boundary IoU equals its mask IoU, at 0.250. The signboard is only two pixels across. Every pixel in it is boundary, so the two measures coincide. Boundary metrics stop being informative for objects smaller than a few times the boundary width, and that is worth knowing before you report them on small objects.

Doing this at scale

Never accumulate per-image IoU and average it. Accumulate the confusion matrix and compute IoU once at the end.

python
totals = np.zeros((n_classes, n_classes), dtype=np.int64)
for t, p in dataset:                    # one image at a time
    totals += np.bincount(
        (t.ravel() * n_classes + p.ravel()), minlength=n_classes ** 2
    ).reshape(n_classes, n_classes)

The bincount trick is the fast version of the double loop above and is what every serious implementation uses. Per-image averaging breaks in a specific, unpleasant way: an image containing none of class c gives that class a zero-over-zero IoU, and whatever you substitute for it — zero, one, or a skip — silently changes your headline number.

Common mistakes

Ignoring the void label. Cityscapes, ADE20K and most real datasets have pixels marked "unlabelled" or "ignore". They must be dropped before the confusion matrix, not counted as a class. Forgetting this is the most common cause of an mIoU that will not match the published baseline.

Using Dice loss and then reporting IoU as if it were independent. Optimising Dice directly is fine and often helps with class imbalance. It also means your validation IoU is measuring the thing you trained on, so it stops being an independent check.

Comparing mIoU across datasets with different class counts. A nineteen-class benchmark and a hundred-and-fifty-class benchmark produce numbers that have nothing to say to each other.

Resizing predictions with bilinear interpolation. Class indices are labels, not quantities. Interpolating between class 3 and class 7 produces class 5, which may be a completely unrelated object. Resize logits, or resize masks with nearest-neighbour.

Thresholding a probability map at 0.5 by default. For imbalanced classes the threshold that maximises Dice is often well below 0.5. Tune it on validation data and report the value you used.

Try it yourself

Shift the predicted road up by one row so it matches the truth, and watch mask IoU and boundary IoU move by very different amounts. Then make the signboard prediction two pixels wide but one row too low, so its mask IoU is zero while three of its four pixels are "nearly" right. That is the case that motivates tolerance-based boundary metrics.

What to learn next

Researcher — Mathematics and papers.

Definitions

For a class $c$ with prediction set $P$ and ground-truth set $G$ over pixels:

$$ \mathrm{IoU}_c = \frac{|P \cap G|}{|P \cup G|} = \frac{\mathrm{TP}}{\mathrm{TP} + \mathrm{FP} + \mathrm{FN}} $$

$$ \mathrm{Dice}_c = \frac{2|P \cap G|}{|P| + |G|} = \frac{2\,\mathrm{TP}}{2\,\mathrm{TP} + \mathrm{FP} + \mathrm{FN}} $$

Dice is the $F_1$ score computed over pixels. The two are related by a strictly increasing bijection on $[0,1]$:

$$ \mathrm{Dice} = \frac{2\,\mathrm{IoU}}{1 + \mathrm{IoU}}, \qquad \mathrm{IoU} = \frac{\mathrm{Dice}}{2 - \mathrm{Dice}} $$

So they induce identical rankings over models. They differ in aggregation: averaging IoU over images and averaging Dice over images give different orderings, because the mean of a nonlinear function is not that function of the mean.

Aggregation choices, and why they matter more than the metric

$$ \mathrm{mIoU} = \frac{1}{|C|}\sum_{c} \mathrm{IoU}_c, \qquad \mathrm{fwIoU} = \sum_c \frac{|G_c|}{\sum_{c'} |G_{c'}|}\,\mathrm{IoU}_c $$

A third option, sometimes called dataset-level or aggregate IoU, pools TP, FP and FN across all images before dividing. Cityscapes reports this. Per-image averaging is a fourth. On long-tailed data these four can differ by ten points or more on the same predictions, which is why cross-paper comparison requires checking the evaluation script, not the metric name.

Panoptic quality

For panoptic segmentation, Kirillov et al. (2019) define

$$ \mathrm{PQ} = \underbrace{\frac{\sum_{(p,g) \in \mathrm{TP}} \mathrm{IoU}(p,g)}{|\mathrm{TP}|}}{\text{segmentation quality}} \times \underbrace{\frac{|\mathrm{TP}|}{|\mathrm{TP}| + \frac{1}{2}|\mathrm{FP}| + \frac{1}{2}|\mathrm{FN}|}}{\text{recognition quality}} $$

Segments match when IoU exceeds 0.5, which makes the matching unique. The factorisation is the useful part: it separates did you find the object from did you outline it well, which mIoU cannot.

Boundary-aware metrics

Boundary IoU (Cheng et al., 2021) computes IoU restricted to a band of width $d$ around each mask's contour:

$$ \mathrm{BIoU}_d = \frac{|(G_d \cap G) \cap (P_d \cap P)|}{|(G_d \cap G) \cup (P_d \cap P)|} $$

where $X_d$ denotes the set of pixels within distance $d$ of $\partial X$. Unlike Trimap IoU it is sensitive to boundary errors on large objects, and unlike the earlier F-measure of Perazzi et al. (2016) it does not saturate.

The related Hausdorff distance and its 95th-percentile variant HD95 are standard in medical imaging:

$$ d_H(P, G) = \max\left{ \sup_{p \in P} \inf_{g \in G} |p - g|,\; \sup_{g \in G} \inf_{p \in P} |p - g| \right} $$

HD95 exists because the plain supremum is decided by a single outlier pixel and is unusable in practice.

Losses versus metrics

Soft Dice loss (Milletari et al., 2016, V-Net) relaxes the sets to probabilities:

$$ \mathcal{L}_{\text{Dice}} = 1 - \frac{2\sum_i p_i g_i + \epsilon}{\sum_i p_i + \sum_i g_i + \epsilon} $$

It is differentiable and handles foreground-background imbalance far better than pixel-wise cross-entropy. Three cautions carry into evaluation. It is unstable when the ground truth for a class is empty in a batch, and the $\epsilon$ chosen then determines the loss value. Its gradients are non-local, so a single confident wrong pixel moves the whole map. And Bertels et al. (2019) show that optimising soft Dice yields systematically miscalibrated probabilities, which matters if anything downstream consumes the confidence rather than the argmax.

Tversky loss (Salehi et al., 2017) generalises Dice with asymmetric FP and FN weights $\alpha, \beta$, and is the standard tool when a missed lesion costs far more than a false alarm.

The annotation ceiling

Inter-annotator Dice is the honest upper bound. Published figures for organ segmentation commonly sit in the 0.85 to 0.95 range depending on structure and modality, and for lesions considerably lower. Joskowicz et al. (2019) and the various MICCAI challenge reports document this directly. Any model result within the inter-observer band is a measurement of the annotation protocol as much as of the model, and the correct response is to report the variability alongside the score rather than to claim superiority.

Papers

What to learn next