Object Detection in Depth

Intersection over union

IoU is the overlap between two boxes divided by the area they cover together, and it is the number that decides whether a detection counts as correct.

On this page 8
  1. The short answer
  2. The analogy
  3. Why it exists
  4. How it works
  5. The part that trips everyone
  6. Where you have already seen it
  7. Remember this
  8. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

IoU is a score from zero to one saying how well two rectangles cover each other.

The analogy

Think about two people spreading bedsheets on the same floor. If one sheet lands exactly on the other, they match perfectly. If one lands half off, part of the floor is covered twice and part is bare.

To score the match, you compare the doubly-covered patch against the whole area the two sheets touch between them. A big shared patch and a small total area means a good match.

That comparison is intersection over union, usually shortened to IoU. Intersection is the shared patch. Union is everything either sheet covers.

Why it exists

A detector draws a box. The dataset says where the real box is. Somebody has to decide whether the detector was right.

Asking for the exact same four numbers is hopeless. No detector lands on the pixel. Asking for "roughly right" needs a definition of roughly. It has to be the same definition for a small bird and a large bus.

IoU gives that. It does not care about pixel counts or object size. It asks one question: of everything these two rectangles touch, what fraction do they share?

How it works

     +-----------------+
     |  truth          |
     |        +--------+--------+
     |        | shared |        |
     |        | patch  | guess  |
     +--------+--------+        |
              |                 |
              +-----------------+

     IoU  =  the shared patch  compared against
             everything both rectangles cover

A perfect overlap scores one. Two boxes that do not touch score zero. Everything useful happens in between.

Most benchmarks call a detection correct when the score passes one half. That number is a convention, chosen by people, not something the mathematics handed down.

The part that trips everyone

Two boxes that miss each other completely both score zero. A near miss and a wild miss look identical.

That is fine when you are marking answers. It is a real problem when you are teaching a model. Zero gives it no hint about which direction to move. Fixing that is why several improved versions of IoU exist.

Where you have already seen it

  • A photo app deciding two face boxes are the same face.
  • A traffic camera counting each vehicle once instead of five times.
  • A stock-taking tool checking a shelf photo against a planogram.
  • Every leaderboard number ever published for a detector.

Remember this

  • IoU compares what two boxes share against everything they cover together.
  • One means perfect, zero means no contact, and a half is the usual pass mark.
  • Two boxes that do not touch both score zero, which is why plain IoU alone is a poor teacher.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install torch torchvision

Written against torch 2.5.1 and torchvision 0.20.1. generalized_box_iou, distance_box_iou and complete_box_iou all live in torchvision.ops.

IoU by hand, then checked against the library

iou.py
import torch
from torchvision.ops import box_iou, generalized_box_iou, distance_box_iou, complete_box_iou

def iou_by_hand(a, b):
    ix1, iy1 = max(a[0], b[0]), max(a[1], b[1])         # overlap starts at the later start
    ix2, iy2 = min(a[2], b[2]), min(a[3], b[3])         # and ends at the earlier end
    iw, ih = max(0.0, ix2 - ix1), max(0.0, iy2 - iy1)   # clamp, or non-overlap goes negative
    inter = iw * ih
    area_a = (a[2] - a[0]) * (a[3] - a[1])
    area_b = (b[2] - b[0]) * (b[3] - b[1])
    return inter / (area_a + area_b - inter)

truth, guess = [100., 100., 200., 200.], [130., 130., 230., 230.]
print("by hand    :", round(iou_by_hand(truth, guess), 4))
print("torchvision:", round(box_iou(torch.tensor([truth]), torch.tensor([guess])).item(), 4))

gt = torch.tensor([[100., 100., 200., 200.]])
guesses = torch.tensor([
    [100., 100., 200., 200.],   # perfect
    [110., 110., 210., 210.],   # small shift
    [150., 150., 250., 250.],   # half off
    [300., 300., 400., 400.],   # nowhere near
])
names = ["perfect", "small shift", "half off", "miles away"]
print(f"\n{'guess':<12}{'IoU':>8}{'GIoU':>8}{'DIoU':>8}{'CIoU':>8}")
for name, i, g, d, c in zip(names,
                            box_iou(guesses, gt).squeeze(1),
                            generalized_box_iou(guesses, gt).squeeze(1),
                            distance_box_iou(guesses, gt).squeeze(1),
                            complete_box_iou(guesses, gt).squeeze(1)):
    print(f"{name:<12}{i:8.3f}{g:8.3f}{d:8.3f}{c:8.3f}")

# Why plain IoU cannot steer training on its own.
near = torch.tensor([[300., 300., 400., 400.]])
far  = torch.tensor([[500., 500., 600., 600.]])
print("\ntwo wrong boxes, one much worse than the other:")
print("  IoU :", box_iou(near, gt).item(), "and", box_iou(far, gt).item())
print("  GIoU:", round(generalized_box_iou(near, gt).item(), 3),
      "and", round(generalized_box_iou(far, gt).item(), 3))

# Identical IoU, different kinds of wrong.
gt2 = torch.tensor([[0., 0., 100., 100.]])
squashed = torch.tensor([[0., 25., 100., 75.]])    # right centre, wrong shape
shifted  = torch.tensor([[50., 0., 100., 100.]])   # shifted to the right
print("\nboth score IoU 0.5:")
for label, box in [("right centre", squashed), ("shifted right", shifted)]:
    print(f"  {label:<14} IoU {box_iou(box, gt2).item():.3f}"
          f"  GIoU {generalized_box_iou(box, gt2).item():.3f}"
          f"  DIoU {distance_box_iou(box, gt2).item():.3f}"
          f"  CIoU {complete_box_iou(box, gt2).item():.3f}")
Output
by hand    : 0.3245
torchvision: 0.3245

guess            IoU    GIoU    DIoU    CIoU
perfect        1.000   1.000   1.000   1.000
small shift    0.681   0.664   0.672   0.672
half off       0.143  -0.079   0.032   0.032
miles away     0.000  -0.778  -0.444  -0.444

two wrong boxes, one much worse than the other:
  IoU : 0.0 and 0.0
  GIoU: -0.778 and -0.92

both score IoU 0.5:
  right centre   IoU 0.500  GIoU 0.500  DIoU 0.500  CIoU 0.497
  shifted right  IoU 0.500  GIoU 0.500  DIoU 0.469  CIoU 0.466

Reading that output line by line

The hand version and the library agree to four decimals. Write the four-line version once so the operation stops being magic. Then use box_iou in real code, because it is vectorised and computes an entire N-by-M matrix in one call.

The two boxes that miss both score exactly 0.0. One is 100 pixels away and one is 300 pixels away, and IoU says the same thing about both. A gradient of zero teaches nothing. GIoU separates them: -0.778 against -0.920.

GIoU adds a penalty for empty space. It takes the smallest rectangle enclosing both boxes, and subtracts the fraction of that rectangle neither box occupies. Its range is -1 to 1 instead of 0 to 1.

GIoU is not always enough. Look at the last block. Both boxes score IoU 0.5, and both score GIoU 0.5 as well. GIoU collapses to plain IoU when one box sits entirely inside the other. The enclosing rectangle is then the outer box.

DIoU separates them: 0.500 against 0.469. DIoU adds a penalty for the distance between the two centres. The squashed box is centred correctly; the shifted box is not. CIoU adds one more term for aspect-ratio mismatch, which is why the squashed box drops slightly to 0.497.

Which one to use, and when

UsePickWhy
Evaluation, mAPplain IoUIt is the definition the benchmarks use. Do not change it.
Matching anchors to targetsplain IoUAssignment is a comparison, not a gradient.
Box regression lossCIoU or GIoUThey give useful gradients when boxes do not overlap.
Non-maximum suppressionplain IoU or DIoUDIoU-NMS keeps nearby distinct objects better.

Loss versions live under separate names: generalized_box_iou_loss, distance_box_iou_loss, complete_box_iou_loss. They return $1 - \text{metric}$ so lower is better, and they accept a reduction argument like any other loss.

The IoU matrix, which is what you actually use

iou_matrix.py
import torch
from torchvision.ops import box_iou

preds = torch.tensor([[ 10.,  10., 110., 110.],
                      [200., 200., 260., 300.],
                      [ 15.,  12., 105., 108.]])
truth = torch.tensor([[ 12.,   8., 108., 112.],
                      [205., 195., 265., 305.]])

m = box_iou(preds, truth)          # shape (3, 2): every prediction against every truth
print(m.round(decimals=3))
print("\nbest ground-truth match for each prediction:", m.argmax(dim=1).tolist())
print("best IoU per prediction               :", [round(v, 3) for v in m.max(dim=1).values.tolist()])
print("was each ground-truth object found at 0.5?",
      (m.max(dim=0).values >= 0.5).tolist())
Output
tensor([[0.9240, 0.0000],
        [0.0000, 0.7750],
        [0.8650, 0.0000]])

best ground-truth match for each prediction: [0, 1, 0]
best IoU per prediction               : [0.924, 0.775, 0.865]
was each ground-truth object found at 0.5? [True, True]

Predictions 0 and 2 both match ground-truth object 0. Only one of them can be counted correct; the other becomes a false positive. That is the job of non-maximum suppression, and of the matching step inside mAP.

Common mistakes

Passing boxes in the wrong format. box_iou expects xyxy. Hand it COCO xywh and it returns numbers that look plausible and are wrong. Convert first; see bounding box formats.

Forgetting the clamp. Without max(0, ...) on the overlap width and height, boxes that miss on both axes give two negative numbers. Their product is positive. You get a large, entirely fictional IoU.

Dividing by zero. A degenerate box with zero area makes the union zero. Add a small epsilon, or filter degenerate boxes out before this point.

Using a CIoU loss for evaluation. Benchmarks are defined on plain IoU. Reporting mAP computed with a different overlap measure makes your number incomparable with everyone else's.

Assuming 0.5 is meaningful. At IoU 0.5 a box can be visibly wrong. COCO reports the average over thresholds from 0.50 to 0.95 for exactly that reason.

Try it yourself

Take two identical 100-by-100 boxes and slide one of them right, one pixel at a time, printing IoU each step. Find the offset where IoU crosses 0.5 and the offset where it hits zero. Then repeat with 20-by-20 boxes. The pixel offset that ruins a small object is far smaller. That single experiment explains most of small object detection.

What to learn next

Researcher — Mathematics and papers.

Definition

For boxes $A$ and $B$ as subsets of the image plane:

$$ \mathrm{IoU}(A, B) = \frac{|A \cap B|}{|A \cup B|} = \frac{|A \cap B|}{|A| + |B| - |A \cap B|} $$

Where $|\cdot|$ is area. This is the Jaccard index. It is scale-invariant and lies in $[0, 1]$. It is also a proper metric: $1 - \mathrm{IoU}$ satisfies the triangle inequality (Kosub, 2019).

Scale invariance is the property that makes it usable across a benchmark containing both birds and buses. No pixel-distance measure has it.

Why it fails as a loss

Two failures, both structural.

Zero gradient on disjoint boxes. If $A \cap B = \emptyset$ then $\mathrm{IoU} = 0$ for every disjoint configuration, so $\nabla \mathrm{IoU} = 0$ everywhere in that region. Early in training most predictions are disjoint from most targets.

Plateaus at equal overlap. Many distinct configurations share an IoU value, so the loss surface has large flat regions.

Generalized IoU

Rezatofighi et al. (2019) add a term for the smallest axis-aligned box $C$ enclosing both:

$$ \mathrm{GIoU}(A, B) = \mathrm{IoU}(A, B) - \frac{|C \setminus (A \cup B)|}{|C|} $$

The second term is the fraction of the enclosing box that neither input occupies. Range is $(-1, 1]$. It is $-1$ in the limit as the boxes separate to infinity, and equals IoU when the union fills $C$.

Its known degeneracy: when $A \subseteq B$, $C = B$, so $|C \setminus (A \cup B)| = 0$ and $\mathrm{GIoU} = \mathrm{IoU}$. Enclosure carries no extra signal. The output block above is exactly this case.

Distance and Complete IoU

Zheng et al. (2020) replace the enclosure term with a normalised centre distance:

$$ \mathcal{L}_{\mathrm{DIoU}} = 1 - \mathrm{IoU} + \frac{\rho^2(\mathbf{b}, \mathbf{b}^{gt})}{c^2} $$

Where $\rho(\cdot,\cdot)$ is Euclidean distance between box centres and $c$ is the diagonal length of the smallest enclosing box $C$. The ratio is bounded in $[0, 1)$, so the loss stays comparable across scales.

CIoU adds an aspect-ratio term:

$$ \mathcal{L}{\mathrm{CIoU}} = \mathcal{L}{\mathrm{DIoU}} + \alpha v, \qquad v = \frac{4}{\pi^2}\left(\arctan\frac{w^{gt}}{h^{gt}} - \arctan\frac{w}{h}\right)^2, \qquad \alpha = \frac{v}{(1 - \mathrm{IoU}) + v} $$

Where $w, h$ are the predicted width and height and $w^{gt}, h^{gt}$ the target's. $v$ measures aspect-ratio disagreement. $\alpha$ is a trade-off weight. It switches the aspect term off when overlap is poor, so the model fixes position before shape. $\alpha$ is treated as a constant during backpropagation.

DIoU also gives DIoU-NMS, which subtracts the centre-distance term from the suppression criterion. Two distinct objects that overlap heavily but have well-separated centres survive, which helps in crowded scenes.

Later variants

  • EIoU (Zhang et al., 2022) replaces the aspect-ratio term with separate width and height penalties. Their argument is that $v$ is zero whenever the ratio matches, even if both edges are wrong.
  • SIoU (Gevorgyan, 2022) adds an angle term steering the prediction onto the nearest axis before closing the distance.
  • alpha-IoU (He et al., 2021) raises IoU to a power, $\mathcal{L} = 1 - \mathrm{IoU}^{\alpha}$, sharpening the gradient on high-quality boxes.
  • Wise-IoU (Tong et al., 2023) attenuates the loss on low-quality examples, which matters when labels are noisy.

Gains between these are small and dataset-dependent. Treat the choice as a hyperparameter. Be sceptical of any paper claiming a large mAP jump from a loss term alone.

Cost

box_iou on $N$ predictions and $M$ targets is $O(NM)$ time and memory. With $N = 100{,}000$ anchors and $M = 50$ targets that is five million pairs, which is fine. With $N = M = 100{,}000$ it is $10^{10}$ entries, which is not. Chunk the computation, or restrict candidate pairs by centre distance first.

Papers

What to learn next