Object Detection in Depth

Anchor boxes

Anchor boxes are a fixed grid of guessed rectangles that a detector nudges into place, which turns finding objects into correcting guesses.

On this page 8
  1. The short answer
  2. The analogy
  3. Why this had to be invented
  4. How it works
  5. The trade-off nobody tells you
  6. Where you have seen the result
  7. Remember this
  8. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

Anchor boxes are ready-made rectangles scattered across the picture, which the model then nudges onto real objects.

The analogy

Think about a tailor fitting a shirt. He does not cut cloth from nothing while you stand there. He picks a stock size off the rail, holds it against you, and takes it in at the shoulders.

Starting from a stock size is far easier than starting from bare cloth. The tailor only has to answer a small question: how much to let out, how much to take in.

An anchor box is that stock size. The model picks the nearest one and adjusts it.

Why this had to be invented

The first detectors ran a separate search to propose likely object regions, then classified each one. The search was slow and lived outside the network, so it could not learn.

Anchors removed the search. Cover the picture with a fixed grid of rectangles of several sizes and shapes. Ask the network two things about every rectangle: is there an object here, and how should this rectangle move to fit it?

Both questions have short answers. Both can be learned. Nothing outside the network has to run.

How it works

   image grid                 shapes placed at every grid point
   . . . . . . . .            +--------+   +----+   +--+
   . . . . . . . .            |  wide  |   |sqr |   |ta|
   . . . . . . . .            +--------+   +----+   |ll|
   . . . . . . . .                                  +--+

   for every point, for every shape:
        "is an object here?"   ->  yes / no
        "how do I fix this box?" -> move a little, grow a little

Every grid point gets the same set of shapes. A picture with a few hundred grid points and a handful of shapes produces tens of thousands of candidate rectangles.

Almost all of them contain nothing. That imbalance is the source of most of the difficulty in training a detector.

The trade-off nobody tells you

Anchors are guesses baked in before you saw your data.

If your objects are tall and thin, and your anchors are all squarish, the nearest stock size never fits. The tailor has no cloth close enough to alter. Detection quality collapses, and no amount of training rescues it.

This is why anchor sizes and shapes are settings you tune for your dataset. It is also why newer detectors dropped anchors altogether.

Where you have seen the result

  • A face detector drawing tight boxes on every face in a group photo.
  • A parking system that finds each car in a wide camera view.
  • A shop shelf tool locating every product on a rack.
  • A road camera boxing helmets and number plates.

Remember this

  • An anchor box is a pre-made rectangle the model corrects rather than invents.
  • Every grid position carries the same set of shapes, so a picture gets tens of thousands of candidates.
  • Anchors that do not match your objects' shapes cap your accuracy before training starts.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install torch torchvision

Written against torch 2.5.1 and torchvision 0.20.1.

Generating anchors, matching them, and encoding the targets

anchors.py
import torch
from torchvision.models.detection.anchor_utils import AnchorGenerator
from torchvision.models.detection.image_list import ImageList
from torchvision.ops import box_iou

# One feature map at stride 32: a 640x640 image becomes a 20x20 grid of cells.
gen = AnchorGenerator(sizes=((32, 64, 128),), aspect_ratios=((0.5, 1.0, 2.0),))
images = ImageList(torch.zeros(1, 3, 640, 640), [(640, 640)])
feature_maps = [torch.zeros(1, 256, 20, 20)]
anchors = gen(images, feature_maps)[0]

print("anchors from one 20x20 map:", anchors.shape[0], "=", 20 * 20, "cells x 9 shapes")
print("\nthe 9 shapes on the first grid position (torchvision centres it at pixel 0,0):")
for a in anchors[:9]:
    x1, y1, x2, y2 = a.tolist()
    print(f"  {x1:8.1f}{y1:8.1f}{x2:8.1f}{y2:8.1f}   w={x2-x1:6.1f} h={y2-y1:6.1f}"
          f" ratio={(x2-x1)/(y2-y1):.2f}")

# Match anchors to one ground-truth box, the way every anchor detector does.
gt = torch.tensor([[300., 280., 420., 500.]])           # a standing person
ious = box_iou(anchors, gt).squeeze(1)

positive = ious >= 0.5
negative = ious < 0.4
ignored = ~positive & ~negative
print(f"\nagainst one ground-truth box, out of {len(anchors)} anchors:")
print(f"  positives (IoU 0.5 and above): {positive.sum().item()}")
print(f"  ignored   (0.4 to 0.5)       : {ignored.sum().item()}")
print(f"  negatives (below 0.4)        : {negative.sum().item()}")
print(f"  best anchor IoU              : {ious.max().item():.3f}")
print(f"  background to object ratio   : {negative.sum().item() // max(positive.sum().item(), 1)} to 1")

best = anchors[ious.argmax()]
print("  best anchor                  :", [round(v, 1) for v in best.tolist()])

# Training does not predict the box. It predicts a correction to the anchor.
def encode(anchor, box):
    aw, ah = anchor[2] - anchor[0], anchor[3] - anchor[1]
    ax, ay = anchor[0] + aw / 2, anchor[1] + ah / 2
    bw, bh = box[2] - box[0], box[3] - box[1]
    bx, by = box[0] + bw / 2, box[1] + bh / 2
    return torch.stack([(bx - ax) / aw, (by - ay) / ah, torch.log(bw / aw), torch.log(bh / ah)])

def decode(anchor, t):
    aw, ah = anchor[2] - anchor[0], anchor[3] - anchor[1]
    ax, ay = anchor[0] + aw / 2, anchor[1] + ah / 2
    bx, by = t[0] * aw + ax, t[1] * ah + ay
    bw, bh = torch.exp(t[2]) * aw, torch.exp(t[3]) * ah
    return torch.stack([bx - bw / 2, by - bh / 2, bx + bw / 2, by + bh / 2])

t = encode(best, gt[0])
print("\nthe four numbers the network is actually trained to output:")
print("  dx, dy, dw, dh =", [round(v, 3) for v in t.tolist()])
print("  decoded back   :", [round(v, 1) for v in decode(best, t).tolist()])

# An object whose shape no anchor covers is close to undetectable.
thin = torch.tensor([[100., 300., 520., 316.]])          # a cable: 420 wide, 16 tall
print("\na very thin object, 420 wide and 16 tall:")
print(f"  best IoU against any of the {len(anchors)} anchors: {box_iou(anchors, thin).max().item():.3f}")
Output
anchors from one 20x20 map: 3600 = 400 cells x 9 shapes

the 9 shapes on the first grid position (torchvision centres it at pixel 0,0):
     -23.0   -11.0    23.0    11.0   w=  46.0 h=  22.0 ratio=2.09
     -45.0   -23.0    45.0    23.0   w=  90.0 h=  46.0 ratio=1.96
     -91.0   -45.0    91.0    45.0   w= 182.0 h=  90.0 ratio=2.02
     -16.0   -16.0    16.0    16.0   w=  32.0 h=  32.0 ratio=1.00
     -32.0   -32.0    32.0    32.0   w=  64.0 h=  64.0 ratio=1.00
     -64.0   -64.0    64.0    64.0   w= 128.0 h= 128.0 ratio=1.00
     -11.0   -23.0    11.0    23.0   w=  22.0 h=  46.0 ratio=0.48
     -23.0   -45.0    23.0    45.0   w=  46.0 h=  90.0 ratio=0.51
     -45.0   -91.0    45.0    91.0   w=  90.0 h= 182.0 ratio=0.49

against one ground-truth box, out of 3600 anchors:
  positives (IoU 0.5 and above): 7
  ignored   (0.4 to 0.5)       : 7
  negatives (below 0.4)        : 3586
  best anchor IoU              : 0.620
  background to object ratio   : 512 to 1
  best anchor                  : [307.0, 293.0, 397.0, 475.0]

the four numbers the network is actually trained to output:
  dx, dy, dw, dh = [0.089, 0.033, 0.288, 0.19]
  decoded back   : [300.0, 280.0, 420.0, 500.0]

a very thin object, 420 wide and 16 tall:
  best IoU against any of the 3600 anchors: 0.153

Reading that output

The first nine anchors are centred at pixel (0, 0), not at the middle of the first cell. torchvision builds its shift grid as arange(width) * stride, so anchors hang off the top-left corner. Different frameworks make different choices here. A half-cell offset is a real source of small, stubborn accuracy loss. It bites when you port an anchor configuration between codebases.

Widths and heights are 46 by 22, not 45.25 by 22.6. AnchorGenerator rounds the scaled sides. The ratios come out as 2.09 and 0.48 rather than exactly 2.0 and 0.5. This is harmless and worth knowing before you write a test asserting exact ratios.

One object produced 7 positives and 3,586 negatives. That is a 512-to-1 imbalance from a single object on a single feature level. Train with plain cross-entropy on that and the model learns to answer "background" to everything. Two lessons later this becomes the entire motivation for focal loss.

The best anchor only reached IoU 0.620. Nothing in this grid lines up perfectly with a 120 by 220 person. That is normal, and it is what the regression step is for.

The four regression targets are small numbers near zero. dx and dy are fractions of the anchor's own width and height. dw and dh are logarithms of a size ratio. Both choices keep the targets scale-free. A 20-pixel error on a truck and on a bird are then not treated the same. Decoding returns the exact ground-truth box, which is the correctness check for any encoder you write.

The thin cable topped out at IoU 0.153. Every assignment rule in common use would label that anchor background. The object is invisible to this detector, and no training run fixes it. Change the anchor ratios, or use an anchor-free head.

How many anchors a real detector has

anchor_count.py
import torch
from torchvision.models.detection.anchor_utils import AnchorGenerator
from torchvision.models.detection.image_list import ImageList

# RetinaNet's real anchor set: 5 pyramid levels, 3 scales, 3 aspect ratios.
sizes = tuple((s, int(s * 2 ** (1 / 3)), int(s * 2 ** (2 / 3))) for s in (32, 64, 128, 256, 512))
ratios = ((0.5, 1.0, 2.0),) * 5
gen = AnchorGenerator(sizes=sizes, aspect_ratios=ratios)

H = W = 640
levels = [("P3", 8), ("P4", 16), ("P5", 32), ("P6", 64), ("P7", 128)]
feats = [torch.zeros(1, 256, H // s, W // s) for _, s in levels]
anchors = gen(ImageList(torch.zeros(1, 3, H, W), [(H, W)]), feats)[0]

print("scales on the first level:", sizes[0], "- a third of an octave apart")
total = 0
for (name, stride), f in zip(levels, feats):
    n = f.shape[-1] * f.shape[-2] * 9
    total += n
    print(f"  {name} stride {stride:>3}: {f.shape[-1]:>2}x{f.shape[-2]:<2} cells x 9 shapes = {n:>7,}")
print(f"  total for one 640x640 image        = {total:>7,}")
print("anchors actually generated          =", f"{anchors.shape[0]:,}")
Output
scales on the first level: (32, 40, 50) - a third of an octave apart
  P3 stride   8: 80x80 cells x 9 shapes =  57,600
  P4 stride  16: 40x40 cells x 9 shapes =  14,400
  P5 stride  32: 20x20 cells x 9 shapes =   3,600
  P6 stride  64: 10x10 cells x 9 shapes =     900
  P7 stride 128:  5x5  cells x 9 shapes =     225
  total for one 640x640 image        =  76,725
anchors actually generated          = 76,725

Three quarters of the anchors live on the highest-resolution level. That is where small objects are found. It is also why that level dominates memory use and the class imbalance.

Choosing anchors for your own dataset

Do not inherit COCO's anchors without checking. Run this on your labels:

  1. Collect every ground-truth width and height in your training set, in the resolution your model will see.
  2. Cluster them. K-means on width and height with an IoU-based distance is what YOLOv2 introduced, and it still works.
  3. Check coverage: for every ground-truth box, compute the best IoU against your candidate anchor set. Report the fraction below 0.5.

Suppose more than a few percent of your objects cannot reach IoU 0.5 with any anchor. Your ceiling is then set before training begins. Widen the ratio range, add a scale, or move to an anchor-free head.

Common mistakes

Copying anchor sizes but not the input resolution. Anchors are in pixels of the resized input. Change from 640 to 1280 and every anchor is effectively half the size relative to your objects.

Assigning by IoU only, with no fallback. Some ground-truth boxes match no anchor above the threshold. Faster R-CNN and RetinaNet both force the single best anchor per ground-truth box to be positive, whatever its IoU. Without that rule, small objects get no positive sample at all.

Forgetting the ignore band. Anchors between the low and high thresholds are excluded from the loss, not labelled background. Treating them as background injects noisy negatives right where the model is least certain.

Predicting sizes directly instead of in log-space. A raw width prediction can go negative. exp cannot.

Testing with anchors that never saw a match. If a level produces zero positives across your whole validation set, that level is dead weight. Print per-level positive counts once.

Try it yourself

Change aspect_ratios to ((0.1, 0.5, 1.0, 2.0, 10.0),) and re-run the thin-cable check. Watch the best IoU jump. Then count how many anchors you now have, and decide whether that trade is worth it for your data.

What to learn next

Researcher — Mathematics and papers.

Assignment as the real design decision

Anchors are cheap. The rule mapping anchors to targets is what determines performance.

Let $\mathcal{A}$ be the anchor set and $\mathcal{G}$ the ground-truth set. The IoU assignment used by Faster R-CNN and RetinaNet is:

$$ \text{label}(a) = \begin{cases} \text{positive}, & \max_{g \in \mathcal{G}} \mathrm{IoU}(a, g) \geq \tau_{+} \ \text{ignore}, & \tau_{-} \leq \max_g \mathrm{IoU}(a, g) < \tau_{+} \ \text{negative}, & \max_g \mathrm{IoU}(a, g) < \tau_{-} \end{cases} $$

With $(\tau_{-}, \tau_{+}) = (0.3, 0.7)$ in the RPN and $(0.4, 0.5)$ in RetinaNet. A second rule forces $\arg\max_{a} \mathrm{IoU}(a, g)$ positive for every $g$, guaranteeing at least one positive per object.

The ignore band exists because anchors near the threshold carry ambiguous supervision. Labelling them either way injects gradient noise.

Regression parameterisation

The R-CNN encoding, with anchor $(x_a, y_a, w_a, h_a)$ and target $(x, y, w, h)$ in centre-size form:

$$ t_x = \frac{x - x_a}{w_a}, \quad t_y = \frac{y - y_a}{h_a}, \quad t_w = \log\frac{w}{w_a}, \quad t_h = \log\frac{h}{h_a} $$

Detectors additionally divide by fixed standard deviations, conventionally $(0.1, 0.1, 0.2, 0.2)$. The four targets then have comparable variance, so one smooth-$L_1$ loss weights them sensibly. torchvision exposes these as bbox_reg_weights on BoxCoder.

The imbalance this creates

With $|\mathcal{A}| \approx 10^5$ and a handful of objects, the positive fraction is around $10^{-3}$. Three families of fix exist, and they are not mutually exclusive.

  • Sampling. The RPN samples 256 anchors per image at a 1:1 target ratio. Fast, and it discards most of the data.
  • Hard example mining. OHEM (Shrivastava et al., 2016) and SSD's hard-negative mining keep the highest-loss negatives at a fixed ratio. Effective, and it adds a non-differentiable selection step.
  • Loss reweighting. Focal loss (Lin et al., 2017) keeps every anchor and down-weights easy ones. See focal loss and RetinaNet.

Learned and adaptive assignment

Fixed IoU thresholds are a hand-set prior. Several lines of work replace them.

  • ATSS (Zhang et al., 2020) sets a per-object threshold. It takes the mean $\mu$ and standard deviation $\sigma$ of IoU over that object's $k$ nearest candidates, then uses $\mu + \sigma$. It shows the anchor-based versus anchor-free gap is largely an assignment difference.
  • FreeAnchor (Zhang et al., 2019) treats assignment as maximum-likelihood estimation over a bag of candidates.
  • OTA (Ge et al., 2021) casts assignment as optimal transport over all objects at once; SimOTA in YOLOX is its cheaper top-$k$ approximation.
  • Task-aligned assignment (Feng et al., 2021, TOOD) scores candidates by $s^{\alpha} \cdot \mathrm{IoU}^{\beta}$, combining classification confidence with localisation quality. This is what current Ultralytics YOLO models use.

The trajectory is consistent: from a fixed geometric rule, to a per-object adaptive rule, to a global assignment problem solved jointly. DETR's Hungarian matching is the end of that road; see DETR and set prediction.

Anchor shape selection

Redmon and Farhadi (2017), YOLOv2, run k-means over ground-truth dimensions with the distance

$$ d(\text{box}, \text{centroid}) = 1 - \mathrm{IoU}(\text{box}, \text{centroid}) $$

rather than Euclidean distance, because Euclidean distance over-penalises large boxes. Five clusters gave better average IoU than nine hand-picked anchors on their data.

Report average best-IoU coverage when you tune anchors: the mean over ground-truth boxes of the best IoU achievable with the candidate set. It upper-bounds recall at that IoU threshold, and it is measurable before any training.

Papers

What to learn next