Anchor boxes
Anchor boxes are a fixed grid of guessed rectangles that a detector nudges into place, which turns finding objects into correcting guesses.
- 15 min read
- 3 reading levels
- Updated
Read these first
On this page 8
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
Anchor boxes are ready-made rectangles scattered across the picture, which the model then nudges onto real objects.
The analogy
Think about a tailor fitting a shirt. He does not cut cloth from nothing while you stand there. He picks a stock size off the rail, holds it against you, and takes it in at the shoulders.
Starting from a stock size is far easier than starting from bare cloth. The tailor only has to answer a small question: how much to let out, how much to take in.
An anchor box is that stock size. The model picks the nearest one and adjusts it.
Why this had to be invented
The first detectors ran a separate search to propose likely object regions, then classified each one. The search was slow and lived outside the network, so it could not learn.
Anchors removed the search. Cover the picture with a fixed grid of rectangles of several sizes and shapes. Ask the network two things about every rectangle: is there an object here, and how should this rectangle move to fit it?
Both questions have short answers. Both can be learned. Nothing outside the network has to run.
How it works
image grid shapes placed at every grid point
. . . . . . . . +--------+ +----+ +--+
. . . . . . . . | wide | |sqr | |ta|
. . . . . . . . +--------+ +----+ |ll|
. . . . . . . . +--+
for every point, for every shape:
"is an object here?" -> yes / no
"how do I fix this box?" -> move a little, grow a littleEvery grid point gets the same set of shapes. A picture with a few hundred grid points and a handful of shapes produces tens of thousands of candidate rectangles.
Almost all of them contain nothing. That imbalance is the source of most of the difficulty in training a detector.
The trade-off nobody tells you
Anchors are guesses baked in before you saw your data.
If your objects are tall and thin, and your anchors are all squarish, the nearest stock size never fits. The tailor has no cloth close enough to alter. Detection quality collapses, and no amount of training rescues it.
This is why anchor sizes and shapes are settings you tune for your dataset. It is also why newer detectors dropped anchors altogether.
Where you have seen the result
- A face detector drawing tight boxes on every face in a group photo.
- A parking system that finds each car in a wide camera view.
- A shop shelf tool locating every product on a rack.
- A road camera boxing helmets and number plates.
Remember this
- An anchor box is a pre-made rectangle the model corrects rather than invents.
- Every grid position carries the same set of shapes, so a picture gets tens of thousands of candidates.
- Anchors that do not match your objects' shapes cap your accuracy before training starts.
What to learn next
- Non-maximum suppression — cleaning up the duplicates all these anchors produce.
- Focal loss and RetinaNet — the loss built to survive a 500-to-1 imbalance.
- Anchor-free detection — what replaced anchors, and why.
Developer — Code and libraries.
Setup
pip install torch torchvisionWritten against torch 2.5.1 and torchvision 0.20.1.
Generating anchors, matching them, and encoding the targets
import torch
from torchvision.models.detection.anchor_utils import AnchorGenerator
from torchvision.models.detection.image_list import ImageList
from torchvision.ops import box_iou
# One feature map at stride 32: a 640x640 image becomes a 20x20 grid of cells.
gen = AnchorGenerator(sizes=((32, 64, 128),), aspect_ratios=((0.5, 1.0, 2.0),))
images = ImageList(torch.zeros(1, 3, 640, 640), [(640, 640)])
feature_maps = [torch.zeros(1, 256, 20, 20)]
anchors = gen(images, feature_maps)[0]
print("anchors from one 20x20 map:", anchors.shape[0], "=", 20 * 20, "cells x 9 shapes")
print("\nthe 9 shapes on the first grid position (torchvision centres it at pixel 0,0):")
for a in anchors[:9]:
x1, y1, x2, y2 = a.tolist()
print(f" {x1:8.1f}{y1:8.1f}{x2:8.1f}{y2:8.1f} w={x2-x1:6.1f} h={y2-y1:6.1f}"
f" ratio={(x2-x1)/(y2-y1):.2f}")
# Match anchors to one ground-truth box, the way every anchor detector does.
gt = torch.tensor([[300., 280., 420., 500.]]) # a standing person
ious = box_iou(anchors, gt).squeeze(1)
positive = ious >= 0.5
negative = ious < 0.4
ignored = ~positive & ~negative
print(f"\nagainst one ground-truth box, out of {len(anchors)} anchors:")
print(f" positives (IoU 0.5 and above): {positive.sum().item()}")
print(f" ignored (0.4 to 0.5) : {ignored.sum().item()}")
print(f" negatives (below 0.4) : {negative.sum().item()}")
print(f" best anchor IoU : {ious.max().item():.3f}")
print(f" background to object ratio : {negative.sum().item() // max(positive.sum().item(), 1)} to 1")
best = anchors[ious.argmax()]
print(" best anchor :", [round(v, 1) for v in best.tolist()])
# Training does not predict the box. It predicts a correction to the anchor.
def encode(anchor, box):
aw, ah = anchor[2] - anchor[0], anchor[3] - anchor[1]
ax, ay = anchor[0] + aw / 2, anchor[1] + ah / 2
bw, bh = box[2] - box[0], box[3] - box[1]
bx, by = box[0] + bw / 2, box[1] + bh / 2
return torch.stack([(bx - ax) / aw, (by - ay) / ah, torch.log(bw / aw), torch.log(bh / ah)])
def decode(anchor, t):
aw, ah = anchor[2] - anchor[0], anchor[3] - anchor[1]
ax, ay = anchor[0] + aw / 2, anchor[1] + ah / 2
bx, by = t[0] * aw + ax, t[1] * ah + ay
bw, bh = torch.exp(t[2]) * aw, torch.exp(t[3]) * ah
return torch.stack([bx - bw / 2, by - bh / 2, bx + bw / 2, by + bh / 2])
t = encode(best, gt[0])
print("\nthe four numbers the network is actually trained to output:")
print(" dx, dy, dw, dh =", [round(v, 3) for v in t.tolist()])
print(" decoded back :", [round(v, 1) for v in decode(best, t).tolist()])
# An object whose shape no anchor covers is close to undetectable.
thin = torch.tensor([[100., 300., 520., 316.]]) # a cable: 420 wide, 16 tall
print("\na very thin object, 420 wide and 16 tall:")
print(f" best IoU against any of the {len(anchors)} anchors: {box_iou(anchors, thin).max().item():.3f}")anchors from one 20x20 map: 3600 = 400 cells x 9 shapes
the 9 shapes on the first grid position (torchvision centres it at pixel 0,0):
-23.0 -11.0 23.0 11.0 w= 46.0 h= 22.0 ratio=2.09
-45.0 -23.0 45.0 23.0 w= 90.0 h= 46.0 ratio=1.96
-91.0 -45.0 91.0 45.0 w= 182.0 h= 90.0 ratio=2.02
-16.0 -16.0 16.0 16.0 w= 32.0 h= 32.0 ratio=1.00
-32.0 -32.0 32.0 32.0 w= 64.0 h= 64.0 ratio=1.00
-64.0 -64.0 64.0 64.0 w= 128.0 h= 128.0 ratio=1.00
-11.0 -23.0 11.0 23.0 w= 22.0 h= 46.0 ratio=0.48
-23.0 -45.0 23.0 45.0 w= 46.0 h= 90.0 ratio=0.51
-45.0 -91.0 45.0 91.0 w= 90.0 h= 182.0 ratio=0.49
against one ground-truth box, out of 3600 anchors:
positives (IoU 0.5 and above): 7
ignored (0.4 to 0.5) : 7
negatives (below 0.4) : 3586
best anchor IoU : 0.620
background to object ratio : 512 to 1
best anchor : [307.0, 293.0, 397.0, 475.0]
the four numbers the network is actually trained to output:
dx, dy, dw, dh = [0.089, 0.033, 0.288, 0.19]
decoded back : [300.0, 280.0, 420.0, 500.0]
a very thin object, 420 wide and 16 tall:
best IoU against any of the 3600 anchors: 0.153Reading that output
The first nine anchors are centred at pixel (0, 0), not at the middle of the first cell. torchvision builds its shift grid as arange(width) * stride, so anchors hang off the top-left corner. Different frameworks make different choices here. A half-cell offset is a real source of small, stubborn accuracy loss. It bites when you port an anchor configuration between codebases.
Widths and heights are 46 by 22, not 45.25 by 22.6. AnchorGenerator rounds the scaled sides. The ratios come out as 2.09 and 0.48 rather than exactly 2.0 and 0.5. This is harmless and worth knowing before you write a test asserting exact ratios.
One object produced 7 positives and 3,586 negatives. That is a 512-to-1 imbalance from a single object on a single feature level. Train with plain cross-entropy on that and the model learns to answer "background" to everything. Two lessons later this becomes the entire motivation for focal loss.
The best anchor only reached IoU 0.620. Nothing in this grid lines up perfectly with a 120 by 220 person. That is normal, and it is what the regression step is for.
The four regression targets are small numbers near zero. dx and dy are fractions of the anchor's own width and height. dw and dh are logarithms of a size ratio. Both choices keep the targets scale-free. A 20-pixel error on a truck and on a bird are then not treated the same. Decoding returns the exact ground-truth box, which is the correctness check for any encoder you write.
The thin cable topped out at IoU 0.153. Every assignment rule in common use would label that anchor background. The object is invisible to this detector, and no training run fixes it. Change the anchor ratios, or use an anchor-free head.
How many anchors a real detector has
import torch
from torchvision.models.detection.anchor_utils import AnchorGenerator
from torchvision.models.detection.image_list import ImageList
# RetinaNet's real anchor set: 5 pyramid levels, 3 scales, 3 aspect ratios.
sizes = tuple((s, int(s * 2 ** (1 / 3)), int(s * 2 ** (2 / 3))) for s in (32, 64, 128, 256, 512))
ratios = ((0.5, 1.0, 2.0),) * 5
gen = AnchorGenerator(sizes=sizes, aspect_ratios=ratios)
H = W = 640
levels = [("P3", 8), ("P4", 16), ("P5", 32), ("P6", 64), ("P7", 128)]
feats = [torch.zeros(1, 256, H // s, W // s) for _, s in levels]
anchors = gen(ImageList(torch.zeros(1, 3, H, W), [(H, W)]), feats)[0]
print("scales on the first level:", sizes[0], "- a third of an octave apart")
total = 0
for (name, stride), f in zip(levels, feats):
n = f.shape[-1] * f.shape[-2] * 9
total += n
print(f" {name} stride {stride:>3}: {f.shape[-1]:>2}x{f.shape[-2]:<2} cells x 9 shapes = {n:>7,}")
print(f" total for one 640x640 image = {total:>7,}")
print("anchors actually generated =", f"{anchors.shape[0]:,}")scales on the first level: (32, 40, 50) - a third of an octave apart P3 stride 8: 80x80 cells x 9 shapes = 57,600 P4 stride 16: 40x40 cells x 9 shapes = 14,400 P5 stride 32: 20x20 cells x 9 shapes = 3,600 P6 stride 64: 10x10 cells x 9 shapes = 900 P7 stride 128: 5x5 cells x 9 shapes = 225 total for one 640x640 image = 76,725 anchors actually generated = 76,725
Three quarters of the anchors live on the highest-resolution level. That is where small objects are found. It is also why that level dominates memory use and the class imbalance.
Choosing anchors for your own dataset
Do not inherit COCO's anchors without checking. Run this on your labels:
- Collect every ground-truth width and height in your training set, in the resolution your model will see.
- Cluster them. K-means on width and height with an IoU-based distance is what YOLOv2 introduced, and it still works.
- Check coverage: for every ground-truth box, compute the best IoU against your candidate anchor set. Report the fraction below 0.5.
Suppose more than a few percent of your objects cannot reach IoU 0.5 with any anchor. Your ceiling is then set before training begins. Widen the ratio range, add a scale, or move to an anchor-free head.
Common mistakes
Copying anchor sizes but not the input resolution. Anchors are in pixels of the resized input. Change from 640 to 1280 and every anchor is effectively half the size relative to your objects.
Assigning by IoU only, with no fallback. Some ground-truth boxes match no anchor above the threshold. Faster R-CNN and RetinaNet both force the single best anchor per ground-truth box to be positive, whatever its IoU. Without that rule, small objects get no positive sample at all.
Forgetting the ignore band. Anchors between the low and high thresholds are excluded from the loss, not labelled background. Treating them as background injects noisy negatives right where the model is least certain.
Predicting sizes directly instead of in log-space. A raw width prediction can go negative. exp cannot.
Testing with anchors that never saw a match. If a level produces zero positives across your whole validation set, that level is dead weight. Print per-level positive counts once.
Try it yourself
Change aspect_ratios to ((0.1, 0.5, 1.0, 2.0, 10.0),) and re-run the thin-cable check. Watch the best IoU jump. Then count how many anchors you now have, and decide whether that trade is worth it for your data.
What to learn next
- Non-maximum suppression — cleaning up the duplicates all these anchors produce.
- Focal loss and RetinaNet — the loss built to survive a 500-to-1 imbalance.
- Anchor-free detection — what replaced anchors, and why.
Researcher — Mathematics and papers.
Assignment as the real design decision
Anchors are cheap. The rule mapping anchors to targets is what determines performance.
Let $\mathcal{A}$ be the anchor set and $\mathcal{G}$ the ground-truth set. The IoU assignment used by Faster R-CNN and RetinaNet is:
$$ \text{label}(a) = \begin{cases} \text{positive}, & \max_{g \in \mathcal{G}} \mathrm{IoU}(a, g) \geq \tau_{+} \ \text{ignore}, & \tau_{-} \leq \max_g \mathrm{IoU}(a, g) < \tau_{+} \ \text{negative}, & \max_g \mathrm{IoU}(a, g) < \tau_{-} \end{cases} $$
With $(\tau_{-}, \tau_{+}) = (0.3, 0.7)$ in the RPN and $(0.4, 0.5)$ in RetinaNet. A second rule forces $\arg\max_{a} \mathrm{IoU}(a, g)$ positive for every $g$, guaranteeing at least one positive per object.
The ignore band exists because anchors near the threshold carry ambiguous supervision. Labelling them either way injects gradient noise.
Regression parameterisation
The R-CNN encoding, with anchor $(x_a, y_a, w_a, h_a)$ and target $(x, y, w, h)$ in centre-size form:
$$ t_x = \frac{x - x_a}{w_a}, \quad t_y = \frac{y - y_a}{h_a}, \quad t_w = \log\frac{w}{w_a}, \quad t_h = \log\frac{h}{h_a} $$
Detectors additionally divide by fixed standard deviations, conventionally $(0.1, 0.1, 0.2, 0.2)$. The four targets then have comparable variance, so one smooth-$L_1$ loss weights them sensibly. torchvision exposes these as bbox_reg_weights on BoxCoder.
The imbalance this creates
With $|\mathcal{A}| \approx 10^5$ and a handful of objects, the positive fraction is around $10^{-3}$. Three families of fix exist, and they are not mutually exclusive.
- Sampling. The RPN samples 256 anchors per image at a 1:1 target ratio. Fast, and it discards most of the data.
- Hard example mining. OHEM (Shrivastava et al., 2016) and SSD's hard-negative mining keep the highest-loss negatives at a fixed ratio. Effective, and it adds a non-differentiable selection step.
- Loss reweighting. Focal loss (Lin et al., 2017) keeps every anchor and down-weights easy ones. See focal loss and RetinaNet.
Learned and adaptive assignment
Fixed IoU thresholds are a hand-set prior. Several lines of work replace them.
- ATSS (Zhang et al., 2020) sets a per-object threshold. It takes the mean $\mu$ and standard deviation $\sigma$ of IoU over that object's $k$ nearest candidates, then uses $\mu + \sigma$. It shows the anchor-based versus anchor-free gap is largely an assignment difference.
- FreeAnchor (Zhang et al., 2019) treats assignment as maximum-likelihood estimation over a bag of candidates.
- OTA (Ge et al., 2021) casts assignment as optimal transport over all objects at once; SimOTA in YOLOX is its cheaper top-$k$ approximation.
- Task-aligned assignment (Feng et al., 2021, TOOD) scores candidates by $s^{\alpha} \cdot \mathrm{IoU}^{\beta}$, combining classification confidence with localisation quality. This is what current Ultralytics YOLO models use.
The trajectory is consistent: from a fixed geometric rule, to a per-object adaptive rule, to a global assignment problem solved jointly. DETR's Hungarian matching is the end of that road; see DETR and set prediction.
Anchor shape selection
Redmon and Farhadi (2017), YOLOv2, run k-means over ground-truth dimensions with the distance
$$ d(\text{box}, \text{centroid}) = 1 - \mathrm{IoU}(\text{box}, \text{centroid}) $$
rather than Euclidean distance, because Euclidean distance over-penalises large boxes. Five clusters gave better average IoU than nine hand-picked anchors on their data.
Report average best-IoU coverage when you tune anchors: the mean over ground-truth boxes of the best IoU achievable with the candidate set. It upper-bounds recall at that IoU threshold, and it is measurable before any training.
Papers
- Ren et al., Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks, NeurIPS 2015 — arxiv.org/abs/1506.01497
- Liu et al., SSD: Single Shot MultiBox Detector, ECCV 2016 — arxiv.org/abs/1512.02325
- Redmon and Farhadi, YOLO9000: Better, Faster, Stronger, CVPR 2017 — arxiv.org/abs/1612.08242
- Zhang et al., Bridging the Gap Between Anchor-based and Anchor-free Detection via ATSS, CVPR 2020 — arxiv.org/abs/1912.02424
- Ge et al., OTA: Optimal Transport Assignment for Object Detection, CVPR 2021 — arxiv.org/abs/2103.14259
What to learn next
- Non-maximum suppression — cleaning up the duplicates all these anchors produce.
- Focal loss and RetinaNet — the loss built to survive a 500-to-1 imbalance.
- Anchor-free detection — what replaced anchors, and why.