Anchor-free detection
Anchor-free detectors let every grid position say how far the object edges are, which deletes a whole layer of hand-tuned box shapes and most of the candidates that came with them.
- 16 min read
- 3 reading levels
- Updated
Read these first
On this page 9
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
An anchor-free detector asks each grid position how far away the object's four edges are. Nothing is corrected from a ready-made rectangle.
The analogy
Think about standing inside a room in the dark and being asked to describe it. You could be handed a set of standard room sizes and asked which one fits best. That means someone had to guess your room sizes in advance.
Or you could stretch out an arm in each direction and report four distances. The wall in front, behind, left and right. Four numbers, no catalogue, and it works in a room nobody anticipated.
Anchor-free detection is the second approach. Stand at a point, feel outwards, report the distances.
Why it exists
Ready-made rectangles came with a hidden cost.
Someone has to choose their sizes and shapes. Choose badly and objects your model will meet in production never match anything well enough to be found. This is a tuning job that has to be redone for every new dataset.
They also multiply everything. Nine shapes at every grid position means nine times the candidates and nine times the memory. It also worsens the imbalance between object and background.
And they add settings. Overlap thresholds for positive, negative and ignore, per detector, per level.
Removing them removes all three problems at once.
How it works
an object in the picture
+-------------------------------+
| |
| ^ up |
| | |
| left o------> right | o = one grid position inside the object
| | |
| v down |
+-------------------------------+
the model reports four distances from o to the four edgesEvery grid position inside an object is trained to report its four distances. Positions outside all objects are trained to report background.
At prediction time, take a position, read its four distances, and you have a rectangle. No catalogue, no matching rule, no thresholds.
The problem this created, and its fix
Positions near the edge of an object give poor boxes. They can see very little of it, yet they may report high confidence.
The fix is a second small output called centre-ness: how central this position is inside the object it belongs to. It runs from zero at the boundary to one at the exact centre. Multiply it into the confidence, and edge positions become quiet without being deleted.
This one extra number is worth a surprising amount of accuracy. It is a good example of a small honest signal beating a clever architecture.
What about two objects on top of each other?
A position inside two overlapping objects has to pick one. Two rules handle almost all of it.
Different sizes go to different levels of the network, so a small object and a large one rarely compete. When they still do, the smaller object wins, because the larger one has plenty of other positions to spare.
Where you have seen the result
- Detection running smoothly on a phone with no per-dataset tuning.
- Text detectors finding lines of very different shapes.
- Aerial tools locating buildings of wildly varying size.
- Modern YOLO releases, which are all built this way.
Remember this
- Every grid position reports four distances to the object's edges, so no ready-made shapes are needed.
- That removes anchor tuning, cuts candidate boxes, and deletes several thresholds.
- Centre-ness quietens positions near an object's edge, which recovers most of the lost accuracy.
What to learn next
- DETR and set prediction — removing the assignment heuristics as well.
- YOLO versions compared — where this head design ended up in production.
- Anchor boxes — what this replaced, and why it needed replacing.
Developer — Code and libraries.
Setup
pip install torch torchvisionWritten against torch 2.5.1 and torchvision 0.20.1. The first script is plain PyTorch tensor arithmetic on an 8x8 grid, so you can read every number.
Building an FCOS target by hand
import torch
# A 128x128 image, one feature level at stride 16, so an 8x8 grid of cells.
STRIDE, G = 16, 8
gt = torch.tensor([20., 30., 100., 110.]) # one object, xyxy
ys, xs = torch.meshgrid(torch.arange(G), torch.arange(G), indexing="ij")
cx = (xs.float() + 0.5) * STRIDE # each cell's centre, in pixels
cy = (ys.float() + 0.5) * STRIDE
# The four numbers an anchor-free head predicts: distance to each edge.
l, t = cx - gt[0], cy - gt[1]
r, b = gt[2] - cx, gt[3] - cy
inside = (l > 0) & (t > 0) & (r > 0) & (b > 0) # a cell is positive if it is inside
print("cells inside the box (1 = trained as a positive):")
print(inside.int())
print("positives:", inside.sum().item(), "out of", G * G)
# Centre-ness: how central is this cell inside the object?
lr = (torch.minimum(l, r) / torch.maximum(l, r).clamp(min=1e-6)).clamp(min=0)
tb = (torch.minimum(t, b) / torch.maximum(t, b).clamp(min=1e-6)).clamp(min=0)
centerness = torch.sqrt(lr * tb) * inside
print("\ncentre-ness per cell (0 outside the box):")
for row in centerness:
print(" " + " ".join(f"{v:5.2f}" for v in row))
# Decode: four distances become a box, with no anchor anywhere.
i, j = divmod(centerness.argmax().item(), G)
pred = torch.tensor([cx[i, j] - l[i, j], cy[i, j] - t[i, j],
cx[i, j] + r[i, j], cy[i, j] + b[i, j]])
print(f"\nmost central cell: row {i}, col {j}, centre ({cx[i,j]:.0f}, {cy[i,j]:.0f})")
print("its l, t, r, b :", [round(v.item(), 1) for v in (l[i,j], t[i,j], r[i,j], b[i,j])])
print("decoded box :", [round(v, 1) for v in pred.tolist()])
print("ground truth :", [round(v, 1) for v in gt.tolist()])
# Centre sampling: FCOS trains only cells near the object centre, not the whole box.
gcx, gcy = (gt[0] + gt[2]) / 2, (gt[1] + gt[3]) / 2
radius = 1.5 * STRIDE
central = inside & ((cx - gcx).abs() < radius) & ((cy - gcy).abs() < radius)
print(f"\nwith centre sampling (radius {radius:.0f}px): "
f"{central.sum().item()} positives instead of {inside.sum().item()}")
# Ambiguity: two overlapping objects fighting over the same cell.
gt2 = torch.tensor([60., 60., 160., 160.])
in2 = (cx > gt2[0]) & (cy > gt2[1]) & (cx < gt2[2]) & (cy < gt2[3])
print(f"\ncells claimed by both objects: {(inside & in2).sum().item()}")
a1 = (gt[2] - gt[0]) * (gt[3] - gt[1])
a2 = (gt2[2] - gt2[0]) * (gt2[3] - gt2[1])
print(f"FCOS breaks the tie by area: object 1 = {a1:.0f}, object 2 = {a2:.0f}")
print("the smaller object wins the shared cells")
# The other tie-breaker: send different sizes to different pyramid levels.
print("\nFCOS size ranges per level (largest edge distance the level may regress):")
for name, stride, lo, hi in [("P3", 8, 0, 64), ("P4", 16, 64, 128), ("P5", 32, 128, 256),
("P6", 64, 256, 512), ("P7", 128, 512, None)]:
top = "no limit" if hi is None else str(hi)
print(f" {name} stride {stride:>3}: {lo} to {top}")
biggest = max(l[i, j], t[i, j], r[i, j], b[i, j]).item()
print(f"our object's largest edge distance is {biggest:.0f}px, so it belongs on P3")cells inside the box (1 = trained as a positive):
tensor([[0, 0, 0, 0, 0, 0, 0, 0],
[0, 0, 0, 0, 0, 0, 0, 0],
[0, 1, 1, 1, 1, 1, 0, 0],
[0, 1, 1, 1, 1, 1, 0, 0],
[0, 1, 1, 1, 1, 1, 0, 0],
[0, 1, 1, 1, 1, 1, 0, 0],
[0, 1, 1, 1, 1, 1, 0, 0],
[0, 0, 0, 0, 0, 0, 0, 0]], dtype=torch.int32)
positives: 25 out of 64
centre-ness per cell (0 outside the box):
0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00
0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00
0.00 0.09 0.22 0.34 0.28 0.16 0.00 0.00
0.00 0.16 0.40 0.63 0.51 0.29 0.00 0.00
0.00 0.22 0.55 0.86 0.70 0.40 0.00 0.00
0.00 0.14 0.36 0.56 0.45 0.26 0.00 0.00
0.00 0.07 0.16 0.26 0.21 0.12 0.00 0.00
0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00
most central cell: row 4, col 3, centre (56, 72)
its l, t, r, b : [36.0, 42.0, 44.0, 38.0]
decoded box : [20.0, 30.0, 100.0, 110.0]
ground truth : [20.0, 30.0, 100.0, 110.0]
with centre sampling (radius 24px): 9 positives instead of 25
cells claimed by both objects: 6
FCOS breaks the tie by area: object 1 = 6400, object 2 = 10000
the smaller object wins the shared cells
FCOS size ranges per level (largest edge distance the level may regress):
P3 stride 8: 0 to 64
P4 stride 16: 64 to 128
P5 stride 32: 128 to 256
P6 stride 64: 256 to 512
P7 stride 128: 512 to no limit
our object's largest edge distance is 44px, so it belongs on P3Reading that output
One object produced 25 positive cells out of 64. Compare that with the anchor case, where one object produced 7 positives out of 3,600. Anchor-free assignment is far denser in positives. That is much of why it trains stably with no anchors tuned to the dataset.
The centre-ness map peaks at 0.86, not at 1.00. No cell centre lands exactly on the object centre, because cell centres sit on a stride-16 grid. That is honest: centre-ness measures a real geometric quantity, not a label you assign.
The gradient in that map is the whole idea. 0.86 at the middle, 0.07 at a corner. Multiply the classification score by centre-ness and the corner cells, whose boxes are the least reliable, stop producing confident detections. FCOS reports this as worth several AP points.
The decode is exact. Four distances from a known cell centre reconstruct the ground-truth box to the decimal. No anchor, no log-space encoding, no per-dataset shape catalogue.
Centre sampling cut positives from 25 to 9. The naive "inside the box" rule includes corner cells that see mostly background, especially for large or L-shaped objects. Restricting to a radius around the centre is a small change with a measurable effect. It was added to FCOS after the first version.
Only 6 cells were contested between two overlapping objects. Two rules resolve them: level assignment by object size handles most cases, and the smaller-area rule handles the rest. Neither is principled; both work, and the paper measures the residual ambiguity as small on COCO.
The same thing in a real model
import torch
from torchvision.models.detection import fcos_resnet50_fpn
torch.manual_seed(0)
model = fcos_resnet50_fpn(weights=None, weights_backbone=None, num_classes=21)
model.train()
losses = model([torch.rand(3, 320, 320)],
[{"boxes": torch.tensor([[50., 60., 180., 240.]]),
"labels": torch.tensor([3])}])
print("FCOS returns three losses:")
for k, v in losses.items():
print(f" {k:<18}{v.item():.4f}")
print("\nparameters:", f"{sum(p.numel() for p in model.parameters()):,}")
print("anchors per position:", model.anchor_generator.num_anchors_per_location())FCOS returns three losses: classification 1.1004 bbox_regression 0.9999 bbox_ctrness 0.7215 parameters: 32,161,370 anchors per position: [1, 1, 1, 1, 1]
Three losses: class, box, and centre-ness. RetinaNet has two. That third one is the extra head this design adds.
num_anchors_per_location() returns [1, 1, 1, 1, 1] rather than nine. That is the anchor-free property, stated in code. Loss values come from random initialisation with a fixed seed. Read them as a shape check, not a measurement.
The family, and what each member does differently
| Model | Year | Represents an object as | Notes |
|---|---|---|---|
| DenseBox | 2015 | four edge distances per pixel | predates the term; used for faces |
| CornerNet | 2018 | top-left and bottom-right keypoints | needs embeddings to pair corners |
| CenterNet | 2019 | one centre point plus width and height | no NMS, peak-picking on a heatmap |
| FCOS | 2019 | four edge distances plus centre-ness | the design most models copied |
| YOLOX | 2021 | FCOS-style head, decoupled towers | brought this into the YOLO line |
| YOLOv8 onward | 2023 | four edge distances as distributions | with task-aligned assignment |
CenterNet is worth a second look. It predicts a heatmap of object centres and reads off local maxima. Duplicate removal then becomes a 3x3 max-pool rather than NMS. That is the earliest widely-used NMS-free detector, and its idea reappears in YOLOv10's one-to-one head.
Common mistakes
Regressing distances without a non-negative activation. Distances cannot be negative. Use exp or relu on the raw output. Without it the model can emit boxes with inverted edges early in training.
Forgetting the per-level scale factor. Distances at stride 32 are numerically much larger than at stride 8. FCOS multiplies each level's output by a learnable scalar so one shared head can serve all levels. Without it the head has to learn a different output range per level through shared weights, and it trains poorly.
Skipping centre sampling. The naive inside-the-box rule adds many low-quality positives. Add the radius restriction.
Applying centre-ness at training but not at inference. Multiply it into the score at inference too. That is where the accuracy gain shows.
Assuming anchor-free means assignment-free. You still choose which cells are positive. FCOS uses geometry; YOLOv8 uses task-aligned scores; DETR uses Hungarian matching. The decision did not disappear, it moved.
Try it yourself
Change gt to a very wide, thin object such as [10., 60., 120., 72.] and re-run. Count the positives, and look at the centre-ness map. Then compare against the anchor coverage for the same shape in anchor boxes. The difference is the strongest practical argument for going anchor-free.
What to learn next
- DETR and set prediction — removing the assignment heuristics as well.
- YOLO versions compared — where this head design ended up in production.
- Anchor boxes — what this replaced, and why it needed replacing.
Researcher — Mathematics and papers.
The FCOS formulation
Take a location $(x, y)$ on a feature map of stride $s$, mapped back to image coordinates. Its regression target against ground-truth box $B = (x_0, y_0, x_1, y_1)$ is:
$$ l^* = x - x_0, \quad t^* = y - y_0, \quad r^* = x_1 - x, \quad b^* = y_1 - y $$
A location is positive when all four are strictly positive, meaning it falls inside $B$. The journal version adds the centre-sampling restriction on top. Centre-ness is:
$$ \text{centerness}^* = \sqrt{\frac{\min(l^, r^)}{\max(l^, r^)} \cdot \frac{\min(t^, b^)}{\max(t^, b^)}} $$
Trained with binary cross-entropy, and multiplied into the classification score at inference. The square root slows the decay away from the centre; without it, the down-weighting is too aggressive.
Multi-level assignment restricts each level $i$ to targets with $\max(l^, t^, r^, b^) \in [m_{i-1}, m_i)$, with thresholds $0, 64, 128, 256, 512, \infty$ across P3 to P7. This is a size-based routing rule, and it removes most ambiguity between overlapping objects of different scale.
Why centre-ness works
Classification confidence and localisation quality are only loosely correlated in a dense detector. A location near an object's boundary can be confidently classified while producing a poor box. NMS ranks by classification score alone. The result is that a bad box outranks a good one.
Centre-ness is a cheap proxy for localisation quality. Li et al. (2020) argue the separate branch is the wrong solution. Their Quality Focal Loss trains the classification score directly against IoU. That removes the inconsistency instead of patching it. Modern YOLO heads follow that route and have no centre-ness branch.
Keypoint-based alternatives
CornerNet (Law and Deng, 2018) predicts two heatmaps of corners plus an embedding per corner. Corners whose embeddings are close get grouped into one box. It also introduces corner pooling, a directional max over rows and columns. That is needed because a box corner often sits on no object pixel at all.
CenterNet (Zhou et al., 2019, Objects as Points) predicts a centre heatmap plus size and offset regressions. Peak extraction by 3x3 max-pooling replaces NMS. The same representation extends to 3D boxes and pose by adding output channels, which is its most attractive property.
The Gaussian target on these heatmaps has a carefully chosen radius. Any point inside it still yields an IoU above a threshold with the ground-truth box. That is a neat piece of design worth reading.
The anchor-free versus anchor-based question, settled
Zhang et al. (2020), ATSS, ran the controlled experiment. They took RetinaNet and FCOS and aligned every training detail. The remaining gap came almost entirely from how positive and negative samples are defined. Whether anchors exist mattered very little.
Their Adaptive Training Sample Selection works per ground-truth object. It takes the mean $\mu$ and standard deviation $\sigma$ of IoU over that object's $k$ closest candidates per level. The object's positive threshold becomes $\mu + \sigma$. Applying it to both detectors brought them to essentially the same AP.
The practical conclusion is that "anchor-free" is a statement about parameterisation, and the accuracy lives in assignment. That reframing is why later work concentrates on assignment: SimOTA, task-aligned assignment, and eventually Hungarian matching in DETR.
Remaining weaknesses
- Very elongated objects. The four-distance representation handles arbitrary aspect ratios. The level-routing rule, however, uses the largest edge distance. A long thin object therefore routes to a coarse level, where its short dimension is under-resolved.
- Crowded identical objects. Level routing cannot separate two same-sized overlapping objects, and the smaller-area tie-break is arbitrary between equals.
- Rotated objects. The representation is inherently axis-aligned. Oriented variants add an angle and inherit its periodicity problem.
Papers
- Tian et al., FCOS: Fully Convolutional One-Stage Object Detection, ICCV 2019 — arxiv.org/abs/1904.01355
- Law and Deng, CornerNet: Detecting Objects as Paired Keypoints, ECCV 2018 — arxiv.org/abs/1808.01244
- Zhou et al., Objects as Points, 2019 — arxiv.org/abs/1904.07850
- Zhang et al., Bridging the Gap Between Anchor-based and Anchor-free Detection via ATSS, CVPR 2020 — arxiv.org/abs/1912.02424
- Ge et al., YOLOX: Exceeding YOLO Series in 2021 — arxiv.org/abs/2107.08430
What to learn next
- DETR and set prediction — removing the assignment heuristics as well.
- YOLO versions compared — where this head design ended up in production.
- Anchor boxes — what this replaced, and why it needed replacing.