DETR and set prediction
DETR asks a transformer for a fixed-size set of objects and matches predictions to truth one-to-one, which removes anchors and duplicate removal from the pipeline entirely.
- 16 min read
- 3 reading levels
- Updated
Read these first
On this page 9
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
DETR asks a transformer for a fixed list of objects. Training pairs each answer with at most one real object.
The analogy
Think about a wedding with a hundred name cards on the tables. Every guest gets exactly one card, and every card holds at most one guest. Nobody sits twice, and spare cards say "empty seat".
Because the pairing is one-to-one, nobody has to walk around afterwards removing duplicate guests. The seating plan prevented duplicates in the first place.
That is what DETR does. Fix the number of slots. Pair each one with at most one real object. Duplicates never get created, so nothing has to delete them.
Why it exists
Every detector before this had a pile of hand-written machinery around the network.
Ready-made box shapes, chosen in advance. Rules deciding which candidates count as objects during training. A clean-up step deleting overlapping outputs afterwards. None of it learned anything, and every piece had settings.
DETR removed all of it in one move. The network outputs a set. The training procedure enforces that a set has no duplicates. That is the whole design.
How it works
picture -> backbone -> transformer encoder (looks at the whole scene)
|
v
a fixed list of empty slots, called object queries
|
v
transformer decoder fills them in
|
v
slot 1: "dog, here" slot 2: "nothing" slot 3: "person, here" ...You choose the number of slots in advance, usually one hundred. Every slot answers, and most answer "nothing here".
During training, the software works out the best possible pairing between slots and real objects. Paired slots learn the object. Unpaired slots learn to say "nothing".
The clever part
The pairing is not first-come-first-served. It looks at every possible pairing and picks the one that is best overall.
That matters. One slot might be the closest match for two different objects. Another slot might be the only one near a third object. Grabbing the closest pair first can leave that third object with nothing. Solving the whole puzzle at once avoids it.
The honest problem
DETR was slow to train. The original needed a very long schedule to match established detectors. It was also weaker on small objects.
A slot has to learn to attend to the right part of the image, starting from nothing. That takes many passes over the data. A follow-up called Deformable DETR reported better results with a tenth of the training epochs. The family has kept improving since.
Where you have seen the result
- Detection running on video with steady, predictable timing.
- Tools that count crowded objects without double-counting.
- Systems combining detection with outlines and text in one model.
- Newer real-time detectors that came out of this line of work.
Remember this
- DETR outputs a fixed list of slots, most of which say "nothing here".
- Training pairs slots to real objects one-to-one, so duplicates are never produced.
- That removes ready-made shapes, matching rules and the clean-up step in one design.
What to learn next
- Open-vocabulary detection — attaching language to this architecture.
- Transformers — the encoder and decoder underneath.
- Attention — the mechanism the queries use to find their object.
Developer — Code and libraries.
Setup
pip install torch torchvision scipyWritten against torch 2.5.1, torchvision 0.20.1 and scipy 1.14.1. The matching demo is pure CPU arithmetic with no model weights involved.
The matching step, which is the whole idea
import torch
from scipy.optimize import linear_sum_assignment
from torchvision.ops import generalized_box_iou
NUM_QUERIES = 6 # the real model uses 100 or 300
CLASSES = ["person", "dog", "car", "no-object"]
# Two real objects in the picture, in normalised centre-size form.
gt_boxes = torch.tensor([[0.20, 0.30, 0.30, 0.50],
[0.70, 0.60, 0.20, 0.30]])
gt_labels = torch.tensor([0, 1]) # person, dog
# What the decoder produced for its six queries.
pred_boxes = torch.tensor([[0.72, 0.58, 0.22, 0.28], # close to the dog
[0.50, 0.50, 0.90, 0.90], # a lazy whole-image box
[0.22, 0.31, 0.28, 0.52], # close to the person
[0.21, 0.29, 0.31, 0.49], # ALSO close to the person
[0.10, 0.90, 0.05, 0.05], # nothing
[0.90, 0.10, 0.05, 0.05]]) # nothing
pred_probs = torch.tensor([[0.05, 0.80, 0.05, 0.10],
[0.10, 0.10, 0.10, 0.70],
[0.75, 0.05, 0.05, 0.15],
[0.70, 0.05, 0.05, 0.20],
[0.02, 0.02, 0.02, 0.94],
[0.03, 0.02, 0.02, 0.93]])
def to_xyxy(b):
cx, cy, w, h = b.unbind(-1)
return torch.stack([cx - w / 2, cy - h / 2, cx + w / 2, cy + h / 2], -1)
# The DETR matching cost: class term, L1 box term, GIoU term.
cost_class = -pred_probs[:, gt_labels] # 6 x 2
cost_l1 = torch.cdist(pred_boxes, gt_boxes, p=1) # 6 x 2
cost_giou = -generalized_box_iou(to_xyxy(pred_boxes), to_xyxy(gt_boxes)) # 6 x 2
cost = 1.0 * cost_class + 5.0 * cost_l1 + 2.0 * cost_giou
print("cost matrix (rows = queries, columns = ground-truth objects):")
print(f"{'':>9}{'person':>10}{'dog':>10}")
for q in range(NUM_QUERIES):
print(f"query {q} " + "".join(f"{v:10.3f}" for v in cost[q].tolist()))
rows, cols = linear_sum_assignment(cost.numpy())
print("\nHungarian assignment:")
for q, g in zip(rows, cols):
print(f" query {q} -> {CLASSES[gt_labels[g]]:<8} (cost {cost[q, g]:.3f})")
unmatched = sorted(set(range(NUM_QUERIES)) - set(rows.tolist()))
print(f" queries {unmatched} are all trained to predict 'no-object'")
print(f" total cost of the optimal assignment: {cost[rows, cols].sum():.3f}")
person_query = rows[list(cols).index(0)]
duplicate = [q for q in (2, 3) if q != person_query][0]
print(f"\nqueries 2 and 3 both found the person.")
print(f" query {person_query} is trained towards 'person'")
print(f" query {duplicate} is trained towards 'no-object', so its score collapses")
print(" that is why DETR needs no NMS at inference")cost matrix (rows = queries, columns = ground-truth objects):
person dog
query 0 6.640 -1.831
query 1 7.030 7.752
query 2 -2.074 6.484
query 3 -2.299 6.643
query 4 7.818 8.117
query 5 9.273 6.910
Hungarian assignment:
query 0 -> dog (cost -1.831)
query 3 -> person (cost -2.299)
queries [1, 2, 4, 5] are all trained to predict 'no-object'
total cost of the optimal assignment: -4.130
queries 2 and 3 both found the person.
query 3 is trained towards 'person'
query 2 is trained towards 'no-object', so its score collapses
that is why DETR needs no NMS at inferenceReading that output
The matched costs are negative and the unmatched ones are large and positive. The cost mixes three terms, so its absolute scale means nothing. What matters is the gap: -2.299 for a good match against 6.640 for a bad one.
Query 2 and query 3 both found the person, and only one was rewarded. Their costs were -2.074 and -2.299, close enough that they are visually the same detection. The matcher picks the better one; the other is trained to say "no-object".
Run this on the same image every epoch and query 2's confidence for "person" drops towards zero. That is duplicate suppression implemented in the loss instead of in post-processing.
Four queries out of six learn "no-object". With 100 queries and an average of 7 objects per COCO image, 93 percent of queries learn background every step. DETR down-weights the no-object class by a factor of 10 in the classification loss to stop it dominating. This is the same imbalance problem as focal loss, handled with a blunter tool.
Why the matching has to be global, not greedy
import numpy as np
from scipy.optimize import linear_sum_assignment
# A cost matrix where picking the cheapest pair first is a mistake.
cost = np.array([[1.0, 2.0], # query A: good for both objects
[1.5, 9.0]]) # query B: only usable for object 1
taken_r, taken_c, greedy_total = set(), set(), 0.0
print("greedy, cheapest pair first:")
for _ in range(2):
best = min(((i, j) for i in range(2) for j in range(2)
if i not in taken_r and j not in taken_c),
key=lambda p: cost[p])
print(f" query {best[0]} -> object {best[1]} cost {cost[best]}")
taken_r.add(best[0]); taken_c.add(best[1]); greedy_total += cost[best]
print(" total:", greedy_total)
r, c = linear_sum_assignment(cost)
print("\nHungarian, minimising the total:")
for i, j in zip(r, c):
print(f" query {i} -> object {j} cost {cost[i, j]}")
print(" total:", cost[r, c].sum())greedy, cheapest pair first: query 0 -> object 0 cost 1.0 query 1 -> object 1 cost 9.0 total: 10.0 Hungarian, minimising the total: query 0 -> object 1 cost 2.0 query 1 -> object 0 cost 1.5 total: 3.5
Greedy grabbed the single cheapest pair first. That left query 1, which can only handle object 0, stuck with object 1 at cost 9. The optimal assignment sacrifices a little on the first pair to avoid a disaster on the second.
In a detector this decides whether a small object is learned at all. Greedy matching hands its one nearby query to a larger, easier object. The small one is abandoned every epoch.
Running a real DETR
transformers ships DETR and its descendants. Loading a checkpoint downloads weights, so no output block is shown here rather than an invented one.
# Requires: pip install transformers pillow requests
# Loading facebook/detr-resnet-50 downloads roughly 160 MB of weights.
import torch
from PIL import Image
from transformers import DetrImageProcessor, DetrForObjectDetection
processor = DetrImageProcessor.from_pretrained("facebook/detr-resnet-50")
model = DetrForObjectDetection.from_pretrained("facebook/detr-resnet-50")
model.eval()
image = Image.open("street.jpg").convert("RGB")
inputs = processor(images=image, return_tensors="pt")
with torch.no_grad():
outputs = model(**inputs)
# logits: (1, 100, num_labels + 1); pred_boxes: (1, 100, 4) in cxcywh, normalised
print("queries:", outputs.logits.shape[1])
results = processor.post_process_object_detection(
outputs, target_sizes=torch.tensor([image.size[::-1]]), threshold=0.9
)[0]
for score, label, box in zip(results["scores"], results["labels"], results["boxes"]):
print(f"{model.config.id2label[label.item()]:<15} {score:.3f} "
f"{[round(v, 1) for v in box.tolist()]}")Three things about that code are worth stating plainly.
The output always has 100 rows, whatever is in the image. Filtering by score is how you get from 100 to the real objects, and there is no NMS call anywhere.
post_process_object_detection converts normalised cxcywh to absolute xyxy and applies the threshold. It is easy to forget and then wonder why every box is under 1 pixel wide.
The exact class of the API moves between major transformers releases. This was checked against transformers 5.6.2, where DetrImageProcessor.post_process_object_detection exists and OWLv2's equivalent is named post_process_grounded_object_detection. Check the method names against your installed version rather than trusting a blog post.
The lineage after DETR
| Model | Year | The problem it attacked |
|---|---|---|
| DETR | 2020 | the original: set prediction, no anchors, no NMS |
| Deformable DETR | 2021 | slow convergence and small objects, via sparse multi-scale attention |
| Conditional DETR | 2021 | decoder cross-attention spatial focus |
| DAB-DETR | 2022 | queries as explicit 4D anchor boxes, refined layer by layer |
| DN-DETR | 2022 | unstable matching, via denoising auxiliary queries |
| DINO | 2022 | combines the above; contrastive denoising and mixed query selection |
| RT-DETR | 2023 | real-time throughput, efficient hybrid encoder |
| D-FINE, DEIM | 2024-2025 | box refinement and matching quality |
| RT-DETRv4 | 2026 | training-time distillation from a vision foundation model |
The consistent theme is what happened to "the query". It started as an abstract learned vector. It became a concrete, interpretable box that the decoder refines. That change is what fixed convergence.
Common mistakes
Setting num_queries too low. If an image can contain 200 objects and you have 100 queries, you have capped recall at 100. Count the maximum objects per image in your data first.
Expecting fast convergence on a small dataset. Set prediction needs a lot of gradient steps to teach queries where to look. Fine-tune from a pretrained checkpoint; do not train from scratch on a few thousand images.
Adding NMS anyway. It usually lowers AP slightly and always hides whether the model actually learned one-to-one behaviour. If duplicates persist, the matching or the schedule is the problem.
Forgetting the box format. DETR predicts normalised cxcywh. Feeding those straight into box_iou gives silent nonsense. See bounding box formats.
Ignoring auxiliary losses. DETR applies the loss after every decoder layer. Disabling that to save memory measurably slows convergence.
Try it yourself
Add a third ground-truth object to the first script, positioned so that query 3 is also its best match. Watch the assignment change: query 3 can only take one, so a worse query is forced onto the other object. That forced compromise, repeated every batch, is the pressure that spreads queries across the image.
What to learn next
- Open-vocabulary detection — attaching language to this architecture.
- Transformers — the encoder and decoder underneath.
- Attention — the mechanism the queries use to find their object.
Researcher — Mathematics and papers.
The set prediction loss
DETR predicts a fixed set $\hat{y} = {\hat{y}i}{i=1}^{N}$ with $N$ larger than the maximum object count. The ground-truth set $y$ is padded with $\varnothing$ (no-object) to size $N$.
The optimal permutation $\hat{\sigma}$ minimises a total matching cost:
$$ \hat{\sigma} = \arg\min_{\sigma \in \mathfrak{S}N} \sum{i=1}^{N} \mathcal{L}_{\text{match}}\big(y_i, \hat{y}_{\sigma(i)}\big) $$
Where $\mathfrak{S}_N$ is the set of permutations of $N$ elements. The matching cost for a non-empty target with class $c_i$ and box $b_i$ is:
$$ \mathcal{L}{\text{match}} = -\mathbb{1}{{c_i \neq \varnothing}} \hat{p}_{\sigma(i)}(c_i)
- \mathbb{1}_{{c_i \neq \varnothing}} \mathcal{L}_{\text{box}}\big(b_i, \hat{b}_{\sigma(i)}\big) $$
Note it uses the raw probability $\hat p$, not its log. The Hungarian algorithm (Kuhn, 1955) solves this in $O(N^3)$; scipy.optimize.linear_sum_assignment implements a modern variant of it. With $N = 100$ this is negligible next to the forward pass.
The training loss then applies negative log-likelihood over classes and the box loss on matched pairs:
$$ \mathcal{L}{\text{Hungarian}} = \sum{i=1}^{N}\Big[-\log \hat{p}_{\hat\sigma(i)}(c_i)
- \mathbb{1}_{{c_i \neq \varnothing}} \mathcal{L}_{\text{box}}\big(b_i, \hat{b}_{\hat\sigma(i)}\big)\Big] $$
With $\mathcal{L}{\text{box}} = \lambda{L1}|b - \hat b|1 + \lambda{\text{giou}}\mathcal{L}_{\text{giou}}$, weights 5 and 2 in the paper. The log-probability weight for $c_i = \varnothing$ is divided by 10 to counter the class imbalance. The matching cost and the training loss are deliberately different functions.
Why the matching is a permutation, not a soft assignment
The one-to-one constraint is what removes NMS. Any relaxation that lets two predictions share a target reintroduces duplicates.
This is also the source of the instability that DN-DETR (Li et al., 2022) identified. The optimal permutation can flip between epochs for the same image. A query then receives contradictory targets across steps. Their fix adds noised copies of ground-truth boxes as extra queries. Those have a fixed, known assignment. The model gets a stable denoising signal alongside the unstable matching one.
The convergence problem, diagnosed
Zhu et al. (2021), Deformable DETR, state the cause directly. DETR "suffers from slow convergence and limited feature spatial resolution". Transformer attention over image features starts nearly uniform, and needs many epochs to become selective. Dense attention over a high-resolution map is also quadratic in cost.
Their fix is multi-scale deformable attention. Each query predicts $K$ sampling offsets per level and attends only to those points:
$$ \mathrm{MSDeformAttn}(z_q, \hat{p}q, {x^l}) = \sum{m=1}^{M} W_m \left[\sum_{l=1}^{L}\sum_{k=1}^{K} A_{mlqk} \cdot W'_m x^l\big(\phi_l(\hat p_q) + \Delta p_{mlqk}\big)\right] $$
Here $z_q$ is the query feature and $\hat p_q$ its normalised reference point. $\phi_l$ rescales that point to level $l$. $\Delta p_{mlqk}$ are learned sampling offsets. $A_{mlqk}$ are learned attention weights that sum to one over $l$ and $k$. The index $m$ runs over attention heads. Cost becomes linear in feature-map size, and multi-scale features become affordable, which is what fixes small-object accuracy.
The paper reports better performance than DETR, especially on small objects, with ten times fewer training epochs.
Query design, the thread that runs through the whole family
- DETR: queries are free learned embeddings with no spatial meaning.
- Conditional DETR (Meng et al., 2021): decodes a spatial query from the reference point, decoupling content and position.
- DAB-DETR (Liu et al., 2022): queries are 4D boxes $(x, y, w, h)$, refined at every decoder layer, with width and height modulating the attention's spatial extent.
- DINO (Zhang et al., 2022): mixed query selection initialises positional queries from encoder outputs, keeping content queries learnable. It adds contrastive denoising with positive and negative noise levels.
Read them in that order and one idea gets progressively more explicit. It converges on something close to a learned, iteratively-refined proposal set. Sparse R-CNN arrived at a similar place from the two-stage direction.
Where the line stands
RT-DETR (Zhao et al., 2023) made the family real-time. Its hybrid encoder decouples intra-scale interaction from cross-scale fusion, and query selection becomes IoU-aware. The line continued through RT-DETRv2, D-FINE and DEIM.
RT-DETRv4 (ECCV 2026) reports 55.4 AP at 124 FPS for its L variant. That is against DEIM-L at 54.7 and YOLOv13-L at 53.4. Its X variant reports 57.0 AP. Its contribution is a training-only Deep Semantic Injector. It aligns the deepest feature map with a frozen DINOv3 backbone, so inference cost is unchanged.
That last detail points at the most interesting current trend in detection. Gains come from what supervises the model during training, not from what runs at inference.
Papers
- Carion et al., End-to-End Object Detection with Transformers, ECCV 2020 — arxiv.org/abs/2005.12872
- Zhu et al., Deformable DETR, ICLR 2021 — arxiv.org/abs/2010.04159
- Liu et al., DAB-DETR: Dynamic Anchor Boxes are Better Queries for DETR, ICLR 2022 — arxiv.org/abs/2201.12329
- Li et al., DN-DETR: Accelerate DETR Training by Introducing Query DeNoising, CVPR 2022 — arxiv.org/abs/2203.01305
- Zhang et al., DINO: DETR with Improved DeNoising Anchor Boxes, ICLR 2023 — arxiv.org/abs/2203.03605
- Zhao et al., DETRs Beat YOLOs on Real-time Object Detection, CVPR 2024 — arxiv.org/abs/2304.08069
What to learn next
- Open-vocabulary detection — attaching language to this architecture.
- Transformers — the encoder and decoder underneath.
- Attention — the mechanism the queries use to find their object.