Non-maximum suppression
NMS keeps the highest-scoring box in a cluster and deletes its overlapping neighbours, which is how thousands of raw detections become a handful of answers.
- 13 min read
- 3 reading levels
- Updated
Read these first
On this page 8
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
NMS keeps the most confident box in a pile of overlapping boxes and throws the rest away.
The analogy
Think about a crowded room where six people all shout the answer to the same question. They are saying almost the same thing at almost the same time. You do not need six answers, you need one.
So you listen for the loudest, confident voice, write that down, and ignore everyone who was saying the same thing. Then you look for the next voice saying something different, and repeat.
Non-maximum suppression does exactly this with rectangles. Loudest first, then silence anyone overlapping them.
Why it exists
A detector does not output one box per object. It scores tens of thousands of candidate rectangles, and an obvious car lights up dozens of them.
Every one of those is a real prediction with a real score. Reporting all of them would count one car as thirty cars. Something has to reduce the pile before a human or a downstream system sees it.
How it works
raw detections after NMS
+----+ +----+
|car | 0.92 |car | 0.92
+----+ +----+
+----+ 0.88 (overlaps 0.92)
+----+
+-----+ 0.75 (overlaps 0.92)
+-----+
+----+ 0.81 +----+ 0.81
|car2| |car2|
+----+ +----+Three steps, repeated until nothing is left.
Take the highest-scoring box that has not been dealt with. Keep it. Delete every remaining box that overlaps it more than a chosen amount. Go again.
The "chosen amount" is a setting you pick. It has a large effect, and the next section explains why.
The setting that ruins people's day
Set the overlap limit too low and NMS becomes greedy. Two motorbikes parked side by side genuinely overlap. A strict setting decides the second one is a duplicate and deletes a real vehicle.
Set it too high and the duplicates come back. You count one bus three times.
There is no correct value. It depends on how crowded your scenes are. Anyone who tells you the value without asking about your data has not thought about it.
Where you have already seen it
- A phone camera drawing one box per face, not twelve.
- Traffic counting software reporting one number, not a wild overcount.
- A shop shelf scanner listing each bottle once.
- Any detection demo video where the boxes look clean and stable.
Remember this
- NMS keeps the most confident box and deletes its overlapping neighbours.
- It runs after the model, on the output, and it learns nothing.
- The overlap setting trades duplicate boxes against deleting real, crowded objects.
What to learn next
- Feature pyramid networks — where all those candidate boxes come from.
- DETR and set prediction — the architecture that makes NMS unnecessary.
- YOLO — a detector whose speed depends on this step.
Developer — Code and libraries.
Setup
pip install torch torchvisionWritten against torch 2.5.1 and torchvision 0.20.1. torchvision.ops.nms runs on CPU and CUDA and is the one you should ship.
NMS by hand, then the library, then the failure modes
import torch
from torchvision.ops import nms, batched_nms, box_iou
# Two scooters parked close together. The detector fired five times.
boxes = torch.tensor([
[100., 100., 200., 260.], # scooter A, best guess
[108., 106., 194., 258.], # A again
[ 90., 94., 214., 268.], # A again, looser
[140., 100., 240., 260.], # scooter B, best guess
[147., 104., 235., 256.], # B again
])
scores = torch.tensor([0.92, 0.88, 0.75, 0.81, 0.79])
def nms_by_hand(boxes, scores, thresh):
order = scores.argsort(descending=True) # loudest claim goes first
keep = []
while order.numel() > 0:
i = order[0].item()
keep.append(i)
rest = order[1:]
if rest.numel() == 0:
break
ious = box_iou(boxes[i:i + 1], boxes[rest]).squeeze(0)
order = rest[ious <= thresh] # anything hugging the winner is deleted
return torch.tensor(keep)
print("kept by hand :", nms_by_hand(boxes, scores, 0.5).tolist())
print("kept by torchvision:", nms(boxes, scores, 0.5).tolist())
print("\nthe two boxes that survive at 0.5:")
for i in nms(boxes, scores, 0.5).tolist():
print(f" box {i} score {scores[i]:.2f} {[int(v) for v in boxes[i].tolist()]}")
print("\nthe same five boxes at different thresholds:")
for t in [0.3, 0.4, 0.5, 0.7, 0.9]:
print(f" iou_threshold={t} -> kept {nms(boxes, scores, t).tolist()}")
print("\nhow much each box overlaps scooter A's best box:")
print(" ", [round(v, 3) for v in box_iou(boxes[:1], boxes).squeeze(0).tolist()])
# Different classes stacked on top of each other.
stacked = torch.tensor([[10., 10., 110., 210.], # a person
[12., 60., 108., 200.]]) # the bag they carry
ssc = torch.tensor([0.90, 0.85])
print(f"\nperson and bag overlap at IoU {box_iou(stacked[:1], stacked[1:]).item():.2f}")
print(" class-blind nms keeps:", nms(stacked, ssc, 0.5).tolist(), " <- the bag is deleted")
print(" batched_nms keeps :",
sorted(batched_nms(stacked, ssc, torch.tensor([0, 26]), 0.5).tolist()))
# Soft-NMS: lower the neighbour's score instead of deleting it.
def soft_nms(boxes, scores, sigma=0.5, min_score=0.05):
s, out = scores.clone(), []
while (s > min_score).any():
i = torch.where(s > min_score, s, torch.full_like(s, -1.0)).argmax()
out.append((i.item(), round(s[i].item(), 3)))
ious = box_iou(boxes[i:i + 1], boxes).squeeze(0)
s = s * torch.exp(-(ious ** 2) / sigma) # gaussian penalty, nothing is ever deleted
s[i] = 0.0
return out
print("\nsoft-nms, index and surviving score:")
for i, sc in soft_nms(boxes, scores):
print(f" box {i} {sc}")kept by hand : [0, 3] kept by torchvision: [0, 3] the two boxes that survive at 0.5: box 0 score 0.92 [100, 100, 200, 260] box 3 score 0.81 [140, 100, 240, 260] the same five boxes at different thresholds: iou_threshold=0.3 -> kept [0] iou_threshold=0.4 -> kept [0, 4] iou_threshold=0.5 -> kept [0, 3] iou_threshold=0.7 -> kept [0, 3] iou_threshold=0.9 -> kept [0, 1, 3, 4, 2] how much each box overlaps scooter A's best box: [1.0, 0.817, 0.742, 0.429, 0.378] person and bag overlap at IoU 0.67 class-blind nms keeps: [0] <- the bag is deleted batched_nms keeps : [0, 1] soft-nms, index and surviving score: box 0 0.92 box 4 0.594 box 2 0.178 box 3 0.091 box 1 0.063
Reading that output, which is more interesting than it looks
The threshold sweep contains three separate failures.
At 0.3, only one box survives. Scooter B overlaps scooter A at IoU 0.429, so a strict threshold deleted a real vehicle. This is the crowded-scene failure, and it is the reason pedestrian detection papers care so much about NMS.
At 0.4, the kept boxes are [0, 4]. Box 3 was the better detection of scooter B, and it was suppressed by box 0 at IoU 0.429. Box 4 overlapped box 0 by only 0.378, so it survived. NMS is greedy, and greed can keep the worse member of a cluster. Nothing in the algorithm notices.
At 0.9 every box comes back, including all three duplicates of scooter A.
Class-blind NMS deleted the bag. A person and the bag they carry overlap at 0.67. nms does not look at labels, so it treats them as duplicates. batched_nms takes an idxs tensor and offsets the coordinates per class internally. Boxes of different classes can then never suppress each other. Use batched_nms in any multi-class detector.
Soft-NMS keeps everything and reorders it. Box 0 stays at 0.92. Box 4 decays to 0.594. Box 3, the better scooter-B detection, decays to 0.091 because it overlapped both box 0 and box 4.
That last number is a warning. Soft-NMS raises recall on crowded benchmarks, but it returns far more boxes and it can demote the box you wanted. You still need a score threshold afterwards, and now that threshold matters more.
The full post-processing pipeline
NMS is the last of four steps, not the only one.
import torch
from torchvision.ops import batched_nms, clip_boxes_to_image, remove_small_boxes
torch.manual_seed(0)
N = 20_000
boxes = torch.rand(N, 4) * 640
boxes[:, 2:] += boxes[:, :2] # make x2 > x1 and y2 > y1
scores = torch.rand(N) ** 8 # most scores near zero, a few high
labels = torch.randint(0, 80, (N,))
print("raw candidates :", len(boxes))
keep = scores > 0.05 # 1. score threshold, does most of the work
boxes, scores, labels = boxes[keep], scores[keep], labels[keep]
print("after score threshold 0.05 :", len(boxes))
order = scores.argsort(descending=True)[:1000] # 2. keep only the top 1000 into NMS
boxes, scores, labels = boxes[order], scores[order], labels[order]
print("after top-k truncation :", len(boxes))
boxes = clip_boxes_to_image(boxes, (640, 640)) # 3. tidy the geometry
keep = remove_small_boxes(boxes, min_size=2)
boxes, scores, labels = boxes[keep], scores[keep], labels[keep]
print("after clipping and size filter:", len(boxes))
keep = batched_nms(boxes, scores, labels, 0.5)[:100] # 4. NMS, then cap the output
print("after per-class NMS, capped :", len(keep))raw candidates : 20000 after score threshold 0.05 : 6265 after top-k truncation : 1000 after clipping and size filter: 991 after per-class NMS, capped : 100
These counts come from fake random scores with torch.manual_seed(0). They are reproducible, and they are not measurements of a real model. With a trained detector the number surviving the score threshold depends entirely on your score distribution. What does not change is the shape of the pipeline: threshold first, truncate second, NMS last. NMS is the expensive step, so handing it 1,000 boxes instead of 20,000 is where the latency saving lives.
Common mistakes
Running class-blind NMS on a multi-class model. Use batched_nms. This is the most common post-processing bug. It shows up as missing objects in exactly the scenes people care about.
Skipping the score threshold. NMS on 20,000 boxes is slow and pointless. Nearly all of them will be deleted by a threshold anyway.
Tuning the NMS threshold on the training set. It interacts with the score threshold and with how crowded your images are. Tune both together on a validation set that looks like production.
Assuming NMS is differentiable. It is not. It sits outside the graph. A detector whose accuracy depends heavily on NMS settings has moved decision-making into code that cannot learn. That is exactly the argument DETR makes.
Forgetting NMS on the deployed graph. Export a model to ONNX and NMS may or may not come with it, depending on the exporter. Test the exported artefact end to end, not the PyTorch model.
Try it yourself
Add a third scooter overlapping scooter B at IoU 0.45, with a score of 0.6. Find a single iou_threshold that keeps all three and no duplicates. If you cannot find one, you have discovered the real limitation of NMS by yourself. You are then ready for anchor-free detection and DETR.
What to learn next
- Feature pyramid networks — where all those candidate boxes come from.
- DETR and set prediction — the architecture that makes NMS unnecessary.
- YOLO — a detector whose speed depends on this step.
Researcher — Mathematics and papers.
The algorithm
Given detections $\mathcal{B} = {b_i}$ with scores ${s_i}$ and threshold $N_t$:
$$ \text{repeat: } m = \arg\max_i s_i, \quad \mathcal{D} \leftarrow \mathcal{D} \cup {b_m}, \quad \mathcal{B} \leftarrow \mathcal{B} \setminus {b_i : \mathrm{IoU}(b_m, b_i) \geq N_t} $$
Greedy NMS is $O(n^2)$ in the worst case and $O(n \log n + nk)$ in practice, where $k$ is the number of survivors. It is sequential by construction, which is why GPU implementations parallelise the IoU matrix but keep the selection loop serial.
Greedy NMS does not solve any stated optimisation problem. It is a heuristic that happens to work. Duplicate removal as maximum weight independent set on the overlap graph is NP-hard. Greedy selection is a standard approximation to it.
Soft-NMS
Bodla et al. (2017) replace deletion with a score decay:
$$ s_i \leftarrow \begin{cases} s_i, & \mathrm{IoU}(b_m, b_i) < N_t \ s_i \, f(\mathrm{IoU}(b_m, b_i)), & \text{otherwise} \end{cases} $$
With linear $f(x) = 1 - x$ or Gaussian $f(x) = e^{-x^2 / \sigma}$, $\sigma$ conventionally 0.5. The Gaussian form is continuous, which avoids a discontinuity at $N_t$.
Reported gains are around 1 AP on COCO with no retraining, concentrated at high IoU thresholds and in crowded scenes. The costs are real: the output set is much larger, and the score distribution is no longer calibrated.
Variants worth knowing
- DIoU-NMS (Zheng et al., 2020) suppresses on $\mathrm{IoU} - \mathcal{R}_{\mathrm{DIoU}}$, so two boxes with distant centres survive even at high overlap. Useful for occluded pedestrians.
- Matrix NMS (Wang et al., 2020, SOLOv2) computes all decay factors in one parallel matrix operation, removing the sequential loop.
- Cluster-NMS batches the iteration over independent clusters of the overlap graph.
- Weighted NMS / box voting replaces the winner with a score-weighted average of its suppressed neighbours, improving localisation at negligible cost.
- Non-maximum merging, used by SAHI for tiled inference, takes the union of overlapping boxes instead of picking one. This matters when both boxes are truncated; see detecting very small objects.
Learning to remove NMS
The deeper objection is architectural. NMS encodes a prior the network could learn, and it is not differentiable.
Hosang et al. (2017) trained a network to perform NMS. DETR (Carion et al., 2020) removed the need by training with one-to-one bipartite matching: only one query is ever rewarded per object, so duplicates are never produced. YOLOv10 (Wang et al., 2024) keeps a one-to-many head for training signal and a one-to-one head for inference. That gives NMS-free deployment without DETR's convergence cost. Ultralytics YOLO26, released January 2026, is end-to-end by default on the same principle.
The measurable payoff is latency variance rather than mean latency. NMS cost depends on the number of detections, so a crowded frame costs more than an empty one. Removing it makes inference time constant. For a real-time pipeline that matters far more than a few tenths of a millisecond of average.
Evaluation interaction
mAP is computed after NMS, and it is sensitive to how many boxes you return. COCO evaluation caps detections at 100 per image. Returning fewer than that costs recall; returning more than the cap does nothing.
This creates a mild perverse incentive. A slightly permissive NMS threshold often raises mAP while making the output visibly worse to a human. Report a precision-recall curve or a fixed-score-threshold precision alongside mAP if the output is going in front of people.
Papers
- Bodla et al., Soft-NMS: Improving Object Detection With One Line of Code, ICCV 2017 — arxiv.org/abs/1704.04503
- Hosang et al., Learning Non-Maximum Suppression, CVPR 2017 — arxiv.org/abs/1705.02950
- Zheng et al., Distance-IoU Loss, AAAI 2020 — arxiv.org/abs/1911.08287
- Wang et al., YOLOv10: Real-Time End-to-End Object Detection, NeurIPS 2024 — arxiv.org/abs/2405.14458
What to learn next
- Feature pyramid networks — where all those candidate boxes come from.
- DETR and set prediction — the architecture that makes NMS unnecessary.
- YOLO — a detector whose speed depends on this step.