Object detection
Object detection finds every object in a picture and draws a box around each one, so you get what it is and where it is, not only what is present.
- 24 min read
- 3 reading levels
- Updated
Read these first
On this page 8
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Object detection finds each thing in a picture, draws a box around it, and names it.
Stand on a railway platform waiting for your family. You do not look at the crowd and think "there are people here". You point. Amma, near the tea stall. Your brother, by the second pillar. A stranger's suitcase, about to be left behind.
Every one of those thoughts has two halves: what it is, and where it is. Object detection is a machine doing exactly that.
What it adds to the last lesson
Image classification gives one label for a whole picture. That is enough for many questions and useless for many others.
CLASSIFICATION -> "this photo contains a dog"
DETECTION -> "a dog, in this rectangle, 92 out of 100 sure"
"a cat, in this rectangle, 85 out of 100 sure"
"a chair, in this rectangle, 61 out of 100 sure"Classification answers one question about the whole frame. Detection answers three questions about every object it can find. It does not know in advance how many objects there are.
The three things you get back
Each detection is a small package of three items.
+--------------------------------------------------+
| label "dog" |
| box left 120, top 40, right 300, bottom 260 |
| confidence 0.92 |
+--------------------------------------------------+The box is called a bounding box — the smallest rectangle that fits around the object. It is stored as four numbers, marking its left, top, right and bottom edges.
The confidence is how sure the model is, running from nothing to complete certainty. It is not a probability of being correct in any honest sense. Treat it as a dial you can turn, not as a promise.
Why the model shouts the same answer five times
Here is the part that always surprises people, and it explains half of how detection actually works.
A detector does not look once and produce a tidy list. It checks thousands of possible rectangles across the picture, at many sizes and positions. When there really is a dog, the rectangles near the dog all fire.
So one dog produces five, ten or twenty overlapping boxes, all saying "dog", all a bit different.
Picture a classroom where you ask a question and eight children answer at once, all saying nearly the same thing. You listen to the loudest, confident one, and gently quiet everyone who is repeating it. Then you look for a hand across the room saying something different, and listen to that one too.
That is exactly the cleanup step. Keep the strongest box. Delete anything that overlaps it heavily. Repeat with whatever is left.
The proper name for this is non-maximum suppression, usually shortened to NMS. It means: keep the maximum, suppress the rest. It runs after the model on every detector you will ever use.
Measuring whether two boxes agree
To decide "overlaps it heavily", you need a way to score how much two rectangles agree.
The measure everyone uses looks at the part the two boxes share. It compares that with the total space they cover between them. Two boxes stacked exactly on each other share everything, so they score the highest. Two boxes in opposite corners share nothing, so they score zero.
Its name is intersection over union, usually shortened to IoU. Intersection is the shared part. Union is everything either box covers. That is the whole idea, and the Developer section works through real numbers.
Where you have already seen it
- A phone camera drawing yellow squares on faces before you press the button.
- Toll gates and parking gates finding the number plate rectangle before reading it.
- Cricket coverage tracking the ball and the players through a frame.
- Warehouse cameras counting boxes on a pallet.
- A driver-assist car marking other vehicles and pedestrians on its display.
What is honestly hard about this
Detection is a real step up in difficulty from classification, and it is fair to know why before you start.
Labelling is slow and expensive. For classification, a person clicks one button per image. For detection, a person drags a rectangle around every single object in every single image. A crowded street scene can take five minutes to label properly. This, not the model, is usually what stops a project.
Small objects are genuinely hard. A face fifteen dots across has almost no detail left. Every detector is much worse at small things, and published scores usually hide this behind one average number.
Crowds break it. When twenty objects overlap heavily, the cleanup step cannot tell "two people standing close together" from "one person detected twice". It will delete a real detection. This is an unsolved problem, not a tuning issue.
The headline score is confusing. Detection results are reported as something called mAP, and it is not a percentage of correct answers. A model with a score of 40 is not "40 percent right". It is explained properly in the Researcher section, and until then, do not read it as accuracy.
Remember this
- Detection gives what, where and how sure, for every object it finds.
- One object produces many overlapping boxes, and a cleanup step keeps the best one.
- IoU measures how much two boxes agree. NMS uses it to remove duplicates.
- Labelling effort, small objects and crowds are the real difficulties, not the model.
What to learn next
- YOLO — the one-stage family that made detection fast enough for video.
- Image segmentation — from rectangles to exact object outlines.
- Model evaluation — precision, recall and the curves mAP is built from.
Developer — Code and libraries.
Two ideas make a detector work, and neither needs a neural network to understand. Build both by hand first, then run a real pretrained model on a real photograph.
Everything here is CPU-only. Nothing needs CUDA.
Setup
pip install torch torchvision matplotlibBe honest with yourself about the download before you start. The CPU-only PyTorch build is a few hundred megabytes. On a metered connection, do this once, on wifi, and keep the environment.
The pretrained detector itself is small: 13.4 MB, downloaded once and cached in your home directory. matplotlib is here only because it ships a photograph inside the package, so you do not need to find or download one.
IoU, written out
def iou(a, b):
"""Overlap of two boxes, each written as [left, top, right, bottom]."""
x1 = max(a[0], b[0])
y1 = max(a[1], b[1])
x2 = min(a[2], b[2])
y2 = min(a[3], b[3])
overlap_w = max(0.0, x2 - x1) # zero when the boxes miss each other
overlap_h = max(0.0, y2 - y1)
overlap = overlap_w * overlap_h
area_a = (a[2] - a[0]) * (a[3] - a[1])
area_b = (b[2] - b[0]) * (b[3] - b[1])
return overlap / (area_a + area_b - overlap)
truth = [10, 10, 60, 60]
guesses = {
"almost perfect": [12, 12, 62, 62],
"half right ": [35, 35, 85, 85],
"corner touch ": [55, 55, 105, 105],
"nowhere near ": [120, 120, 170, 170],
}
print("truth box:", truth)
print()
for name, box in guesses.items():
print(f"{name} {str(box):22s} IoU = {iou(truth, box):.3f}")truth box: [10, 10, 60, 60] almost perfect [12, 12, 62, 62] IoU = 0.855 half right [35, 35, 85, 85] IoU = 0.143 corner touch [55, 55, 105, 105] IoU = 0.005 nowhere near [120, 120, 170, 170] IoU = 0.000
Read those four numbers carefully, because they calibrate your intuition for the rest of your career.
A box shifted by two pixels in each direction scores 0.855. A box shifted by twenty-five pixels scores 0.143. That box is still on the same object, and still a reasonable guess by eye. IoU punishes shift far harder than people expect.
The standard "a detection counts as correct" bar is IoU above 0.5. That is a much tighter box than the phrase "half right" suggests.
The two max(0.0, ...) lines are the whole trick for non-overlapping boxes. Without them, two distant boxes produce a negative width and a negative height. Those multiply to a positive area, giving a confident, meaningless answer.
NMS, written out
def iou(a, b):
x1, y1 = max(a[0], b[0]), max(a[1], b[1])
x2, y2 = min(a[2], b[2]), min(a[3], b[3])
overlap = max(0.0, x2 - x1) * max(0.0, y2 - y1)
area_a = (a[2] - a[0]) * (a[3] - a[1])
area_b = (b[2] - b[0]) * (b[3] - b[1])
return overlap / (area_a + area_b - overlap)
def nms(boxes, scores, threshold=0.5):
"""Keep the best box, delete everything that overlaps it too much, repeat."""
order = sorted(range(len(boxes)), key=lambda i: scores[i], reverse=True)
kept = []
while order:
best = order.pop(0)
kept.append(best)
order = [i for i in order if iou(boxes[best], boxes[i]) < threshold]
return kept
# What a detector really produces: several boxes per object, all slightly different.
boxes = [
[ 10, 10, 60, 60], # dog, good
[ 12, 8, 58, 62], # dog, nearly the same box
[ 14, 14, 64, 58], # dog, again
[100, 30, 150, 80], # cat, good
[ 98, 33, 152, 77], # cat, nearly the same box
]
scores = [0.92, 0.88, 0.71, 0.85, 0.64]
labels = ["dog", "dog", "dog", "cat", "cat"]
print("boxes coming out of the model:", len(boxes))
kept = nms(boxes, scores, threshold=0.5)
print("boxes surviving NMS :", len(kept))
print()
for i in kept:
print(f" kept {labels[i]} score {scores[i]:.2f} box {boxes[i]}")
print()
for i in range(len(boxes)):
if i not in kept:
winner = max(kept, key=lambda k: iou(boxes[k], boxes[i]))
print(f" dropped box {i} (score {scores[i]:.2f}) - "
f"IoU {iou(boxes[winner], boxes[i]):.2f} with a stronger box")boxes coming out of the model: 5 boxes surviving NMS : 2 kept dog score 0.92 box [10, 10, 60, 60] kept cat score 0.85 box [100, 30, 150, 80] dropped box 1 (score 0.88) - IoU 0.86 with a stronger box dropped box 2 (score 0.71) - IoU 0.76 with a stronger box dropped box 4 (score 0.64) - IoU 0.82 with a stronger box
Five boxes in, two objects out. That is the entire algorithm, in eight lines.
Two details in this code carry real weight. order.pop(0) takes the highest-scoring survivor — NMS is greedy and never revisits a decision. And the list comprehension deletes overlapping boxes permanently. That is the failure mode from the beginner section. Two people standing shoulder to shoulder produce boxes with high IoU, and one of them is silently erased.
Production code does this per class, not across all classes at once. A dog box and a chair box that overlap are both correct, and suppressing one because of the other is a bug. In PyTorch, torchvision.ops.batched_nms handles the per-class bookkeeping for you.
A real detector, on a real photograph
Now the actual thing. This downloads 13.4 MB of weights the first time and nothing afterwards.
import torch
import matplotlib.cbook as cbook
from torchvision.io import decode_image
from torchvision.models.detection import (
ssdlite320_mobilenet_v3_large,
SSDLite320_MobileNet_V3_Large_Weights,
)
# A photograph that ships inside matplotlib, so no image is downloaded.
image = decode_image(cbook.get_sample_data("grace_hopper.jpg").name)
print("image tensor:", tuple(image.shape), image.dtype)
weights = SSDLite320_MobileNet_V3_Large_Weights.DEFAULT
print("weights file:", weights.meta["_file_size"], "MB")
print("parameters :", f'{weights.meta["num_params"]:,}')
print("COCO box mAP:", weights.meta["_metrics"]["COCO-val2017"]["box_map"])
model = ssdlite320_mobilenet_v3_large(weights=weights)
model.eval() # turns off training-only behaviour
with torch.no_grad(): # no gradients needed, so do not build them
result = model([weights.transforms()(image)])[0]
names = weights.meta["categories"]
scores = result["scores"]
print()
print("boxes the model returned:", len(scores))
for cut in (0.5, 0.3, 0.1, 0.05):
print(f" boxes scoring above {cut:.2f}: {int((scores > cut).sum())}")
print()
print("everything above 0.05:")
for box, label, score in zip(result["boxes"], result["labels"], result["scores"]):
if score <= 0.05:
break # results arrive sorted by score
left, top, right, bottom = (int(v) for v in box)
print(f" {names[label]:8s} {score:.2f} "
f"left={left:3d} top={top:3d} right={right:3d} bottom={bottom:3d}")image tensor: (3, 600, 512) torch.uint8 weights file: 13.418 MB parameters : 3,440,060 COCO box mAP: 21.3 boxes the model returned: 300 boxes scoring above 0.50: 1 boxes scoring above 0.30: 1 boxes scoring above 0.10: 2 boxes scoring above 0.05: 7 everything above 0.05: person 1.00 left= 1 top= 6 right=510 bottom=596 tie 0.14 left=228 top=422 right=286 bottom=517 tie 0.07 left=222 top=391 right=298 bottom=525 tie 0.07 left=218 top=450 right=276 bottom=568 tie 0.06 left=267 top=419 right=293 bottom=453 tie 0.05 left=228 top=421 right=275 bottom=468 tie 0.05 left=245 top=419 right=279 bottom=459
On the first run, torchvision also prints a download progress bar. It goes to the error stream rather than the normal output, so it is not part of the block above, and it does not appear again once the file is cached.
Scores may move in the last decimal place on a different PyTorch build. The shape of the result will not.
On the laptop CPU used to write this lesson, the model call took well under a tenth of a second. A 3.4 million parameter detector is genuinely practical without a GPU.
Read that output properly — it is the whole lesson
Three hundred boxes came back. NMS has already run inside the model, and torchvision then caps the result at 300 detections per image. Almost all of them score near zero.
One box is confident and correct. A person, filling nearly the whole frame, at maximum confidence.
Six low-scoring boxes are all pointing at the same real object. The photograph is a portrait of Grace Hopper in naval uniform, and she is wearing a black tie. Every one of those six boxes sits over it. The model has genuinely found a second object. It has also scored it so low that any sensible threshold throws it away.
Look at why NMS left six copies of one tie. The box [267, 419, 293, 453] sits entirely inside the box [222, 391, 298, 525]. Their shared area is the whole of the small box, about 884 pixels. Their union is the whole of the large box, about 10,184. That is an IoU near 0.09, nowhere close to any suppression threshold. A small box swallowed by a large one survives NMS every time. The algorithm did exactly what it was told and still left you a mess.
The threshold is doing all the work. Above 0.5 you get one object and miss a real one. Above 0.05 you find both and get five redundant copies of the second. There is no setting that gives you two clean boxes on this model.
That is the honest shape of detection. The score threshold is not a detail you tune at the end. It is the trade-off between missing things and cluttering the output, and only your own data can tell you where to put it.
Line by line, the parts worth stopping on
weights.transforms() returns the exact preprocessing this checkpoint was trained with — resize rule, scaling and channel order. Never rebuild it from memory. A bilinear-versus-bicubic mismatch alone measurably costs accuracy.
model.eval() switches batch normalisation to its stored running statistics and disables dropout. Forgetting it on a detector produces results that are wrong in a way that looks like a bad model rather than a bug.
torch.no_grad() stops PyTorch recording operations for backpropagation. It cuts memory sharply and speeds inference up. There is no reason to omit it when you are not training.
The model takes a list and returns a list. model([one_image])[0] looks awkward and is deliberate. Detection inputs can have different sizes, so they cannot be stacked into one tensor the way classification inputs can.
box_map is 21.3, and that is a decent score. A classifier reporting 21.3 would be worthless. Detection scores live on a different scale entirely, explained in the Researcher block.
Common mistakes
Reading the confidence as a probability of correctness. A 0.9 does not mean nine out of ten such detections are right. Detector scores are badly calibrated out of the box. If a downstream decision depends on the number, measure the relationship on your own labelled data.
Leaving the score threshold at the library default. Defaults are chosen to look good on a benchmark. Your application has its own cost for a miss and its own cost for a false alarm, and those set the threshold.
Running NMS across all classes at once. Overlapping objects of different classes are both correct. Suppress within a class, not across classes.
Forgetting that boxes are in the resized coordinate space. Some pipelines return boxes for the resized input rather than the original image. torchvision's detection models map them back for you. Many other codebases do not. Print a box and check it against the original width and height before you trust it.
Judging a detector on one photograph. This page uses one image to teach the shape of the output. It tells you nothing about whether the model works. Twenty images from your real setting tell you more than any benchmark number.
Assuming a bigger model fixes everything. Sometimes it does. A stronger detector on this same photograph reports the tie once, above 0.5, with a single box. You will see that happen in the YOLO lesson. But a bigger backbone costs memory and latency, and it does not remove the need to choose a threshold. Measure both.
Try it yourself
Change the loop cutoff from 0.05 to 0.02 and count what appears. Then write down, before running it, how many boxes you expect at 0.01.
After that, do the more useful experiment. Point decode_image at a photograph from your own phone — a street, a kitchen table, a family gathering. Print everything above 0.3.
Then look at the picture and count what the model missed. That gap, between what you can see and what it reported, is the honest measure of any detector. It is also the thing no benchmark score will ever tell you.
What to learn next
- YOLO — the one-stage family that made detection fast enough for video.
- Image segmentation — from rectangles to exact object outlines.
- Model evaluation — precision, recall and the curves mAP is built from.
Researcher — Mathematics and papers.
The task
Given an image I, produce a set of predictions of unknown size:
D = { (b_i, c_i, s_i) } for i = 1..Nb_i— a box, conventionally(x1, y1, x2, y2)in image coordinates, or(cx, cy, w, h)in centre form.c_i— a class label from{1, ..., K}.s_i— a confidence score in[0, 1].N— not fixed, and not known in advance.
That last point is the structural difficulty. Classification maps an image to a fixed-size output. Detection maps an image to a set of varying cardinality, which does not fit the standard supervised template. Every detector architecture is, at bottom, a different answer to "how do I turn set prediction into fixed-size prediction".
Three answers, historically
Two-stage. Propose regions, then classify each. R-CNN (Girshick et al., 2014) used selective search plus a CNN per crop, at roughly 47 seconds per image. Fast R-CNN (2015) shared the convolutional computation and introduced RoI pooling. Faster R-CNN (Ren et al., 2015) replaced selective search with a learned Region Proposal Network, making the whole system end-to-end and roughly 200 times faster than R-CNN.
One-stage. Skip proposals; predict boxes and classes directly from a dense grid of anchors. SSD (Liu et al., 2016) and YOLO (Redmon et al., 2016) established the family. The accuracy gap against two-stage methods was diagnosed by Lin et al. (2017) as a foreground-background imbalance problem — roughly 10^4 to 10^5 easy negatives per positive — and addressed with focal loss:
FL(p_t) = -alpha_t * (1 - p_t)^gamma * log(p_t)p_t— predicted probability of the true class.gamma— focusing parameter, 2 in the paper. It down-weights already-easy examples.alpha_t— per-class balancing weight, 0.25 for the foreground in RetinaNet.
Set prediction. DETR (Carion et al., 2020) removed anchors and NMS entirely. A transformer decoder emits a fixed set of N queries, and training uses bipartite matching between predictions and ground truth, solved by the Hungarian algorithm. The matching makes the loss permutation-invariant, so duplicate suppression becomes part of the objective rather than a post-process.
DETR's original weakness was convergence — 500 epochs against Faster R-CNN's 12 — and poor small-object performance. Deformable DETR (Zhu et al., 2021) fixed both with sparse attention over multi-scale features. DINO (Zhang et al., 2022) added denoising queries and now anchors the top of the COCO leaderboard.
Anchors and box parameterisation
Anchor-based detectors regress an offset from a prior box rather than absolute coordinates. The R-CNN parameterisation is still the standard:
t_x = (x - x_a) / w_a t_w = log(w / w_a)
t_y = (y - y_a) / h_a t_h = log(h / h_a)(x, y, w, h)— the target box, centre form.(x_a, y_a, w_a, h_a)— the anchor.t_*— the regressed targets.
Two design choices are doing real work here. Dividing the centre offsets by anchor size makes the target scale-invariant. Taking the log of the size ratio makes the target symmetric — halving and doubling are equal-magnitude errors — and keeps predicted widths positive without a constraint.
Anchor-free detectors (FCOS, Tian et al., 2019; CenterNet, Zhou et al., 2019) predict distances to the four box edges from each location instead. That removes the anchor hyperparameters which dominated tuning effort in the anchor era: scales, aspect ratios and assignment IoU thresholds.
Box regression losses
Smooth L1 on the t_* targets was the original choice. It optimises four numbers independently, which does not correspond to the IoU that evaluation actually uses.
| Loss | Adds | Reference |
|---|---|---|
| IoU loss | 1 - IoU, directly | Yu et al., 2016 |
| GIoU | penalty from the smallest enclosing box; non-zero gradient when boxes are disjoint | Rezatofighi et al., CVPR 2019 |
| DIoU | normalised centre distance term; faster convergence | Zheng et al., AAAI 2020 |
| CIoU | DIoU plus an aspect-ratio consistency term | Zheng et al., AAAI 2020 |
The gradient problem GIoU solves is worth stating. When two boxes do not overlap, IoU is 0 and its gradient is 0 everywhere. Plain IoU loss therefore gives no learning signal in exactly the case that needs one.
NMS and its variants
Greedy NMS sorts by score and removes any box whose IoU with a kept box exceeds N_t:
s_i = 0 if IoU(M, b_i) >= N_tM— the currently highest-scoring box.N_t— the suppression threshold, typically 0.5 to 0.7.
The hard cut is the flaw. A genuinely distinct object overlapping a detected one is deleted outright. Soft-NMS (Bodla et al., ICCV 2017) decays the score instead of zeroing it:
s_i = s_i * exp( -IoU(M, b_i)^2 / sigma )sigma— decay width, 0.5 in the paper.
This gives a consistent gain of roughly 1 point of AP on COCO with no retraining, and it particularly helps crowded scenes. Crowd-specific alternatives exist — CrowdDet (Chu et al., 2020) predicts a set per proposal — and the general problem of dense overlapping instances remains open.
Cost: greedy NMS is O(n^2) in the worst case over n surviving boxes, though n is small after score thresholding. It is also sequential, which is why it shows up as a latency floor in optimised inference graphs and why NMS-free detectors are attractive for deployment.
Evaluation: what mAP is, and is not
Average precision for one class is the area under the precision-recall curve, computed by ranking all detections by score and sweeping the threshold. A detection counts as a true positive if its IoU with an unmatched ground-truth box of the same class exceeds a threshold. Each ground-truth box can be matched at most once; extra matches are false positives.
COCO mAP — reported as AP or mAP@[.5:.95] — averages AP over ten IoU thresholds from 0.50 to 0.95 in steps of 0.05, and over all 80 categories. The precision-recall curve is interpolated at 101 recall points.
Three things follow, and each is routinely misread:
- It is not accuracy. A COCO AP of 21.3, as in the model above, is a respectable score for a 3.4M-parameter mobile detector. State-of-the-art large models sit around 60 to 65. Nothing about these numbers maps onto "percent correct".
- Averaging over IoU thresholds means localisation quality dominates. A model that finds everything but boxes it loosely scores far below its
AP@0.5. The gap betweenAP@0.5andAP@[.5:.95]is a direct readout of box tightness. - The size breakdown is where the information is. COCO reports
AP_small,AP_medium,AP_largefor objects under32^2, between32^2and96^2, and above96^2pixels.AP_smallis typically less than half ofAP_largefor the same model. If your application is small objects, the headline number is actively misleading.
Cost
For a single-stage detector at input H x W, the dense head evaluates A anchors per location over the feature pyramid:
total anchors ≈ A * sum over levels l of (H / s_l) * (W / s_l)s_l— the stride of pyramid levell, typically 8, 16, 32, 64, 128.A— anchors per location, 9 in RetinaNet.
At 640 x 640 with RetinaNet's configuration this is roughly 76,000 anchors, essentially all negative. That imbalance is the reason focal loss exists.
Backbone FLOPs dominate compute; NMS dominates the non-parallelisable tail. On mobile targets the practical lever is input resolution, which scales cost quadratically and small-object AP roughly linearly.
Papers
- Girshick, R. et al. (2014). Rich feature hierarchies for accurate object detection and semantic segmentation. arXiv:1311.2524
- Ren, S. et al. (2015). Faster R-CNN. arXiv:1506.01497
- Liu, W. et al. (2016). SSD: Single Shot MultiBox Detector. arXiv:1512.02325
- Lin, T.-Y. et al. (2017). Feature Pyramid Networks for Object Detection. arXiv:1612.03144
- Lin, T.-Y. et al. (2017). Focal Loss for Dense Object Detection. arXiv:1708.02002
- Bodla, N. et al. (2017). Soft-NMS. arXiv:1704.04503
- Tian, Z. et al. (2019). FCOS: Fully Convolutional One-Stage Object Detection. arXiv:1904.01355
- Rezatofighi, H. et al. (2019). Generalized Intersection over Union. arXiv:1902.09630
- Carion, N. et al. (2020). End-to-End Object Detection with Transformers. arXiv:2005.12872
- Zhu, X. et al. (2021). Deformable DETR. arXiv:2010.04159
- Zhang, H. et al. (2022). DINO: DETR with Improved DeNoising Anchor Boxes. arXiv:2203.03605
- Liu, S. et al. (2023). Grounding DINO. arXiv:2303.05499
- Howard, A. et al. (2019). Searching for MobileNetV3. arXiv:1905.02244 — the source of the SSDLite model used above.
Open problems
Small objects. The gap between AP_small and AP_large has narrowed but not closed, and it is the binding constraint in aerial imagery, retail shelf analysis and long-range driving perception. Higher input resolution is the reliable fix and it is expensive.
Crowds and heavy occlusion. Greedy NMS cannot distinguish duplicate detections from genuinely overlapping instances, and set-prediction methods still degrade sharply in dense scenes.
Long-tail categories. LVIS (Gupta et al., 2019) exposed how badly COCO-tuned methods handle rare classes with a handful of examples. Copy-paste augmentation and decoupled classifier training help; the problem is not solved.
Open vocabulary. Grounding DINO, GLIP and OWL-ViT detect classes named in text at inference time, removing the fixed-K constraint. Accuracy trails a fine-tuned closed-set detector on any specific category, and the flexibility frequently wins in practice.
Calibration. Detection confidences are poorly calibrated, and unlike classification the literature on fixing it is thin. Any system that gates an action on a score threshold needs its own reliability measurement on its own data.
What to learn next
- YOLO — the one-stage family that made detection fast enough for video.
- Image segmentation — from rectangles to exact object outlines.
- Model evaluation — precision, recall and the curves mAP is built from.