Open-vocabulary detection
Open-vocabulary detectors replace the fixed list of class names with text embeddings, so you can ask for a class the model was never trained on by typing its name.
- 16 min read
- 3 reading levels
- Updated
Read these first
On this page 9
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
An open-vocabulary detector finds things you name in words, including things it was never trained on.
The analogy
Think about a new stock clerk on their first day. Hand them a printed list of twenty product names and they can find those twenty. Ask for anything not on the list and they stare at you.
Now imagine a clerk who understands language. Ask for "the blue bottle with the broken cap" and they work it out. They understand the words, rather than matching a fixed list.
An ordinary detector is the first clerk. An open-vocabulary detector is the second.
Why it exists
An ordinary detector has its class list frozen into its last layer. Eighty classes means eighty slots. Adding "cricket bat" means collecting labelled pictures of cricket bats, retraining, and testing again.
That is weeks of work for one class. For a shop with thousands of products, or a factory with a new defect type every month, it never finishes.
The fix came from a different corner of the field. Models had been trained to place pictures and their captions near each other in the same number space. Suppose a picture patch and the words describing it end up close together. You can then score a box against a word, rather than against a fixed slot.
How it works
"auto-rickshaw" -> text encoder -> a list of numbers
|
region of picture -> vision model -> a list of numbers
|
how close are they? -> the scoreThe last layer stops being a fixed table of classes. It becomes a comparison against whatever words you supply, right at the moment you ask.
Change the words, change what the detector looks for. Nothing is retrained.
What it is genuinely good at
Trying an idea in an afternoon instead of a month. You can test whether "cracked tile" is even findable before commissioning a labelling project.
Long lists of rare things. A model can carry thousands of names at once without a separate detector for each.
Labelling help. Point it at unlabelled pictures, get rough boxes, and have people correct them. That is far faster than drawing from scratch.
What it is honestly bad at
Close names. Ask for "cup" and "mug" and it will struggle, because the words themselves sit close together.
Things that are none of your names. Ask for five classes when the picture contains a sixth thing. The model still has to give an answer. It often gives a confident wrong one.
Wording sensitivity. "Dog" and "a photo of a dog" produce different results. The prompt is now part of your system, and it needs testing like code.
Accuracy on a fixed set. If your classes never change, a normal detector trained on your own data will beat it. Open vocabulary buys flexibility, and pays for it in precision.
Where you have seen it
- Photo search that finds "a red umbrella" without those tags existing.
- Labelling tools that pre-fill boxes for a human to correct.
- Robots asked to pick up a named object.
- Stock tools handling new products without retraining.
Remember this
- The fixed class list is replaced by text you supply at the moment of asking.
- New classes cost a sentence instead of a labelling project.
- Close names, unlisted objects and prompt wording are its real, unsolved weaknesses.
What to learn next
- CLIP — the shared picture-and-text space this depends on.
- Vision-language models — the wider family.
- Detecting very small objects — a failure mode these models inherit.
Developer — Code and libraries.
Setup
pip install torchThe first script needs only torch 2.5.1. The real-model section further down needs transformers, and loading a checkpoint downloads weights.
The mechanism, in numbers small enough to read
import torch
import torch.nn.functional as F
# A closed-vocabulary head is a matrix with one ROW PER CLASS, fixed at training time.
closed = torch.nn.Linear(256, 80)
print("closed head weight shape:", tuple(closed.weight.shape))
print("adding one class means changing that shape and retraining the layer\n")
# An open-vocabulary head replaces the matrix with TEXT EMBEDDINGS you can swap.
# Hand-made 4-dimensional embeddings so the numbers stay readable.
# dimensions: [has-wheels, is-animal, is-furry, is-metal]
text = {
"car": torch.tensor([1.0, 0.0, 0.0, 0.9]),
"bicycle": torch.tensor([1.0, 0.0, 0.0, 0.7]),
"dog": torch.tensor([0.0, 1.0, 1.0, 0.0]),
"cat": torch.tensor([0.0, 1.0, 1.0, 0.0]),
}
def cosine(a, b):
return F.cosine_similarity(a.unsqueeze(0), b.unsqueeze(0)).item()
def score(region, table, temp=0.07):
names = list(table)
sims = torch.tensor([cosine(region, table[n]) for n in names])
return names, sims, F.softmax(sims / temp, dim=0)
region = torch.tensor([0.9, 0.1, 0.0, 0.8]) # what the detector saw in one box
names, sims, probs = score(region, text)
print("one region embedding scored against four class names:")
for n, s, p in zip(names, sims, probs):
print(f" {n:<9} similarity {s:6.3f} score {p:6.3f}")
# Adding a class at inference time costs one more row. No retraining.
text["auto-rickshaw"] = torch.tensor([1.0, 0.0, 0.0, 0.6])
names, sims, probs = score(region, text)
print("\nafter adding 'auto-rickshaw' with no training at all:")
for n, s, p in zip(names, sims, probs):
print(f" {n:<14} similarity {s:6.3f} score {p:6.3f}")
print(f"\ncosine('dog', 'cat') = {cosine(text['dog'], text['cat']):.3f}")
print(f"cosine('car', 'auto-rickshaw') = {cosine(text['car'], text['auto-rickshaw']):.3f}")
print("names this close cannot be told apart by the text head alone")
# The failure that matters most: a region that is none of your classes.
pole = torch.tensor([0.0, 0.0, 0.0, 1.0]) # a metal pole. Not in the vocabulary.
names, sims, probs = score(pole, text)
print("\na metal pole, which is none of the five class names:")
for n, s, p in zip(names, sims, probs):
print(f" {n:<14} similarity {s:6.3f} softmax {p:6.3f}")
print("softmax must sum to one, so SOMETHING gets a confident score")
sig = torch.sigmoid(sims / 0.07 - 8)
print("independent sigmoid instead of softmax:",
{n: round(v, 3) for n, v in zip(names, sig.tolist())})closed head weight shape: (80, 256)
adding one class means changing that shape and retraining the layer
one region embedding scored against four class names:
car similarity 0.997 score 0.524
bicycle similarity 0.990 score 0.476
dog similarity 0.059 score 0.000
cat similarity 0.059 score 0.000
after adding 'auto-rickshaw' with no training at all:
car similarity 0.997 score 0.372
bicycle similarity 0.990 score 0.338
dog similarity 0.059 score 0.000
cat similarity 0.059 score 0.000
auto-rickshaw similarity 0.979 score 0.291
cosine('dog', 'cat') = 1.000
cosine('car', 'auto-rickshaw') = 0.982
names this close cannot be told apart by the text head alone
a metal pole, which is none of the five class names:
car similarity 0.669 softmax 0.732
bicycle similarity 0.573 softmax 0.187
dog similarity 0.000 softmax 0.000
cat similarity 0.000 softmax 0.000
auto-rickshaw similarity 0.514 softmax 0.081
softmax must sum to one, so SOMETHING gets a confident score
independent sigmoid instead of softmax: {'car': 0.826, 'bicycle': 0.548, 'dog': 0.0, 'cat': 0.0, 'auto-rickshaw': 0.343}Reading that output
Adding a class changed only the number of rows compared against. The vision side was untouched. This is the entire architectural difference: Linear(256, 80) becomes a dot product against a matrix you assemble at request time.
Adding 'auto-rickshaw' moved every other score. Car fell from 0.524 to 0.372. Under a softmax the scores compete. Adding a name to your prompt list changes the confidence of names already there. If you have tuned a score threshold for five classes, adding a sixth invalidates it.
Dog and cat have cosine similarity 1.000 in this toy space. They are indistinguishable, because the four hand-made dimensions carry no feature that separates them. Real CLIP embeddings are 512-dimensional and do separate them. The same failure still appears for genuinely close names, such as "cup" and "mug", or "bolt" and "screw".
The metal pole is the most important block. It is none of the five classes. Under softmax it gets a 0.732 score for "car". This is not a bug in the demo; it is what a softmax over a closed set does when the true answer is outside the set.
The sigmoid row is the practical fix. Independent sigmoids do not have to sum to one, so all scores can be low together. Real open-vocabulary detectors including OWL-ViT and YOLO-World use per-query sigmoid scoring for this reason. Softmax over prompts is a common home-made mistake.
Running a real open-vocabulary model
Every model below downloads weights on first use, so no output block is shown here rather than an invented one. All of them run on CPU, slowly.
# pip install transformers pillow
# google/owlv2-base-patch16-ensemble downloads roughly 600 MB.
import torch
from PIL import Image
from transformers import Owlv2Processor, Owlv2ForObjectDetection
processor = Owlv2Processor.from_pretrained("google/owlv2-base-patch16-ensemble")
model = Owlv2ForObjectDetection.from_pretrained("google/owlv2-base-patch16-ensemble")
model.eval()
image = Image.open("street.jpg").convert("RGB")
texts = [["a photo of an auto-rickshaw", "a photo of a helmet", "a photo of a stray dog"]]
inputs = processor(text=texts, images=image, return_tensors="pt")
with torch.no_grad():
outputs = model(**inputs)
results = processor.post_process_grounded_object_detection(
outputs, threshold=0.2, target_sizes=torch.tensor([image.size[::-1]])
)[0]
for score, label, box in zip(results["scores"], results["labels"], results["boxes"]):
print(f"{texts[0][label]:<32} {score:.3f} {[round(v, 1) for v in box.tolist()]}")Two API notes, both checked against transformers 5.6.2 on the machine this lesson was written on. The method is post_process_grounded_object_detection, not post_process_object_detection, which is the DETR name. Changing a class list means changing texts and calling again; there is no retraining step anywhere.
This part of the ecosystem moves. Check the method names against your installed version before copying any example, including this one. Use [m for m in dir(Owlv2Processor) if "post" in m].
The families, and what to reach for
| Model | Year | How it connects text to boxes | Reach for it when |
|---|---|---|---|
| ViLD | 2021 | distils CLIP into a Mask R-CNN classifier | historical interest |
| GLIP | 2022 | reframes detection as phrase grounding | you have caption-style data |
| OWL-ViT | 2022 | ViT with per-patch box and text heads | simple, well-supported baseline |
| OWLv2 | 2023 | OWL-ViT scaled with self-training | best of the OWL line for rare classes |
| Grounding DINO | 2023 | DETR plus deep text fusion in the encoder | free-form phrases, strong quality |
| YOLO-World | 2024 | YOLO backbone with a text-guided neck | real-time with a fixed prompt set |
| YOLOE | 2025 | text, visual or prompt-free with a built-in vocabulary | real-time, and you may not know the names |
Two numbers that give the shape of the field. OWLv2's self-training raised LVIS rare-class AP from 31.2 to 44.6 with an L/14 backbone. Those are classes for which it saw no human box annotations. Ultralytics report YOLOE-v8s beating YOLO-Worldv2-S by 3.5 AP on LVIS, at a third of the training cost. A 4,585-name built-in vocabulary supports prompt-free use.
Take both as vendor and author claims measured on their benchmarks. The number that matters is the one you measure on your own images.
Prompt engineering, which is now part of your system
The text encoder is a language model, and wording changes the embedding.
- Prefix templates help. "a photo of a {class}" was CLIP's finding and it carries over. Averaging embeddings over several templates is a cheap gain.
- Be specific about the visual thing, not the concept. "a yellow auto-rickshaw" beats "public transport".
- Add distractor names on purpose. If your model keeps calling poles "cars", add "pole" to your prompt list so the pole has somewhere to go. This is the practical fix for the failure shown above.
- Version your prompts. They are configuration that changes model behaviour. Put them in a file, not in a notebook cell.
Common mistakes
Softmax over your prompt list. Use independent sigmoids. Otherwise the scores of your existing classes shift every time you add a prompt.
No background name. Without something for unlisted objects to match, they are forced onto your real classes.
Reusing a threshold across prompt sets. Scores are not calibrated and not comparable between prompt lists. Re-tune when the list changes.
Expecting closed-set accuracy. On a fixed class list with real training data, a fine-tuned closed detector wins. Use open vocabulary for flexibility, prototyping and labelling, and fine-tune for production if the list ever settles.
Not measuring. Label 200 images, then measure. Open-vocabulary demos look excellent and hide their failure rate.
Try it yourself
Add "metal pole": torch.tensor([0.0, 0.0, 0.0, 1.0]) to the text table and re-score the pole. It stops being a car. That single change is the whole practical technique for handling out-of-vocabulary objects: give them a name to land on.
What to learn next
- CLIP — the shared picture-and-text space this depends on.
- Vision-language models — the wider family.
- Detecting very small objects — a failure mode these models inherit.
Researcher — Mathematics and papers.
The problem statement
Closed-set detection learns $p(c \mid \text{region})$ over a fixed $\mathcal{C}$ of size $K$, realised as a $K \times d$ weight matrix. Open-vocabulary detection replaces that matrix with $\mathbf{T} \in \mathbb{R}^{|\mathcal{Q}| \times d}$. Each row is a text embedding of a query phrase, produced at inference:
$$ s_{iq} = \frac{\langle \mathbf{v}_i, \mathbf{t}_q \rangle}{\tau \, |\mathbf{v}_i| |\mathbf{t}_q|} $$
Where $\mathbf{v}_i$ is the region embedding for detection $i$, $\mathbf{t}_q$ the embedding of query $q$, and $\tau$ a temperature. Scores pass through a per-query sigmoid, not a softmax over $\mathcal{Q}$. That keeps the query set genuinely open, and assumes no mutually exclusive competition.
The hard part is not the head. It is producing region embeddings that live in the same space as sentence embeddings. Detection data is scarce, while image-text pairs are abundant.
Three strategies for that alignment
Distillation. ViLD (Gu et al., 2022) crops proposals and embeds each with a frozen CLIP image encoder. It then regresses the detector's region embeddings onto those. Bounded above by CLIP's crop-level quality, and slow to train because every proposal needs a CLIP pass.
Grounding as pre-training. GLIP (Li et al., 2022) reformulates detection as phrase grounding: predict the alignment between region features and words in a caption. This unifies detection data and grounding data, and lets caption data teach localisation. Grounding DINO (Liu et al., 2023) extends this with feature fusion at three points. It adds a language-guided query selection module and cross-modality decoder layers. Text now influences the detector rather than only scoring its output.
Self-training. OWLv2 (Minderer et al., 2023) uses an existing detector to pseudo-annotate boxes on web image-text pairs. It then trains on the result.
The paper reports LVIS rare-class AP rising from 31.2 to 44.6 with an L/14 architecture. That is a 43 percent relative improvement on classes with no human box annotations.
They identify three engineering problems: label-space choice, pseudo-annotation filtering, and training efficiency. Those are the levers anyone reproducing this work will spend their time on.
Real-time variants
YOLO-World (Cheng et al., 2024) attaches a re-parameterisable vision-language path aggregation network to a YOLOv8 backbone. Its practical trick is prompt re-parameterisation. Once the prompt set is fixed, the text embeddings fold into the network weights offline. Inference then costs the same as a closed-set YOLO. Open vocabulary at development time, closed at deployment time, which is the right trade for most production systems.
YOLOE extends this to three prompting modes, including a prompt-free mode over a large built-in vocabulary. Ultralytics report YOLOE-L at 52.6 mAP on closed-set COCO against YOLO11-L's 53.5. Latency and parameter count are identical. The open-vocabulary capability costs about one point of closed-set accuracy, not a large fraction.
Evaluation, and its problems
The standard protocol is LVIS with its frequent, common and rare splits, reporting $\mathrm{AP}_r$ on classes excluded from box supervision.
Three caveats deserve stating.
- Leakage through pre-training. The vision-language backbone saw web data. A "rare" LVIS class almost certainly appeared in it with text. This is not zero-shot in the strict sense, and the field's terminology overstates the claim.
- Prompt dependence. Reported numbers depend on the prompt templates used, which are not always published in full.
- Vocabulary size effects. Scores shift with the number of queries evaluated. Comparing a method evaluated over 1,203 LVIS names against one evaluated over 80 COCO names is not a like-for-like comparison.
The honest evaluation is a labelled slice of your own data with your own prompt list. Measure precision and recall at the threshold you will actually ship.
Open problems
- Attribute and relation queries. "The cracked tile, not the whole floor" requires compositional grounding that current models handle unreliably.
- Negation. "A car without a roof rack" is close to unusable in every model in this family.
- Calibration. Scores are not probabilities and are not comparable across prompt sets, which makes thresholding a per-deployment tuning job.
- Fine-grained distinctions. Separating visually similar categories whose names sit close in text space is largely unsolved. That is exactly where industrial applications live.
Papers
- Gu et al., Open-vocabulary Object Detection via Vision and Language Knowledge Distillation (ViLD), ICLR 2022 — arxiv.org/abs/2104.13921
- Li et al., Grounded Language-Image Pre-training (GLIP), CVPR 2022 — arxiv.org/abs/2112.03857
- Minderer et al., Simple Open-Vocabulary Object Detection with Vision Transformers (OWL-ViT), ECCV 2022 — arxiv.org/abs/2205.06230
- Minderer et al., Scaling Open-Vocabulary Object Detection (OWLv2), NeurIPS 2023 — arxiv.org/abs/2306.09683
- Liu et al., Grounding DINO, 2023 — arxiv.org/abs/2303.05499
- Cheng et al., YOLO-World: Real-Time Open-Vocabulary Object Detection, CVPR 2024 — arxiv.org/abs/2401.17270
What to learn next
- CLIP — the shared picture-and-text space this depends on.
- Vision-language models — the wider family.
- Detecting very small objects — a failure mode these models inherit.