YOLO versions compared
A decade of YOLO releases from three different research groups, what each one actually changed in the head and the loss, and how to pick a version without believing a marketing table.
- 15 min read
- 3 reading levels
- Updated
Read these first
On this page 9
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
YOLO is a family of fast detectors. The version number tells you which group released it and when. It does not tell you how good it is for you.
The analogy
Think about a scooter model that gets a new badge every year. Some years the engine changes. Some years only the sticker changes. Some years a different company buys the name and puts it on a different scooter.
You would not pick one on the badge number alone. You would ask what changed, and whether the change helps the roads you ride on.
YOLO version numbers work the same way. The number is a release label, not a ranking.
Why there are so many
The first YOLO came from one researcher in 2016 and had a memorable idea. Look at the whole picture once. Split it into a grid, and let every cell answer for what it sees. No separate proposal step, no running a network thousands of times.
That idea was fast enough to run on video, and the name became famous. Since then the name has passed between several groups. Those include one company with a popular open-source package, and several university labs.
They do not share a codebase or a numbering plan. That is why the versions are not a straight line.
How the family changed, in plain terms
The early versions used ready-made box shapes. The model picked one and adjusted it. Picking the shapes was a tuning job.
Later versions dropped the ready-made shapes. Each grid cell says how far the object edge is in each of four directions. Nothing has to be tuned in advance.
The newest versions removed the clean-up step too. Older versions produce many overlapping boxes and delete duplicates afterwards. The newest ones are trained so only one box fires per object, which makes them faster and steadier on video.
How to actually pick one
Do you need it on a phone or a small board? -> smallest size of a recent version
Do you need commercial licensing certainty? -> read the licence FIRST
Do you have a strange dataset? -> test two versions on YOUR data
Do you already have a working pipeline? -> keep it until it stops workingOn your own data, the difference between recent versions is usually small. Better labels usually matter more. That sentence will save you more time than any comparison table.
The licence, which people forget
Some YOLO releases are under a licence that requires you to open-source your own software if you distribute it. Others are more permissive. This has ended real projects late in development.
Check the licence of the exact version and weights you plan to ship, before you build anything on it.
Where you have seen these
- Shop and warehouse cameras counting stock.
- Traffic systems reading vehicles and helmets.
- Farm drones counting plants and spotting disease.
- Sports tools tracking players through a match.
Remember this
- YOLO is a family name used by several different groups, so version numbers are not a ranking.
- The big real changes were dropping ready-made box shapes, then dropping the duplicate clean-up step.
- Pick by your constraints and your data, and read the licence before you build.
What to learn next
- Anchor-free detection — the head design every modern YOLO uses.
- Focal loss and RetinaNet — the loss family DFL grew out of.
- YOLO — running one end to end on a real image.
Developer — Code and libraries.
Setup
pip install ultralytics torchThe snippets below were run against ultralytics 8.3.40 and torch 2.5.1. That package version ships configs up to YOLO11 and YOLOv10. YOLO26 and YOLOv12 need a newer ultralytics release. The inspection below covers the models this pinned version can build. Everything runs on CPU and nothing is downloaded.
What each generation's head actually outputs
import torch
import torch.nn.functional as F
from torchvision.ops import nms
# --- 1. How many candidate boxes each generation of head produces at 640x640.
old = sum(s * s * 3 * (5 + 80) for s in (80, 40, 20)) # v3/v5: 3 anchors, 4 box + 1 objectness + 80 classes
new = sum(s * s * (4 * 16 + 80) for s in (80, 40, 20)) # v8 onward: no anchors, 16-bin box sides
print(f"anchor-based head (v3/v5 style): {old:>9,} numbers, {(80*80+40*40+20*20)*3:>6,} candidate boxes")
print(f"anchor-free head (v8 style) : {new:>9,} numbers, {80*80+40*40+20*20:>6,} candidate boxes")
# --- 2. One box side, predicted as a 16-bin distribution, then integrated (DFL).
logits = torch.tensor([0., 0., 0., 0., 0., 1., 4., 6., 4., 1., 0., 0., 0., 0., 0., 0.])
p = F.softmax(logits, dim=0)
bins = torch.arange(16, dtype=torch.float32)
distance = (p * bins).sum()
print("\none side of one box as a 16-bin distribution:")
print(" probabilities:", [round(v, 3) for v in p.tolist()])
print(f" peak bin {p.argmax().item()}, expected value {distance:.3f} cells")
print(f" at stride 8 that is {distance * 8:.1f} pixels from the cell centre")
# --- 3. Anchor-free decode: four distances from the cell centre to the four edges.
cx, cy, stride = 12.5, 7.5, 8.0
left, top, right, bottom = 2.0, 3.5, 4.0, 6.5
box = torch.tensor([(cx - left) * stride, (cy - top) * stride,
(cx + right) * stride, (cy + bottom) * stride])
print("\ndecoded box in image pixels:", [round(v, 1) for v in box.tolist()])
# --- 4. Why v10 and YOLO26 can drop NMS: one-to-one assignment during training.
boxes = torch.tensor([[100., 100., 200., 200.],
[ 98., 104., 205., 198.],
[104., 102., 196., 203.],
[ 96., 99., 202., 201.]])
one_to_many = torch.tensor([0.90, 0.86, 0.83, 0.79])
one_to_one = torch.tensor([0.91, 0.04, 0.03, 0.02])
print(f"\none-to-many head, above 0.25 before NMS: {(one_to_many > 0.25).sum().item()} boxes")
print(f"one-to-many head, after NMS : {nms(boxes, one_to_many, 0.5).numel()} boxes")
print(f"one-to-one head, above 0.25, no NMS : {(one_to_one > 0.25).sum().item()} boxes")anchor-based head (v3/v5 style): 2,142,000 numbers, 25,200 candidate boxes anchor-free head (v8 style) : 1,209,600 numbers, 8,400 candidate boxes one side of one box as a 16-bin distribution: probabilities: [0.002, 0.002, 0.002, 0.002, 0.002, 0.005, 0.103, 0.763, 0.103, 0.005, 0.002, 0.002, 0.002, 0.002, 0.002, 0.002] peak bin 7, expected value 7.015 cells at stride 8 that is 56.1 pixels from the cell centre decoded box in image pixels: [84.0, 32.0, 132.0, 112.0] one-to-many head, above 0.25 before NMS: 4 boxes one-to-many head, after NMS : 1 boxes one-to-one head, above 0.25, no NMS : 1 boxes
Reading that output
Dropping anchors cut candidate boxes from 25,200 to 8,400. Three anchors per cell became one prediction per cell. Fewer candidates means a cheaper NMS step and fewer negatives to drown the loss.
The 16-bin distribution is the piece most people have never looked at. Since YOLOv8, each of the four box sides is a distribution over 16 discrete distances. It is not a single number. The final value is the expected value of that distribution, 7.015 here from a peak at bin 7.
This is Distribution Focal Loss, from the Generalized Focal Loss paper. It buys two things. A blurry or occluded edge produces a wide distribution rather than a confidently wrong number. The softmax also gives sub-cell precision, with no separate refinement stage. It costs 64 output channels per cell instead of 4.
YOLO26 removes DFL again. The argument is that the extra channels cost more at export time than they return, especially on integer-quantised edge hardware.
The one-to-one comparison is the end-to-end story. Four overlapping boxes above threshold need NMS to become one. A one-to-one trained head puts almost all its confidence on a single cell, so the output is already clean. The saving is not only average latency; it is that latency stops depending on how crowded the frame is.
Inspecting the real architectures
import ultralytics
from ultralytics import YOLO
print("ultralytics", ultralytics.__version__)
for cfg in ["yolov8n.yaml", "yolov9t.yaml", "yolov10n.yaml", "yolo11n.yaml"]:
m = YOLO(cfg) # builds from config; no weights are downloaded
head = m.model.model[-1]
n = sum(p.numel() for p in m.model.parameters())
one2one = hasattr(head, "one2one_cv2")
print(f"{cfg:<14} head={type(head).__name__:<10} params={n:>9,}"
f" reg_max={head.reg_max} end2end={head.end2end} one-to-one branch={one2one}")ultralytics 8.3.40 yolov8n.yaml head=Detect params=3,157,200 reg_max=16 end2end=False one-to-one branch=False yolov9t.yaml head=Detect params=2,128,720 reg_max=16 end2end=False one-to-one branch=False yolov10n.yaml head=v10Detect params=2,775,520 reg_max=16 end2end=True one-to-one branch=True yolo11n.yaml head=Detect params=2,624,080 reg_max=16 end2end=False one-to-one branch=False
Three real facts fall out of that, none of which you would get from a benchmark table.
reg_max=16 is identical across all four. The distribution-based box representation has been stable since v8.
Only yolov10n reports end2end=True and carries a one2one_cv2 branch. That extra branch is the architectural difference behind the NMS-free claim. v10 trains two heads. A one-to-many head gives rich gradients, and a one-to-one head is used at inference.
YOLO11n has fewer parameters than YOLOv8n (2.62M against 3.16M). Smaller is a design goal in this family, not an accident.
One warning about this package. ultralytics also ships yolov3.yaml and yolov5.yaml, but they are re-implementations built with the modern anchor-free Detect head. Building them and reporting "YOLOv5 is anchor-free" would be wrong. The original v3 and v5 are anchor-based. Do not use these configs as evidence about the original models.
The family, and what each release actually changed
| Version | Year | From | The change that mattered |
|---|---|---|---|
| YOLOv1 | 2016 | Redmon et al. | One pass, grid cells predict boxes directly |
| YOLOv2 | 2017 | Redmon and Farhadi | Anchors chosen by k-means on the dataset |
| YOLOv3 | 2018 | Redmon and Farhadi | Three scales, per-class sigmoid instead of softmax |
| YOLOv4 | 2020 | Bochkovskiy et al. | A catalogue of training tricks: mosaic, CIoU, CSP backbone |
| YOLOv5 | 2020 | Ultralytics | Engineering: PyTorch, export paths, usable tooling |
| YOLOX | 2021 | Megvii | Anchor-free head, decoupled classification and box branches, SimOTA |
| YOLOv6 | 2022 | Meituan | Reparameterised backbone aimed at industrial deployment |
| YOLOv7 | 2022 | Wang et al. | Trainable bag-of-freebies, extended efficient layer aggregation |
| YOLOv8 | 2023 | Ultralytics | Anchor-free, DFL box representation, task-aligned assignment |
| YOLOv9 | 2024 | Wang and Liao | Programmable gradient information, reversible branches |
| YOLOv10 | 2024 | Tsinghua University | Dual assignment, NMS-free inference |
| YOLO11 | 2024 | Ultralytics | Fewer parameters at similar or better accuracy |
| YOLOv12 | 2025 | Tian et al. | Attention-centric blocks with FlashAttention |
| YOLO26 | 2026 | Ultralytics | End-to-end by default, DFL removed, new optimiser and loss schedule |
Ultralytics documents YOLO26 as released in January 2026. It is natively end-to-end with no NMS, and Distribution Focal Loss is removed. The release also adds a progressive loss schedule, small-target-aware label assignment and a hybrid optimiser. Their published COCO table lists YOLO26n at 40.9 mAP and YOLO26x at 57.5 mAP, measured on a T4 with TensorRT.
Treat every vendor table, including that one, as a claim measured under a specific recipe on one dataset. It is not a prediction about your data.
Choosing a version, honestly
Ask these in order.
What is the licence? Ultralytics YOLO models are AGPL-3.0 with a paid commercial option. YOLOX and RTMDet are Apache-2.0 and MIT respectively. RF-DETR is Apache-2.0. If your product ships to customers, this question outranks accuracy.
What runs on your hardware? An export path that works beats two extra mAP points. Test ONNX or TensorRT export early, not at the end.
Does end-to-end help you? If you process video with tight latency budgets, an NMS-free model gives you steadier frame times. If you run batch jobs, it changes little.
Have you measured on your own data? Train two candidates for a short schedule on your dataset and compare. Published COCO numbers rank models on 80 everyday classes at moderate resolution. That may have nothing to do with cracks on a weld.
Common mistakes
Believing a version number is a ranking. YOLOv9, v10 and v12 came from three different groups. They are not sequential improvements on one codebase.
Comparing at different input sizes. A model at 1280 will beat the same model at 640 on small objects and be four times slower. Fix the resolution before comparing anything.
Comparing PyTorch latency to TensorRT latency. Published speed tables are usually exported and half-precision. Your unexported PyTorch loop is not the same measurement.
Ignoring the licence until launch. This is the single most expensive mistake in this list.
Skipping the export test. Some architectures export cleanly to ONNX, some need work. Find out on day one.
Try it yourself
Build yolov10n.yaml and yolo11n.yaml, and print head.one2one_cv2 for the one that has it. Then count how many output tensors each head returns in eval mode. That difference is the whole NMS-free mechanism, visible in ten lines of code.
What to learn next
- Anchor-free detection — the head design every modern YOLO uses.
- Focal loss and RetinaNet — the loss family DFL grew out of.
- YOLO — running one end to end on a real image.
Researcher — Mathematics and papers.
What actually changed, mechanism by mechanism
The YOLO line is best read as four independent threads rather than one sequence.
Assignment. v1 assigned one grid cell per object. v2 to v5 assigned by anchor IoU. YOLOX introduced SimOTA, a top-$k$ approximation of the optimal transport assignment of Ge et al. (2021). v8 onward uses task-aligned assignment (Feng et al., 2021, TOOD), scoring candidates by
$$ t = s^{\alpha} \cdot u^{\beta} $$
Where $s$ is the classification score for the target class, $u$ the predicted-to-target IoU, and $\alpha, \beta$ control the balance. Candidates are ranked by $t$ and the top $k$ per object become positives. The point is that classification confidence and localisation quality are selected jointly. The head does not learn to be confident about badly-placed boxes.
Box representation. Single scalars, then log-space anchor offsets, then $(l,t,r,b)$ distances. Then the discrete distribution of Li et al. (2020), Generalized Focal Loss:
$$ \hat{y} = \sum_{i=0}^{n} P(y_i) \, y_i $$
Where $P$ is a softmax over $n+1$ bins covering the plausible distance range. DFL supervises the two bins adjacent to the continuous target:
$$ \mathrm{DFL}(S_i, S_{i+1}) = -\big((y_{i+1} - y)\log S_i + (y - y_i)\log S_{i+1}\big) $$
This keeps the predicted distribution sharp around the true value, and still a proper probability distribution. That is what makes the expectation sub-cell accurate.
Duplicate removal. v1 to YOLO11 use NMS. YOLOv10 (Wang et al., 2024) trains two heads from a shared backbone, with consistent matching metrics. One is one-to-many and one is one-to-one. The one-to-many head is discarded at inference. The one-to-many branch supplies the dense supervision that makes DETR-style training slow to converge; the one-to-one branch supplies the clean output.
Backbone and neck. CSP connections (v4, v5), reparameterisable blocks (v6, v7), programmable gradient information (v9), attention blocks with area attention and FlashAttention (v12). These deliver smaller gains than the assignment and representation changes. They are also the parts most sensitive to the training recipe.
The benchmarking problem
Published YOLO comparisons are unusually hard to trust, for reasons that are structural rather than dishonest.
- Recipe coupling. Epoch count, mosaic schedule, EMA, and augmentation differ between releases. Liu et al. (2022) studied CNNs against transformers. Recipe differences accounted for most of the claimed architectural gap. The same caution applies here.
- Latency measurement. With or without NMS, with or without pre-processing, at what batch size, in what precision, on what hardware. A table without all five is not comparable.
- Parameter count is not latency. Attention blocks and depthwise convolutions are memory-bandwidth-bound, so FLOPs and parameters correlate poorly with wall-clock time on edge devices.
The defensible comparison is one you run yourself. Same data, same resolution, same augmentation budget, same export path. Measure latency end to end, including pre- and post-processing.
Where the frontier sits
As of mid-2026 the real-time frontier is contested between the YOLO line and the real-time DETR line.
- Ultralytics reports YOLO26x at 57.5 COCO mAP with T4 TensorRT latency of 11.8 ms.
- RT-DETRv4 (arXiv 2510.25257, ECCV 2026) reports 55.4 AP at 124 FPS for its L variant, and 57.0 AP for X. The baselines are DEIM-L at 54.7 and YOLOv13-L at 53.4. Its contribution is a training-only distillation from a frozen DINOv3 model, so inference cost is unchanged.
- RF-DETR is reported as the first real-time detector past 60 mAP on COCO, and is Apache-2.0 licensed.
Two things follow. First, the architectural gap between the families is now small. Licence, tooling and export path decide most real deployments. Second, the interesting recent gains come from training-time supervision (foundation-model distillation, better matching) rather than from inference-time architecture.
Papers
- Redmon et al., You Only Look Once: Unified, Real-Time Object Detection, CVPR 2016 — arxiv.org/abs/1506.02640
- Ge et al., YOLOX: Exceeding YOLO Series in 2021 — arxiv.org/abs/2107.08430
- Li et al., Generalized Focal Loss, NeurIPS 2020 — arxiv.org/abs/2006.04388
- Wang et al., YOLOv10: Real-Time End-to-End Object Detection, NeurIPS 2024 — arxiv.org/abs/2405.14458
- RT-DETRv4: Painlessly Furthering Real-Time Object Detection with Vision Foundation Models, ECCV 2026 — arxiv.org/abs/2510.25257
- Ultralytics YOLO26 documentation — docs.ultralytics.com/models/yolo26
What to learn next
- Anchor-free detection — the head design every modern YOLO uses.
- Focal loss and RetinaNet — the loss family DFL grew out of.
- YOLO — running one end to end on a real image.