Object Detection in Depth

YOLO versions compared

A decade of YOLO releases from three different research groups, what each one actually changed in the head and the loss, and how to pick a version without believing a marketing table.

Read these first

On this page 9
  1. The short answer
  2. The analogy
  3. Why there are so many
  4. How the family changed, in plain terms
  5. How to actually pick one
  6. The licence, which people forget
  7. Where you have seen these
  8. Remember this
  9. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

YOLO is a family of fast detectors. The version number tells you which group released it and when. It does not tell you how good it is for you.

The analogy

Think about a scooter model that gets a new badge every year. Some years the engine changes. Some years only the sticker changes. Some years a different company buys the name and puts it on a different scooter.

You would not pick one on the badge number alone. You would ask what changed, and whether the change helps the roads you ride on.

YOLO version numbers work the same way. The number is a release label, not a ranking.

Why there are so many

The first YOLO came from one researcher in 2016 and had a memorable idea. Look at the whole picture once. Split it into a grid, and let every cell answer for what it sees. No separate proposal step, no running a network thousands of times.

That idea was fast enough to run on video, and the name became famous. Since then the name has passed between several groups. Those include one company with a popular open-source package, and several university labs.

They do not share a codebase or a numbering plan. That is why the versions are not a straight line.

How the family changed, in plain terms

The early versions used ready-made box shapes. The model picked one and adjusted it. Picking the shapes was a tuning job.

Later versions dropped the ready-made shapes. Each grid cell says how far the object edge is in each of four directions. Nothing has to be tuned in advance.

The newest versions removed the clean-up step too. Older versions produce many overlapping boxes and delete duplicates afterwards. The newest ones are trained so only one box fires per object, which makes them faster and steadier on video.

How to actually pick one

   Do you need it on a phone or a small board?  ->  smallest size of a recent version
   Do you need commercial licensing certainty?  ->  read the licence FIRST
   Do you have a strange dataset?               ->  test two versions on YOUR data
   Do you already have a working pipeline?      ->  keep it until it stops working

On your own data, the difference between recent versions is usually small. Better labels usually matter more. That sentence will save you more time than any comparison table.

The licence, which people forget

Some YOLO releases are under a licence that requires you to open-source your own software if you distribute it. Others are more permissive. This has ended real projects late in development.

Check the licence of the exact version and weights you plan to ship, before you build anything on it.

Where you have seen these

  • Shop and warehouse cameras counting stock.
  • Traffic systems reading vehicles and helmets.
  • Farm drones counting plants and spotting disease.
  • Sports tools tracking players through a match.

Remember this

  • YOLO is a family name used by several different groups, so version numbers are not a ranking.
  • The big real changes were dropping ready-made box shapes, then dropping the duplicate clean-up step.
  • Pick by your constraints and your data, and read the licence before you build.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install ultralytics torch

The snippets below were run against ultralytics 8.3.40 and torch 2.5.1. That package version ships configs up to YOLO11 and YOLOv10. YOLO26 and YOLOv12 need a newer ultralytics release. The inspection below covers the models this pinned version can build. Everything runs on CPU and nothing is downloaded.

What each generation's head actually outputs

head_math.py
import torch
import torch.nn.functional as F
from torchvision.ops import nms

# --- 1. How many candidate boxes each generation of head produces at 640x640.
old = sum(s * s * 3 * (5 + 80) for s in (80, 40, 20))     # v3/v5: 3 anchors, 4 box + 1 objectness + 80 classes
new = sum(s * s * (4 * 16 + 80) for s in (80, 40, 20))    # v8 onward: no anchors, 16-bin box sides
print(f"anchor-based head (v3/v5 style): {old:>9,} numbers, {(80*80+40*40+20*20)*3:>6,} candidate boxes")
print(f"anchor-free head  (v8 style)   : {new:>9,} numbers, {80*80+40*40+20*20:>6,} candidate boxes")

# --- 2. One box side, predicted as a 16-bin distribution, then integrated (DFL).
logits = torch.tensor([0., 0., 0., 0., 0., 1., 4., 6., 4., 1., 0., 0., 0., 0., 0., 0.])
p = F.softmax(logits, dim=0)
bins = torch.arange(16, dtype=torch.float32)
distance = (p * bins).sum()
print("\none side of one box as a 16-bin distribution:")
print("  probabilities:", [round(v, 3) for v in p.tolist()])
print(f"  peak bin {p.argmax().item()}, expected value {distance:.3f} cells")
print(f"  at stride 8 that is {distance * 8:.1f} pixels from the cell centre")

# --- 3. Anchor-free decode: four distances from the cell centre to the four edges.
cx, cy, stride = 12.5, 7.5, 8.0
left, top, right, bottom = 2.0, 3.5, 4.0, 6.5
box = torch.tensor([(cx - left) * stride, (cy - top) * stride,
                    (cx + right) * stride, (cy + bottom) * stride])
print("\ndecoded box in image pixels:", [round(v, 1) for v in box.tolist()])

# --- 4. Why v10 and YOLO26 can drop NMS: one-to-one assignment during training.
boxes = torch.tensor([[100., 100., 200., 200.],
                      [ 98., 104., 205., 198.],
                      [104., 102., 196., 203.],
                      [ 96.,  99., 202., 201.]])
one_to_many = torch.tensor([0.90, 0.86, 0.83, 0.79])
one_to_one = torch.tensor([0.91, 0.04, 0.03, 0.02])
print(f"\none-to-many head, above 0.25 before NMS: {(one_to_many > 0.25).sum().item()} boxes")
print(f"one-to-many head, after NMS            : {nms(boxes, one_to_many, 0.5).numel()} boxes")
print(f"one-to-one head, above 0.25, no NMS    : {(one_to_one > 0.25).sum().item()} boxes")
Output
anchor-based head (v3/v5 style): 2,142,000 numbers, 25,200 candidate boxes
anchor-free head  (v8 style)   : 1,209,600 numbers,  8,400 candidate boxes

one side of one box as a 16-bin distribution:
  probabilities: [0.002, 0.002, 0.002, 0.002, 0.002, 0.005, 0.103, 0.763, 0.103, 0.005, 0.002, 0.002, 0.002, 0.002, 0.002, 0.002]
  peak bin 7, expected value 7.015 cells
  at stride 8 that is 56.1 pixels from the cell centre

decoded box in image pixels: [84.0, 32.0, 132.0, 112.0]

one-to-many head, above 0.25 before NMS: 4 boxes
one-to-many head, after NMS            : 1 boxes
one-to-one head, above 0.25, no NMS    : 1 boxes

Reading that output

Dropping anchors cut candidate boxes from 25,200 to 8,400. Three anchors per cell became one prediction per cell. Fewer candidates means a cheaper NMS step and fewer negatives to drown the loss.

The 16-bin distribution is the piece most people have never looked at. Since YOLOv8, each of the four box sides is a distribution over 16 discrete distances. It is not a single number. The final value is the expected value of that distribution, 7.015 here from a peak at bin 7.

This is Distribution Focal Loss, from the Generalized Focal Loss paper. It buys two things. A blurry or occluded edge produces a wide distribution rather than a confidently wrong number. The softmax also gives sub-cell precision, with no separate refinement stage. It costs 64 output channels per cell instead of 4.

YOLO26 removes DFL again. The argument is that the extra channels cost more at export time than they return, especially on integer-quantised edge hardware.

The one-to-one comparison is the end-to-end story. Four overlapping boxes above threshold need NMS to become one. A one-to-one trained head puts almost all its confidence on a single cell, so the output is already clean. The saving is not only average latency; it is that latency stops depending on how crowded the frame is.

Inspecting the real architectures

yolo_heads.py
import ultralytics
from ultralytics import YOLO

print("ultralytics", ultralytics.__version__)
for cfg in ["yolov8n.yaml", "yolov9t.yaml", "yolov10n.yaml", "yolo11n.yaml"]:
    m = YOLO(cfg)                      # builds from config; no weights are downloaded
    head = m.model.model[-1]
    n = sum(p.numel() for p in m.model.parameters())
    one2one = hasattr(head, "one2one_cv2")
    print(f"{cfg:<14} head={type(head).__name__:<10} params={n:>9,}"
          f"  reg_max={head.reg_max}  end2end={head.end2end}  one-to-one branch={one2one}")
Output
ultralytics 8.3.40
yolov8n.yaml   head=Detect     params=3,157,200  reg_max=16  end2end=False  one-to-one branch=False
yolov9t.yaml   head=Detect     params=2,128,720  reg_max=16  end2end=False  one-to-one branch=False
yolov10n.yaml  head=v10Detect  params=2,775,520  reg_max=16  end2end=True  one-to-one branch=True
yolo11n.yaml   head=Detect     params=2,624,080  reg_max=16  end2end=False  one-to-one branch=False

Three real facts fall out of that, none of which you would get from a benchmark table.

reg_max=16 is identical across all four. The distribution-based box representation has been stable since v8.

Only yolov10n reports end2end=True and carries a one2one_cv2 branch. That extra branch is the architectural difference behind the NMS-free claim. v10 trains two heads. A one-to-many head gives rich gradients, and a one-to-one head is used at inference.

YOLO11n has fewer parameters than YOLOv8n (2.62M against 3.16M). Smaller is a design goal in this family, not an accident.

One warning about this package. ultralytics also ships yolov3.yaml and yolov5.yaml, but they are re-implementations built with the modern anchor-free Detect head. Building them and reporting "YOLOv5 is anchor-free" would be wrong. The original v3 and v5 are anchor-based. Do not use these configs as evidence about the original models.

The family, and what each release actually changed

VersionYearFromThe change that mattered
YOLOv12016Redmon et al.One pass, grid cells predict boxes directly
YOLOv22017Redmon and FarhadiAnchors chosen by k-means on the dataset
YOLOv32018Redmon and FarhadiThree scales, per-class sigmoid instead of softmax
YOLOv42020Bochkovskiy et al.A catalogue of training tricks: mosaic, CIoU, CSP backbone
YOLOv52020UltralyticsEngineering: PyTorch, export paths, usable tooling
YOLOX2021MegviiAnchor-free head, decoupled classification and box branches, SimOTA
YOLOv62022MeituanReparameterised backbone aimed at industrial deployment
YOLOv72022Wang et al.Trainable bag-of-freebies, extended efficient layer aggregation
YOLOv82023UltralyticsAnchor-free, DFL box representation, task-aligned assignment
YOLOv92024Wang and LiaoProgrammable gradient information, reversible branches
YOLOv102024Tsinghua UniversityDual assignment, NMS-free inference
YOLO112024UltralyticsFewer parameters at similar or better accuracy
YOLOv122025Tian et al.Attention-centric blocks with FlashAttention
YOLO262026UltralyticsEnd-to-end by default, DFL removed, new optimiser and loss schedule

Ultralytics documents YOLO26 as released in January 2026. It is natively end-to-end with no NMS, and Distribution Focal Loss is removed. The release also adds a progressive loss schedule, small-target-aware label assignment and a hybrid optimiser. Their published COCO table lists YOLO26n at 40.9 mAP and YOLO26x at 57.5 mAP, measured on a T4 with TensorRT.

Treat every vendor table, including that one, as a claim measured under a specific recipe on one dataset. It is not a prediction about your data.

Choosing a version, honestly

Ask these in order.

What is the licence? Ultralytics YOLO models are AGPL-3.0 with a paid commercial option. YOLOX and RTMDet are Apache-2.0 and MIT respectively. RF-DETR is Apache-2.0. If your product ships to customers, this question outranks accuracy.

What runs on your hardware? An export path that works beats two extra mAP points. Test ONNX or TensorRT export early, not at the end.

Does end-to-end help you? If you process video with tight latency budgets, an NMS-free model gives you steadier frame times. If you run batch jobs, it changes little.

Have you measured on your own data? Train two candidates for a short schedule on your dataset and compare. Published COCO numbers rank models on 80 everyday classes at moderate resolution. That may have nothing to do with cracks on a weld.

Common mistakes

Believing a version number is a ranking. YOLOv9, v10 and v12 came from three different groups. They are not sequential improvements on one codebase.

Comparing at different input sizes. A model at 1280 will beat the same model at 640 on small objects and be four times slower. Fix the resolution before comparing anything.

Comparing PyTorch latency to TensorRT latency. Published speed tables are usually exported and half-precision. Your unexported PyTorch loop is not the same measurement.

Ignoring the licence until launch. This is the single most expensive mistake in this list.

Skipping the export test. Some architectures export cleanly to ONNX, some need work. Find out on day one.

Try it yourself

Build yolov10n.yaml and yolo11n.yaml, and print head.one2one_cv2 for the one that has it. Then count how many output tensors each head returns in eval mode. That difference is the whole NMS-free mechanism, visible in ten lines of code.

What to learn next

Researcher — Mathematics and papers.

What actually changed, mechanism by mechanism

The YOLO line is best read as four independent threads rather than one sequence.

Assignment. v1 assigned one grid cell per object. v2 to v5 assigned by anchor IoU. YOLOX introduced SimOTA, a top-$k$ approximation of the optimal transport assignment of Ge et al. (2021). v8 onward uses task-aligned assignment (Feng et al., 2021, TOOD), scoring candidates by

$$ t = s^{\alpha} \cdot u^{\beta} $$

Where $s$ is the classification score for the target class, $u$ the predicted-to-target IoU, and $\alpha, \beta$ control the balance. Candidates are ranked by $t$ and the top $k$ per object become positives. The point is that classification confidence and localisation quality are selected jointly. The head does not learn to be confident about badly-placed boxes.

Box representation. Single scalars, then log-space anchor offsets, then $(l,t,r,b)$ distances. Then the discrete distribution of Li et al. (2020), Generalized Focal Loss:

$$ \hat{y} = \sum_{i=0}^{n} P(y_i) \, y_i $$

Where $P$ is a softmax over $n+1$ bins covering the plausible distance range. DFL supervises the two bins adjacent to the continuous target:

$$ \mathrm{DFL}(S_i, S_{i+1}) = -\big((y_{i+1} - y)\log S_i + (y - y_i)\log S_{i+1}\big) $$

This keeps the predicted distribution sharp around the true value, and still a proper probability distribution. That is what makes the expectation sub-cell accurate.

Duplicate removal. v1 to YOLO11 use NMS. YOLOv10 (Wang et al., 2024) trains two heads from a shared backbone, with consistent matching metrics. One is one-to-many and one is one-to-one. The one-to-many head is discarded at inference. The one-to-many branch supplies the dense supervision that makes DETR-style training slow to converge; the one-to-one branch supplies the clean output.

Backbone and neck. CSP connections (v4, v5), reparameterisable blocks (v6, v7), programmable gradient information (v9), attention blocks with area attention and FlashAttention (v12). These deliver smaller gains than the assignment and representation changes. They are also the parts most sensitive to the training recipe.

The benchmarking problem

Published YOLO comparisons are unusually hard to trust, for reasons that are structural rather than dishonest.

  • Recipe coupling. Epoch count, mosaic schedule, EMA, and augmentation differ between releases. Liu et al. (2022) studied CNNs against transformers. Recipe differences accounted for most of the claimed architectural gap. The same caution applies here.
  • Latency measurement. With or without NMS, with or without pre-processing, at what batch size, in what precision, on what hardware. A table without all five is not comparable.
  • Parameter count is not latency. Attention blocks and depthwise convolutions are memory-bandwidth-bound, so FLOPs and parameters correlate poorly with wall-clock time on edge devices.

The defensible comparison is one you run yourself. Same data, same resolution, same augmentation budget, same export path. Measure latency end to end, including pre- and post-processing.

Where the frontier sits

As of mid-2026 the real-time frontier is contested between the YOLO line and the real-time DETR line.

  • Ultralytics reports YOLO26x at 57.5 COCO mAP with T4 TensorRT latency of 11.8 ms.
  • RT-DETRv4 (arXiv 2510.25257, ECCV 2026) reports 55.4 AP at 124 FPS for its L variant, and 57.0 AP for X. The baselines are DEIM-L at 54.7 and YOLOv13-L at 53.4. Its contribution is a training-only distillation from a frozen DINOv3 model, so inference cost is unchanged.
  • RF-DETR is reported as the first real-time detector past 60 mAP on COCO, and is Apache-2.0 licensed.

Two things follow. First, the architectural gap between the families is now small. Licence, tooling and export path decide most real deployments. Second, the interesting recent gains come from training-time supervision (foundation-model distillation, better matching) rather than from inference-time architecture.

Papers

What to learn next