YOLO
YOLO looks at the whole picture in one pass and reports every object at once, which is what made object detection fast enough for live video.
- 19 min read
- 3 reading levels
- Updated
Read these first
On this page 8
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
YOLO is a detector that looks at the whole picture once and reports every object at the same time.
Think of a teacher taking attendance. The slow way is to walk down every row, stop at each desk, and check who is sitting there. Forty desks, forty separate looks.
The fast way is to stand at the front and take the whole class in at one sweep. You know at once who is where. Same information, one glance.
YOLO does the second thing. The name is short for You Only Look Once, and it is a description of the method, not marketing.
Why it had to be invented
The detectors that came before it worked the slow way. They cut out thousands of candidate rectangles from the picture and ran a classifier on each one, separately.
That is thorough and it is painfully slow. Early versions took tens of seconds for a single photograph. Anything involving video was out of the question.
The problem was structural, not a matter of faster computers. Thousands of separate looks will always cost thousands of times one look.
YOLO's move was to stop looking many times. Push the picture through the network once. That single pass produces every box, every label and every confidence together.
How the single look works
The trick is to divide the picture into a grid. Each square of the grid is responsible for its own patch.
+-----+-----+-----+-----+
| | | | |
+-----+-----+-----+-----+
| | | dog | | <- the dog's centre falls here,
+-----+-----+-----+-----+ so this square answers for the dog
| | | | |
+-----+-----+-----+-----+
|chair| | | | <- and this one answers for the chair
+-----+-----+-----+-----+Every square reports the same three things, whether or not anything is there:
- Is there an object whose centre sits in me?
- If so, where exactly are its edges?
- And what is it?
Squares with nothing in them report "nothing here" with high confidence, and get thrown away. Squares with something in them produce a box.
All sixteen squares in that picture are filled in during the same pass. Nothing is looked at twice.
Why that made video possible
Once detection costs one pass instead of thousands, the arithmetic changes completely.
old way: picture -> 2000 crops -> 2000 classifier runs -> boxes seconds
YOLO: picture -> 1 pass -> boxes millisecondsA phone can now do this on a live camera feed. So can a small camera bolted above a factory belt. That shift, from "a research demo" to "runs on the thing you already own", is YOLO's real contribution.
About the version numbers
You will see YOLOv3, YOLOv5, YOLOv8, YOLOv11 and more. This is confusing, and the confusion is real rather than something you have misunderstood.
The first three versions came from one researcher, Joseph Redmon. After that, different teams around the world each built their own and kept the name. There is no single owner and no single roadmap. A higher number does not always mean a better model for your job.
There is also a piece of history worth knowing. Redmon stopped working on computer vision in 2020. He said publicly that he could not ignore how his work was being used. Military purposes, and surveillance. He walked away from a field he had helped define. Whatever you conclude from that, the person who built this thought hard about where it ends up.
Where you have already seen it
- A shop counting how many people entered today, from the CCTV feed.
- Traffic cameras counting vehicles by type at a junction.
- A phone camera drawing live boxes around faces and pets before you press the button.
- Sports analysis tracking players and the ball through a match.
- Farm cameras counting fruit on a tree to estimate the harvest.
The honest parts
Free to learn with is not the same as free to sell with. The most popular YOLO packages carry a licence with a condition attached. Build a product on them, and you must publish your own source code. If your project is a hobby or a class assignment, you are fine. If it is a business, read the licence before you write a line, or buy the commercial one. This catches out a lot of people late and expensively.
Small objects are still hard. YOLO is fast, not magic. A face thirty metres away is a handful of dots. No amount of speed recovers detail that was never captured.
Speed claims are measured on strong hardware. A number like "300 frames per second" comes from a powerful graphics card. On a laptop processor you will get a small fraction of it. That is often still enough — but check on your machine, not on the chart.
Remember this
- YOLO makes one pass over the whole picture and produces every box at once.
- It divides the picture into a grid, and each square answers for the objects centred in it.
- One pass instead of thousands is what made live video detection possible.
- The version numbers come from different teams, and the licence matters if you plan to sell something.
What to learn next
- Object detection — IoU, NMS and mAP, which every YOLO version depends on.
- Image segmentation — from rectangles to exact outlines.
- Convolutional neural networks — the backbone doing the single pass.
Developer — Code and libraries.
Two parts. First decode a YOLO-style grid by hand, so nothing is hidden. Then run a real, current YOLO model on a real photograph, on a CPU.
Setup
pip install ultralytics matplotlibBe honest with yourself about the sizes. ultralytics is a small package that pulls in PyTorch and OpenCV, so the first install is a few hundred megabytes. The model weights are separate and tiny: 6.3 MB for the smallest one, downloaded once and cached.
matplotlib is here only because it ships a photograph inside the package, so nothing else is downloaded.
The grid, decoded by hand
This is the part that makes YOLO click. The network emits a fixed block of numbers per grid cell, and turning that block into boxes is plain arithmetic.
import numpy as np
def sigmoid(x):
return 1.0 / (1.0 + np.exp(-x))
IMG = 64 # the image is 64x64 pixels
GRID = 4 # YOLO splits it into a 4x4 grid
STRIDE = IMG // GRID
ANCHOR_W, ANCHOR_H = 20.0, 30.0 # one prior box shape, chosen before training
CLASSES = ["dog", "cat", "chair"]
# What the network actually emits: 8 numbers per grid cell.
# [ shift_x, shift_y, size_w, size_h, is_there_anything, dog, cat, chair ]
raw = np.full((GRID, GRID, 8), -5.0) # -5 through a sigmoid is about 0.007: "nothing here"
# Cell (row 1, column 2) fires for a dog, dead centre of its cell, at the anchor size.
raw[1, 2] = [0.0, 0.0, 0.0, 0.0, 2.0, 3.0, -5.0, -5.0]
# Cell (row 3, column 0) fires for a chair, offset inside its cell and wider than the anchor.
raw[3, 0] = [1.1, -0.4, 0.5, -0.2, 1.5, -5.0, -5.0, 2.2]
print(f"grid cells: {GRID} x {GRID} = {GRID * GRID} predictions, from one pass over the image")
print()
for row in range(GRID):
for col in range(GRID):
tx, ty, tw, th, obj, *cls = raw[row, col]
# The centre is a shift INSIDE this cell, so a cell can only claim its own patch.
cx = (sigmoid(tx) + col) * STRIDE
cy = (sigmoid(ty) + row) * STRIDE
# The size is a multiplier ON the anchor, so the network never predicts a negative width.
bw = ANCHOR_W * np.exp(tw)
bh = ANCHOR_H * np.exp(th)
score = sigmoid(obj) * sigmoid(np.array(cls))
best = int(score.argmax())
if score[best] < 0.25:
continue
print(f"cell (row {row}, col {col}) -> {CLASSES[best]:5s} score {score[best]:.2f}")
print(f" centre ({cx:.1f}, {cy:.1f}) size {bw:.1f} x {bh:.1f}")
print(f" box [{cx - bw / 2:.1f}, {cy - bh / 2:.1f}, "
f"{cx + bw / 2:.1f}, {cy + bh / 2:.1f}]")grid cells: 4 x 4 = 16 predictions, from one pass over the image
cell (row 1, col 2) -> dog score 0.84
centre (40.0, 24.0) size 20.0 x 30.0
box [30.0, 9.0, 50.0, 39.0]
cell (row 3, col 0) -> chair score 0.74
centre (12.0, 54.4) size 33.0 x 24.6
box [-4.5, 42.1, 28.5, 66.7]Sixteen predictions were computed. Fourteen said "nothing here" and were dropped by one threshold check.
Three design choices worth understanding
The centre is a shift inside the cell, not a position in the image. sigmoid(tx) is squeezed between zero and one, and adding the column index places it inside that cell and nowhere else. A cell physically cannot claim an object in the far corner of the image. That constraint is what makes the grid a division of labour rather than sixteen models arguing.
The size is a multiplier on a prior box, not a raw width. ANCHOR_W * exp(tw) is always positive, so the network never has to learn "widths cannot be negative". And because it is a multiplier, predicting tw = 0 gives exactly the anchor size. The network starts from a sensible guess and learns a correction. That anchor is a fixed box shape decided before training, usually by clustering the shapes in the training data.
Look at the second box: [-4.5, 42.1, 28.5, 66.7]. It runs off the left edge and off the bottom of a 64-pixel image. Nothing in the decode prevents that. Every real implementation clips boxes to the image bounds afterwards. Forget to clip and you get negative coordinates, which crash a cropping step much later, in a place that gives you no clue why.
Now the real thing
import matplotlib.cbook as cbook
from ultralytics import YOLO
# The same photograph used in the object detection lesson, shipped inside matplotlib.
path = cbook.get_sample_data("grace_hopper.jpg").name
model = YOLO("yolov8n.pt") # downloads about 6.3 MB the first time only
print("parameters:", f"{sum(p.numel() for p in model.model.parameters()):,}")
print("classes it knows:", len(model.names))
result = model.predict(path, conf=0.25, verbose=False)[0]
print("image size:", result.orig_shape)
print("detections:", len(result.boxes))
for box in result.boxes:
name = result.names[int(box.cls)]
left, top, right, bottom = (int(v) for v in box.xyxy[0])
print(f" {name:8s} {float(box.conf):.2f} "
f"left={left:3d} top={top:3d} right={right:3d} bottom={bottom:3d}")parameters: 3,157,200 classes it knows: 80 image size: (600, 512) detections: 2 person 0.91 left= 0 top= 28 right=511 bottom=599 tie 0.53 left=217 top=409 right=293 bottom=532
On your very first run, a download progress line appears above this. ultralytics saves yolov8n.pt into the folder you ran from, not into a hidden cache, so the second run is silent. Confidences may also shift in the second decimal place on a different build.
Two detections, both correct: a person, and the tie she is wearing.
Compare that with the previous lesson, carefully
In object detection, SSDLite320 saw this exact photograph. It found the person at 1.00, and the tie six times over, every copy scoring below 0.15. No threshold gave a clean answer.
YOLOv8n finds both objects, once each, with the tie at 0.53. It does it with fewer parameters — 3.16 million against 3.44 million — and a smaller weights file.
Resist the conclusion that YOLO is better. The honest reasons for the gap are specific:
Input resolution. SSDLite runs at 320 pixels. YOLOv8 defaults to 640. Four times the pixels is most of the difference on a medium-sized object like a tie. It also costs roughly four times the compute.
Six years of training recipe. Mosaic augmentation, better label assignment and longer schedules account for a large share of modern detector gains, independent of architecture.
Different reported scores on the same benchmark. SSDLite320 reports a COCO mAP of 21.3. Ultralytics report roughly 37 for YOLOv8n at 640 pixels. Comparing those two numbers directly compares two resolutions as much as two models.
The lesson generalises: when one detector beats another, check the input size before you credit the architecture.
The one-line version
ultralytics installs a command as well as a library.
yolo predict model=yolov8n.pt source="path/to/your/photo.jpg" conf=0.25This writes an annotated copy of your image into a runs/detect/predict folder. No output block here. The printed summary includes timings and paths that vary by machine, and an invented one would teach you to expect the wrong thing.
Common mistakes
Ignoring the licence. Ultralytics YOLO releases are AGPL-3.0. In plain terms: build a product on them, let users reach it over a network, and you owe those users your source code. A commercial licence is available. Decide before you build, not after. Several other YOLO variants carry different licences again, so check the exact repository and the exact weights file.
Leaving conf at the default. The default of 0.25 is a demo setting. A missed defect and a false alarm rarely cost the same, and that ratio, not a library default, sets your threshold.
Changing imgsz at inference and not at training. The model learned at one resolution. Running it at another shifts accuracy in ways that look like a broken model. If you need a different resolution, train at it.
Comparing models across resolutions. Covered above, and worth repeating because it is the single most common way detector benchmarks get misread.
Assuming the classes are yours. A pretrained YOLO knows 80 everyday COCO categories. Cracked welds, diseased leaves and counterfeit labels are not among them. Detecting those means fine-tuning on your own labelled boxes, which brings back all of the labelling effort from the previous lesson.
Treating tracking as free. Detection gives boxes per frame with no identity between frames. Following one person across a video is a separate problem, called tracking. ultralytics exposes model.track(...), and it has its own failure modes when people cross paths.
Try it yourself
In decode_grid.py, set raw[1, 2] and raw[1, 3] to fire for a dog with the same size. Predict what you get before running it: two boxes, in neighbouring cells, overlapping heavily. That is exactly the duplicate problem that non-maximum suppression exists to clean up, reproduced in fifteen lines.
Then run run_yolo.py on a photograph from your own phone. Count the objects you can see. Count the ones it found. Then lower conf to 0.1 and count again.
The gap between those two counts, on your own pictures, is worth more than any published benchmark score.
What to learn next
- Object detection — IoU, NMS and mAP, which every YOLO version depends on.
- Image segmentation — from rectangles to exact outlines.
- Convolutional neural networks — the backbone doing the single pass.
Researcher — Mathematics and papers.
YOLOv1: the reformulation
Redmon et al. (2016) recast detection as a single regression from image pixels to a fixed-size tensor:
output shape: S x S x (B * 5 + C)S— grid size, 7 in the paper.B— boxes predicted per cell, 2.C— number of classes, 20 for PASCAL VOC.- The 5 per box is
(x, y, w, h, confidence).
Confidence was defined as Pr(Object) * IoU(pred, truth), so the target for confidence is not a constant — it is the achieved localisation quality, which the network must therefore predict about itself.
The loss is sum-of-squares throughout, with two corrections:
L = lambda_coord * sum_ij 1^obj_ij [ (x - x_hat)^2 + (y - y_hat)^2 ]
+ lambda_coord * sum_ij 1^obj_ij [ (sqrt(w) - sqrt(w_hat))^2 + (sqrt(h) - sqrt(h_hat))^2 ]
+ sum_ij 1^obj_ij (C_i - C_hat_i)^2
+ lambda_noobj * sum_ij 1^noobj_ij (C_i - C_hat_i)^2
+ sum_i 1^obj_i sum_classes (p_i(c) - p_hat_i(c))^21^obj_ij— indicator that boxjin celliis responsible for a ground-truth object.lambda_coord = 5,lambda_noobj = 0.5.
Both constants exist for the same reason: the overwhelming majority of cells contain nothing, so an unweighted loss is dominated by pushing confidences to zero. The square roots on w and h make a fixed absolute error matter more on a small box than a large one. It is a crude precursor to the IoU-based losses that replaced it.
The published limitations were honest and specific. At most B objects per cell, so tight groups are lost. Coarse localisation from the 7 x 7 grid. And weak generalisation to unusual aspect ratios.
The lineage
| Version | Year | Team | Principal change |
|---|---|---|---|
| v1 | 2016 | Redmon et al. | Single-pass grid regression |
| v2 / YOLO9000 | 2017 | Redmon & Farhadi | Anchor boxes from k-means, batch norm, high-res fine-tune |
| v3 | 2018 | Redmon & Farhadi | Three-scale prediction, Darknet-53, per-class logistic outputs |
| v4 | 2020 | Bochkovskiy et al. | CSPDarknet, PANet neck, mosaic augmentation, CIoU loss |
| v5 | 2020 | Ultralytics | Engineering and tooling; no paper |
| v6 | 2022 | Meituan | Reparameterisable backbone, industry-oriented |
| v7 | 2022 | Wang et al. | E-ELAN, trainable bag-of-freebies |
| v8 | 2023 | Ultralytics | Anchor-free, decoupled head, DFL, task-aligned assignment |
| v9 | 2024 | Wang & Liao | Programmable gradient information, GELAN |
| v10 | 2024 | Wang et al. (Tsinghua) | NMS-free training via consistent dual assignments |
| v11 | 2024 | Ultralytics | Efficiency refinements across tasks |
| v12 | 2025 | Tian et al. | Attention-centric backbone at YOLO latencies |
Three shifts matter more than the version numbers.
Anchor-free prediction. v8 onward drop anchor boxes and regress distances from each location to the four box edges, following FCOS. This removes the anchor scale, aspect-ratio and assignment-IoU hyperparameters that dominated tuning in the v2 to v7 era.
Distribution Focal Loss. Rather than regressing a scalar edge distance, v8 predicts a discrete distribution over candidate distances and takes its expectation. DFL (Li et al., 2020, Generalized Focal Loss) supervises the two bins adjacent to the target. It models localisation ambiguity explicitly, which measurably helps on blurred and occluded boundaries.
Dynamic label assignment. Static IoU-threshold assignment gave way to SimOTA (YOLOX) and Task-Aligned Assignment (TOOD, adopted by v8), which choose positives using a joint classification-and-localisation quality score. This is one of the larger uncredited sources of gain across the v5-to-v8 span.
NMS as the latency floor
Once the backbone is fast, post-processing dominates the tail. NMS is sequential, data-dependent, and its cost scales with the number of surviving boxes, which is why measured end-to-end latency often diverges from reported backbone FLOPs.
YOLOv10 attacks this directly with consistent dual assignments. A one-to-many head supplies rich training signal. A one-to-one head, trained with a matching metric consistent with the first, is what runs at inference. The one-to-one head emits a single box per object, so NMS is unnecessary. Reported latency drops are largest exactly where box counts are highest.
DETR reached the same destination earlier by a different route — bipartite matching — but at a training cost YOLO's users would not accept.
Reading published speed numbers
A YOLO throughput figure is meaningless without five accompanying facts:
- Hardware, exactly. A number from an A100 tells you nothing about a Raspberry Pi.
- Precision. FP16 and INT8 numbers are routinely quoted next to FP32 accuracy.
- Batch size. Batch-32 throughput can be several times batch-1 latency, and interactive systems only care about batch 1.
- Whether pre- and post-processing are included. On CPU with a small model, NMS and box decoding can take longer than the forward pass itself.
ultralyticsexposes the split asresult.speed, and the first time you print it is usually a surprise. - Input resolution. Cost scales roughly quadratically with it.
Papers report all five; blog posts and comparison charts frequently report one.
Licensing, stated plainly
This is engineering-relevant, not a footnote. Ultralytics YOLO releases — v5, v8, v11 and the ultralytics package itself — are AGPL-3.0. Section 13 of that licence extends the source-disclosure obligation to users who interact with the software over a network, which covers essentially any hosted product. Ultralytics sells a commercial licence for this reason.
Other lines differ. YOLOv4 and YOLOv7 are GPL-3.0. Some research releases are Apache-2.0 or MIT. Derivative repositories generally inherit the licence of what they forked. Check the licence of the specific repository and the specific weights file, since pretrained weights can carry their own terms.
Papers
- Redmon, J. et al. (2016). You Only Look Once: Unified, Real-Time Object Detection. arXiv:1506.02640
- Redmon, J. & Farhadi, A. (2017). YOLO9000: Better, Faster, Stronger. arXiv:1612.08242
- Redmon, J. & Farhadi, A. (2018). YOLOv3: An Incremental Improvement. arXiv:1804.02767
- Bochkovskiy, A. et al. (2020). YOLOv4: Optimal Speed and Accuracy of Object Detection. arXiv:2004.10934
- Ge, Z. et al. (2021). YOLOX: Exceeding YOLO Series in 2021. arXiv:2107.08430 — SimOTA assignment.
- Li, X. et al. (2020). Generalized Focal Loss. arXiv:2006.04388 — the DFL used from v8 onward.
- Feng, C. et al. (2021). TOOD: Task-aligned One-stage Object Detection. arXiv:2108.07755
- Wang, C.-Y. et al. (2022). YOLOv7. arXiv:2207.02696
- Wang, C.-Y. & Liao, H.-Y. M. (2024). YOLOv9: Learning What You Want to Learn Using Programmable Gradient Information. arXiv:2402.13616
- Wang, A. et al. (2024). YOLOv10: Real-Time End-to-End Object Detection. arXiv:2405.14458
- Tian, Y. et al. (2025). YOLOv12: Attention-Centric Real-Time Object Detectors. arXiv:2502.12524
Where the family stands
Real-time detection is close to an engineering commodity. The accuracy-latency Pareto frontier moves by small increments each year, and most of the movement now comes from training recipes, label assignment and deployment engineering rather than from architecture.
Three directions remain genuinely open.
NMS-free at scale. YOLOv10 and the DETR family both remove the sequential bottleneck. Whether one-to-one assignment matches one-to-many on dense, crowded data at equal training budget is not settled.
Open vocabulary at real-time latency. YOLO-World and Grounding DINO variants accept text-defined classes at inference. Closing the accuracy gap to a fine-tuned closed-set detector without giving back the latency is an active problem.
Honest evaluation. COCO mAP has been the target for a decade, and a decade of tuning against one benchmark carries its usual risks. Size-disaggregated AP, per-class breakdowns and evaluation on a distribution resembling deployment tell you far more than the headline figure — and are reported far less often.
What to learn next
- Object detection — IoU, NMS and mAP, which every YOLO version depends on.
- Image segmentation — from rectangles to exact outlines.
- Convolutional neural networks — the backbone doing the single pass.