Detecting very small objects
Small objects fail for four separate reasons, and resizing the picture before the model runs is usually the largest of them.
- 18 min read
- 3 reading levels
- Updated
Read these first
On this page 9
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
Small objects fail mostly because the picture was shrunk before the model ever saw it.
The analogy
Think about photographing a page of a book from across the room, then printing it as a passport photo. You still have the page. You cannot read a word of it.
The letters were there in the room. They did not survive the shrinking. Asking a better reader to look at the tiny print will not help, because the information is already gone.
Detection models do this shrinking by default, and most people never notice.
Why it happens
Detectors expect a fixed input size, often around six hundred pixels on the longest side. A drone photo four thousand pixels wide gets shrunk to fit.
That shrink is roughly six times. A person thirty pixels tall in the original becomes five pixels tall. There is nothing left to recognise.
The second reason is the grid. A detector looks at the picture through a grid of cells, and each cell is responsible for what it covers. If an object is smaller than one cell, no cell has a good view of it.
The third reason is the training data. Datasets contain far more medium and large objects than tiny ones. The model spends most of its training on things it can already see.
The fourth is how accuracy is scored. A few pixels of error on a large object barely matters. The same few pixels on a tiny object means the box is counted as wrong.
How to fix it
the whole picture, shrunk cut into overlapping pieces
+-------------------+ +-----+-----+-----+
| . . . . . . . | | | | | each piece is fed at
| everything tiny | --> +-----+-----+-----+ its original size, so
| | | | | | small things stay big
+-------------------+ +-----+-----+-----+Cut the picture into overlapping pieces and run the detector on each one at its full size. Then shift every box back into the original picture's coordinates and merge the duplicates.
The pieces overlap on purpose, so an object sitting on a seam is complete in at least one of them.
This costs time. Forty-eight pieces means forty-eight passes of the model. That is the honest trade, and it is usually worth it.
The other fixes, briefly
Feed a bigger picture. Going from six hundred to twelve hundred pixels doubles every object. It also makes the model roughly four times slower.
Detect from a finer grid. Most detectors have a coarsest and a finest level. Adding a finer one helps small objects and costs memory.
Copy and paste small objects during training. If your data has few tiny objects, paste more of them into training pictures. This is a well-tested trick.
The most important sentence here
Before changing anything about your model, measure the size of your objects in pixels after your pipeline resizes them.
Most teams who believe they have a model problem actually have a resizing problem. Checking takes ten minutes and often ends the investigation.
Where this matters
- Drone surveys counting plants, animals or roof damage.
- Satellite images finding vehicles and small buildings.
- Factory cameras finding tiny surface defects.
- Sports footage tracking a ball across a wide pitch.
Remember this
- The picture is usually shrunk before the model sees it, and that is the biggest cause.
- Cutting the picture into overlapping pieces at full resolution is the most reliable fix.
- Measure your object sizes in pixels after resizing, before blaming the model.
What to learn next
- Anchor-free detection — removing the anchor-IoU wall that blocks tiny objects.
- Feature pyramid networks — where the finest level comes from, and what it costs.
- Data augmentation — the copy-paste and scale-jitter side of the fix.
Developer — Code and libraries.
Setup
pip install torch torchvisionWritten against torch 2.5.1 and torchvision 0.20.1. Everything below is geometry and runs instantly on CPU.
Diagnosing all four causes with numbers
import torch
from torchvision.ops import box_iou, batched_nms
# --- 1. Resizing is where most small objects die, before the model even runs.
W, H = 4000, 3000 # a drone photo
for target in [640, 1024, 1280]:
scale = target / max(W, H)
print(f"resize longest side to {target:>4}: scale {scale:.3f} "
f"a 30px object becomes {30 * scale:4.1f}px, a 12px object becomes {12 * scale:4.1f}px")
# --- 2. How many grid cells a small object gets at each detector stride.
print("\ngrid cells covered by an object at each stride:")
print(f"{'object':>10}{'stride 4':>11}{'stride 8':>11}{'stride 16':>11}{'stride 32':>11}")
for size in [8, 16, 32, 64]:
row = "".join(f"{(size / s) ** 2:>11.2f}" for s in (4, 8, 16, 32))
print(f"{f'{size}x{size}':>10}" + row)
# --- 3. Anchor coverage: a standard anchor set barely touches a tiny box.
anchors = []
for size in (32, 64, 128, 256, 512):
for ratio in (0.5, 1.0, 2.0):
w, h = size * ratio ** 0.5, size / ratio ** 0.5
anchors.append([100 - w / 2, 100 - h / 2, 100 + w / 2, 100 + h / 2])
anchors = torch.tensor(anchors)
print("\nbest IoU between a perfectly centred standard anchor set and one object:")
for s in [8, 12, 16, 24, 32, 64]:
gt = torch.tensor([[100. - s / 2, 100. - s / 2, 100. + s / 2, 100. + s / 2]])
print(f" {s:>2}x{s:<2} object -> best IoU {box_iou(anchors, gt).max():.3f}")
# --- 4. Tiling: cut the picture up, detect on each piece, put the boxes back.
def tiles(width, height, size=640, overlap=0.2):
step = int(size * (1 - overlap))
xs = list(range(0, max(width - size, 0) + 1, step)) or [0]
ys = list(range(0, max(height - size, 0) + 1, step)) or [0]
if xs[-1] + size < width:
xs.append(width - size)
if ys[-1] + size < height:
ys.append(height - size)
return [(x, y) for y in ys for x in xs]
grid = tiles(4000, 3000, size=640, overlap=0.2)
print(f"\n4000x3000 image, 640 tiles, 20 percent overlap: {len(grid)} tiles")
print(f" first four origins: {grid[:4]}")
print(" each tile is fed at native resolution, so a 12px object stays 12px")
# One object sitting on a tile seam gets found twice, in two coordinate systems.
obj = torch.tensor([1290., 500., 1330., 545.]) # 40x45 px, near a seam
found = []
for (tx, ty) in grid:
tile = torch.tensor([tx, ty, tx + 640, ty + 640], dtype=torch.float32)
inter = torch.tensor([max(obj[0], tile[0]), max(obj[1], tile[1]),
min(obj[2], tile[2]), min(obj[3], tile[3])])
if inter[2] > inter[0] and inter[3] > inter[1]:
local = inter - torch.tensor([tx, ty, tx, ty]) # tile-local coordinates
found.append(((tx, ty), local, inter))
print(f"\nthe object at {[int(v) for v in obj.tolist()]} falls inside {len(found)} tiles:")
for (tx, ty), local, glob in found:
print(f" tile ({tx:>4},{ty:>4}) local {[round(v.item(), 1) for v in local]}"
f" -> global {[round(v.item(), 1) for v in glob]}")
boxes = torch.stack([g for _, _, g in found])
scores = torch.tensor([0.81, 0.88])
labels = torch.zeros(len(boxes), dtype=torch.long)
print("\nIoU between the two partial detections:",
round(box_iou(boxes[:1], boxes[1:]).max().item(), 3))
keep = batched_nms(boxes, scores, labels, 0.4)
print("NMS keeps :", [[round(v, 1) for v in boxes[i].tolist()] for i in keep.tolist()])
merged = torch.stack([boxes[:, 0].min(), boxes[:, 1].min(),
boxes[:, 2].max(), boxes[:, 3].max()])
print("union merge:", [round(v.item(), 1) for v in merged])
print("true box :", [round(v, 1) for v in obj.tolist()])resize longest side to 640: scale 0.160 a 30px object becomes 4.8px, a 12px object becomes 1.9px
resize longest side to 1024: scale 0.256 a 30px object becomes 7.7px, a 12px object becomes 3.1px
resize longest side to 1280: scale 0.320 a 30px object becomes 9.6px, a 12px object becomes 3.8px
grid cells covered by an object at each stride:
object stride 4 stride 8 stride 16 stride 32
8x8 4.00 1.00 0.25 0.06
16x16 16.00 4.00 1.00 0.25
32x32 64.00 16.00 4.00 1.00
64x64 256.00 64.00 16.00 4.00
best IoU between a perfectly centred standard anchor set and one object:
8x8 object -> best IoU 0.063
12x12 object -> best IoU 0.141
16x16 object -> best IoU 0.250
24x24 object -> best IoU 0.562
32x32 object -> best IoU 1.000
64x64 object -> best IoU 1.000
4000x3000 image, 640 tiles, 20 percent overlap: 48 tiles
first four origins: [(0, 0), (512, 0), (1024, 0), (1536, 0)]
each tile is fed at native resolution, so a 12px object stays 12px
the object at [1290, 500, 1330, 545] falls inside 2 tiles:
tile (1024, 0) local [266.0, 500.0, 306.0, 545.0] -> global [1290.0, 500.0, 1330.0, 545.0]
tile (1024, 512) local [266.0, 0.0, 306.0, 33.0] -> global [1290.0, 512.0, 1330.0, 545.0]
IoU between the two partial detections: 0.733
NMS keeps : [[1290.0, 512.0, 1330.0, 545.0]]
union merge: [1290.0, 500.0, 1330.0, 545.0]
true box : [1290.0, 500.0, 1330.0, 545.0]Reading that output, block by block
A 30-pixel object becomes 4.8 pixels at input size 640. Under COCO's definition, anything below 32x32 pixels is "small". At this scale nearly everything in a drone photo is below that. No architectural change recovers information the resize destroyed.
Even at 1280 the same object is 9.6 pixels. Raising input resolution helps and does not solve it. That is the argument for tiling in one line.
An 8x8 object covers 0.06 cells at stride 32. Not one cell. Six hundredths of one. The deepest pyramid level has nothing to work with. Detectors for small objects add a finer level, or drop the coarse ones.
Anchor coverage is the quiet killer. A 12x12 object reaches a best IoU of 0.141 against a standard anchor set, even when perfectly centred on it. Every common assignment rule labels that anchor background. The object generates no positive training sample at all, so the model never learns it exists. This is a strong argument for anchor-free heads on small-object work; see anchor-free detection.
48 tiles for one image. That is 48 forward passes for one photo. Tiling is not free, and batching the tiles is the main way to make it tolerable.
The last block is the trap that people ship. The object straddles a seam. The tile at y=0 sees all of it; the tile at y=512 sees the bottom 33 pixels. NMS ranks by score, and the truncated detection happened to score higher. NMS kept the wrong box and lost 12 pixels off the top.
A union merge recovers the true box exactly. This is why SAHI offers non-maximum merging alongside suppression. When two detections of the same object are both truncated, picking one is wrong. Merging is right.
A workable tiled inference loop
import torch
from torchvision.ops import batched_nms
def tiled_detect(detect_fn, image, tile=640, overlap=0.2, batch=8):
"""detect_fn(batch_of_tiles) -> list of (boxes_xyxy, scores, labels) per tile."""
_, H, W = image.shape
step = int(tile * (1 - overlap))
xs = sorted({*range(0, max(W - tile, 0) + 1, step), max(W - tile, 0)})
ys = sorted({*range(0, max(H - tile, 0) + 1, step), max(H - tile, 0)})
origins = [(x, y) for y in ys for x in xs]
all_boxes, all_scores, all_labels = [], [], []
for i in range(0, len(origins), batch):
chunk = origins[i:i + batch]
crops = torch.stack([image[:, y:y + tile, x:x + tile] for x, y in chunk])
for (x, y), (b, s, l) in zip(chunk, detect_fn(crops)):
if len(b) == 0:
continue
b = b + torch.tensor([x, y, x, y], dtype=b.dtype) # tile -> image coords
all_boxes.append(b); all_scores.append(s); all_labels.append(l)
if not all_boxes:
return torch.zeros(0, 4), torch.zeros(0), torch.zeros(0, dtype=torch.long)
boxes = torch.cat(all_boxes); scores = torch.cat(all_scores); labels = torch.cat(all_labels)
keep = batched_nms(boxes, scores, labels, 0.5)
return boxes[keep], scores[keep], labels[keep]
# A fake detector so this runs with no weights: it "finds" one box per tile.
def fake_detect(crops):
n = crops.shape[0]
return [(torch.tensor([[100., 100., 140., 145.]]),
torch.tensor([0.9]),
torch.tensor([0])) for _ in range(n)]
img = torch.zeros(3, 1500, 2000)
b, s, l = tiled_detect(fake_detect, img)
print("tiles processed :", len(sorted({*range(0, 2000 - 640 + 1, 512), 1360})) *
len(sorted({*range(0, 1500 - 640 + 1, 512), 860})))
print("detections kept :", len(b))
print("first box :", [round(v, 1) for v in b[0].tolist()])
print("last box :", [round(v, 1) for v in b[-1].tolist()])tiles processed : 12 detections kept : 12 first box : [100.0, 100.0, 140.0, 145.0] last box : [1460.0, 960.0, 1500.0, 1005.0]
The fake detector returns the identical tile-local box for every tile. The 12 detections should therefore land at 12 different places, with no overlap. All 12 survived NMS, and the first and last boxes are far apart. That is the correctness check for the coordinate shift: identical local boxes must become distinct global boxes.
Swap fake_detect for a real model and the structure is unchanged. The sorted({*range(...), max(W - tile, 0)}) trick guarantees the right and bottom edges are covered. That holds even when the image size is not a multiple of the step.
The fixes, ranked by what they cost you
| Fix | Effect | Cost |
|---|---|---|
| Measure object sizes after resizing | none, but it tells you what to do | ten minutes |
| Raise input resolution | doubles object size for double the side length | around 4x compute |
| Tile the image (SAHI-style) | keeps native resolution | one pass per tile |
| Add a finer pyramid level | more cells for small objects | large memory increase |
| Go anchor-free | removes the anchor-IoU wall | architecture change |
| Copy-paste small objects in training | more positive samples | data pipeline work |
| Fine-tune on tiles as well | inference and training distributions match | a training run |
The SAHI paper reports gains from sliced inference alone of 6.8, 5.1 and 5.3 AP. Those are for FCOS, VFNet and TOOD. Adding slicing-aided fine-tuning took the cumulative gains to 12.7, 13.4 and 14.5 points. Those numbers are from their benchmark, not a promise about yours. The ordering they imply is still worth taking seriously. Tile at inference first, then fine-tune on tiles.
Common mistakes
Tiling at inference but training on whole images. The model then sees a different object-size distribution at test time than it trained on. Fine-tune on tiles as well.
Overlap too small. Overlap must exceed the largest object you expect, or some object is truncated in every tile. 20 percent of the tile size is a reasonable default for genuinely small objects.
Merging with plain NMS. As shown above, this keeps whichever truncated box scored higher. Use a merge that takes the union for boxes touching a tile edge.
Forgetting the coordinate shift. Add the tile origin to every box. This bug produces boxes clustered in the top-left corner, which is at least easy to spot.
Judging with overall mAP. COCO reports AP-small, AP-medium and AP-large separately. Track AP-small, or your improvement will be invisible inside a number dominated by large objects.
Skipping the diagnosis. Print a histogram of your object widths in pixels after resizing. Do this before anything else on this page.
Try it yourself
Take the tiling function and set overlap=0.0. Re-run the seam test with the object at [1290., 500., 1330., 545.]. It is now truncated in every tile that sees it, and no merge recovers it. Then raise overlap until the object is complete somewhere. The value you find is a lower bound on the overlap your data needs.
What to learn next
- Anchor-free detection — removing the anchor-IoU wall that blocks tiny objects.
- Feature pyramid networks — where the finest level comes from, and what it costs.
- Data augmentation — the copy-paste and scale-jitter side of the fix.
Researcher — Mathematics and papers.
Why small objects are hard, in four separable terms
Information. After resizing by factor $s$, an object of side $d$ occupies $sd$ pixels. Detail below the new Nyquist limit is unrecoverable. This term is upstream of the model entirely.
Sampling. At stride $r$, an object of side $d$ covers $(d/r)^2$ feature positions. Below one position, no location is fully responsible for the object. The classification and regression heads then see a feature vector dominated by context.
Assignment. Under IoU assignment, the achievable IoU between a $d \times d$ object and an anchor of side $a$, perfectly centred, is $\min(d,a)^2 / \max(d,a)^2$. For $d = 12$ and $a = 32$ this is $144/1024 \approx 0.14$, below every common positive threshold. The object contributes no positive sample. This is the term the output block above measures.
Evaluation. IoU is scale-invariant in value but not in sensitivity to pixel error. Shifting a $d \times d$ box by $\delta$ pixels along one axis gives $\mathrm{IoU} = (d - \delta)/(d + \delta)$. For $d = 100, \delta = 5$ that is 0.905. For $d = 10, \delta = 5$ it is 0.333. One absolute localisation error moves a small object across the match threshold. The same error leaves a large one comfortably inside it.
That last term means AP-small is partly a measure of sub-pixel regression precision, not only of detection. Xu et al. (2022) propose the Normalized Wasserstein Distance as an assignment metric. It degrades gracefully on small boxes, where IoU falls off a cliff.
Slicing-aided inference
Akyon et al. (2022), SAHI, formalise the tiling pipeline: slice into overlapping patches, run any detector per patch, map boxes back, and merge. The method is detector-agnostic and needs no retraining, which is why it is the standard first intervention.
Their reported results are worth reading precisely. Sliced inference alone gave AP increases of 6.8, 5.1 and 5.3 points for FCOS, VFNet and TOOD. Adding slicing-aided fine-tuning, where the model is also trained on slices, produced cumulative increases of 12.7, 13.4 and 14.5 points. The fine-tuning term is roughly as large as the inference term, and it is the half most people skip.
Two implementation details decide whether it works.
Overlap must exceed the largest expected object, or some object is truncated in every slice.
Merging must not be plain NMS. Two truncated detections of one object are both wrong; greedy selection keeps one wrong box. Non-maximum merging replaces the cluster with the union of its members, which is correct when both members are partial views.
Architectural responses
- Higher-resolution levels. Adding P2 (stride 4) or P1 (stride 2) multiplies feature-map area by 4 or 16. This is the dominant memory cost in high-resolution detectors.
- Anchor-free heads remove the assignment term entirely, since positives are defined by containment rather than IoU. See anchor-free detection.
- Scale-aware assignment. ATSS's per-object adaptive threshold naturally admits smaller objects than a fixed 0.5.
- Alternative assignment metrics. NWD and dot-distance measures replace IoU where IoU is numerically unstable at small scales.
- Query-based detectors with deformable attention sample specific points rather than pooling a region. That suits objects covering a fraction of a feature cell. YOLO26's small-target-aware label assignment is a recent example of the same concern in a convolutional head.
Data-side responses
Copy-paste augmentation (Kisantal et al., 2019, for small objects; Ghiasi et al., 2021, in general) pastes additional instances of small objects into training images. On COCO, small objects appear in a minority of images and are a small fraction of instances. Oversampling those images and duplicating instances within them both help.
Scale jitter with a wide range forces the model to see the same object at many sizes. Large-scale jitter is one of the more reliable gains in modern detection recipes.
Mosaic augmentation, standard in YOLO training, composes four images into one. That shrinks objects and raises the density of small instances per training image.
Measuring it properly
Report AP-small, AP-medium and AP-large separately. COCO defines the boundaries at area $< 32^2$ and $> 96^2$ pixels, in original image coordinates.
Two cautions. Those boundaries are absolute pixel areas defined for COCO's typical resolution. On 4000-pixel-wide imagery almost every object is "small" by that definition. The split stops being informative. Define your own boundaries from your own size histogram.
And report the size distribution after your resize, not from the annotation file. The gap between those two histograms is, in most projects, the entire problem.
Papers
- Akyon et al., Slicing Aided Hyper Inference and Fine-tuning for Small Object Detection, ICIP 2022 — arxiv.org/abs/2202.06934
- Kisantal et al., Augmentation for Small Object Detection, 2019 — arxiv.org/abs/1902.07296
- Xu et al., Detecting Tiny Objects in Aerial Images: A Normalized Wasserstein Distance and a New Benchmark, 2022 — arxiv.org/abs/2206.13996
- Ghiasi et al., Simple Copy-Paste is a Strong Data Augmentation Method, CVPR 2021 — arxiv.org/abs/2012.07177
What to learn next
- Anchor-free detection — removing the anchor-IoU wall that blocks tiny objects.
- Feature pyramid networks — where the finest level comes from, and what it costs.
- Data augmentation — the copy-paste and scale-jitter side of the fix.