Object Detection in Depth

Detecting very small objects

Small objects fail for four separate reasons, and resizing the picture before the model runs is usually the largest of them.

On this page 9
  1. The short answer
  2. The analogy
  3. Why it happens
  4. How to fix it
  5. The other fixes, briefly
  6. The most important sentence here
  7. Where this matters
  8. Remember this
  9. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

Small objects fail mostly because the picture was shrunk before the model ever saw it.

The analogy

Think about photographing a page of a book from across the room, then printing it as a passport photo. You still have the page. You cannot read a word of it.

The letters were there in the room. They did not survive the shrinking. Asking a better reader to look at the tiny print will not help, because the information is already gone.

Detection models do this shrinking by default, and most people never notice.

Why it happens

Detectors expect a fixed input size, often around six hundred pixels on the longest side. A drone photo four thousand pixels wide gets shrunk to fit.

That shrink is roughly six times. A person thirty pixels tall in the original becomes five pixels tall. There is nothing left to recognise.

The second reason is the grid. A detector looks at the picture through a grid of cells, and each cell is responsible for what it covers. If an object is smaller than one cell, no cell has a good view of it.

The third reason is the training data. Datasets contain far more medium and large objects than tiny ones. The model spends most of its training on things it can already see.

The fourth is how accuracy is scored. A few pixels of error on a large object barely matters. The same few pixels on a tiny object means the box is counted as wrong.

How to fix it

   the whole picture, shrunk         cut into overlapping pieces
   +-------------------+             +-----+-----+-----+
   |  . . . . . . .    |             |     |     |     |   each piece is fed at
   |  everything tiny  |    -->      +-----+-----+-----+   its original size, so
   |                   |             |     |     |     |   small things stay big
   +-------------------+             +-----+-----+-----+

Cut the picture into overlapping pieces and run the detector on each one at its full size. Then shift every box back into the original picture's coordinates and merge the duplicates.

The pieces overlap on purpose, so an object sitting on a seam is complete in at least one of them.

This costs time. Forty-eight pieces means forty-eight passes of the model. That is the honest trade, and it is usually worth it.

The other fixes, briefly

Feed a bigger picture. Going from six hundred to twelve hundred pixels doubles every object. It also makes the model roughly four times slower.

Detect from a finer grid. Most detectors have a coarsest and a finest level. Adding a finer one helps small objects and costs memory.

Copy and paste small objects during training. If your data has few tiny objects, paste more of them into training pictures. This is a well-tested trick.

The most important sentence here

Before changing anything about your model, measure the size of your objects in pixels after your pipeline resizes them.

Most teams who believe they have a model problem actually have a resizing problem. Checking takes ten minutes and often ends the investigation.

Where this matters

  • Drone surveys counting plants, animals or roof damage.
  • Satellite images finding vehicles and small buildings.
  • Factory cameras finding tiny surface defects.
  • Sports footage tracking a ball across a wide pitch.

Remember this

  • The picture is usually shrunk before the model sees it, and that is the biggest cause.
  • Cutting the picture into overlapping pieces at full resolution is the most reliable fix.
  • Measure your object sizes in pixels after resizing, before blaming the model.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install torch torchvision

Written against torch 2.5.1 and torchvision 0.20.1. Everything below is geometry and runs instantly on CPU.

Diagnosing all four causes with numbers

small_objects.py
import torch
from torchvision.ops import box_iou, batched_nms

# --- 1. Resizing is where most small objects die, before the model even runs.
W, H = 4000, 3000                                   # a drone photo
for target in [640, 1024, 1280]:
    scale = target / max(W, H)
    print(f"resize longest side to {target:>4}: scale {scale:.3f}  "
          f"a 30px object becomes {30 * scale:4.1f}px, a 12px object becomes {12 * scale:4.1f}px")

# --- 2. How many grid cells a small object gets at each detector stride.
print("\ngrid cells covered by an object at each stride:")
print(f"{'object':>10}{'stride 4':>11}{'stride 8':>11}{'stride 16':>11}{'stride 32':>11}")
for size in [8, 16, 32, 64]:
    row = "".join(f"{(size / s) ** 2:>11.2f}" for s in (4, 8, 16, 32))
    print(f"{f'{size}x{size}':>10}" + row)

# --- 3. Anchor coverage: a standard anchor set barely touches a tiny box.
anchors = []
for size in (32, 64, 128, 256, 512):
    for ratio in (0.5, 1.0, 2.0):
        w, h = size * ratio ** 0.5, size / ratio ** 0.5
        anchors.append([100 - w / 2, 100 - h / 2, 100 + w / 2, 100 + h / 2])
anchors = torch.tensor(anchors)
print("\nbest IoU between a perfectly centred standard anchor set and one object:")
for s in [8, 12, 16, 24, 32, 64]:
    gt = torch.tensor([[100. - s / 2, 100. - s / 2, 100. + s / 2, 100. + s / 2]])
    print(f"  {s:>2}x{s:<2} object -> best IoU {box_iou(anchors, gt).max():.3f}")

# --- 4. Tiling: cut the picture up, detect on each piece, put the boxes back.
def tiles(width, height, size=640, overlap=0.2):
    step = int(size * (1 - overlap))
    xs = list(range(0, max(width - size, 0) + 1, step)) or [0]
    ys = list(range(0, max(height - size, 0) + 1, step)) or [0]
    if xs[-1] + size < width:
        xs.append(width - size)
    if ys[-1] + size < height:
        ys.append(height - size)
    return [(x, y) for y in ys for x in xs]

grid = tiles(4000, 3000, size=640, overlap=0.2)
print(f"\n4000x3000 image, 640 tiles, 20 percent overlap: {len(grid)} tiles")
print(f"  first four origins: {grid[:4]}")
print("  each tile is fed at native resolution, so a 12px object stays 12px")

# One object sitting on a tile seam gets found twice, in two coordinate systems.
obj = torch.tensor([1290., 500., 1330., 545.])          # 40x45 px, near a seam
found = []
for (tx, ty) in grid:
    tile = torch.tensor([tx, ty, tx + 640, ty + 640], dtype=torch.float32)
    inter = torch.tensor([max(obj[0], tile[0]), max(obj[1], tile[1]),
                          min(obj[2], tile[2]), min(obj[3], tile[3])])
    if inter[2] > inter[0] and inter[3] > inter[1]:
        local = inter - torch.tensor([tx, ty, tx, ty])   # tile-local coordinates
        found.append(((tx, ty), local, inter))

print(f"\nthe object at {[int(v) for v in obj.tolist()]} falls inside {len(found)} tiles:")
for (tx, ty), local, glob in found:
    print(f"  tile ({tx:>4},{ty:>4})  local {[round(v.item(), 1) for v in local]}"
          f"  -> global {[round(v.item(), 1) for v in glob]}")

boxes = torch.stack([g for _, _, g in found])
scores = torch.tensor([0.81, 0.88])
labels = torch.zeros(len(boxes), dtype=torch.long)
print("\nIoU between the two partial detections:",
      round(box_iou(boxes[:1], boxes[1:]).max().item(), 3))
keep = batched_nms(boxes, scores, labels, 0.4)
print("NMS keeps  :", [[round(v, 1) for v in boxes[i].tolist()] for i in keep.tolist()])
merged = torch.stack([boxes[:, 0].min(), boxes[:, 1].min(),
                      boxes[:, 2].max(), boxes[:, 3].max()])
print("union merge:", [round(v.item(), 1) for v in merged])
print("true box   :", [round(v, 1) for v in obj.tolist()])
Output
resize longest side to  640: scale 0.160  a 30px object becomes  4.8px, a 12px object becomes  1.9px
resize longest side to 1024: scale 0.256  a 30px object becomes  7.7px, a 12px object becomes  3.1px
resize longest side to 1280: scale 0.320  a 30px object becomes  9.6px, a 12px object becomes  3.8px

grid cells covered by an object at each stride:
    object   stride 4   stride 8  stride 16  stride 32
       8x8       4.00       1.00       0.25       0.06
     16x16      16.00       4.00       1.00       0.25
     32x32      64.00      16.00       4.00       1.00
     64x64     256.00      64.00      16.00       4.00

best IoU between a perfectly centred standard anchor set and one object:
   8x8  object -> best IoU 0.063
  12x12 object -> best IoU 0.141
  16x16 object -> best IoU 0.250
  24x24 object -> best IoU 0.562
  32x32 object -> best IoU 1.000
  64x64 object -> best IoU 1.000

4000x3000 image, 640 tiles, 20 percent overlap: 48 tiles
  first four origins: [(0, 0), (512, 0), (1024, 0), (1536, 0)]
  each tile is fed at native resolution, so a 12px object stays 12px

the object at [1290, 500, 1330, 545] falls inside 2 tiles:
  tile (1024,   0)  local [266.0, 500.0, 306.0, 545.0]  -> global [1290.0, 500.0, 1330.0, 545.0]
  tile (1024, 512)  local [266.0, 0.0, 306.0, 33.0]  -> global [1290.0, 512.0, 1330.0, 545.0]

IoU between the two partial detections: 0.733
NMS keeps  : [[1290.0, 512.0, 1330.0, 545.0]]
union merge: [1290.0, 500.0, 1330.0, 545.0]
true box   : [1290.0, 500.0, 1330.0, 545.0]

Reading that output, block by block

A 30-pixel object becomes 4.8 pixels at input size 640. Under COCO's definition, anything below 32x32 pixels is "small". At this scale nearly everything in a drone photo is below that. No architectural change recovers information the resize destroyed.

Even at 1280 the same object is 9.6 pixels. Raising input resolution helps and does not solve it. That is the argument for tiling in one line.

An 8x8 object covers 0.06 cells at stride 32. Not one cell. Six hundredths of one. The deepest pyramid level has nothing to work with. Detectors for small objects add a finer level, or drop the coarse ones.

Anchor coverage is the quiet killer. A 12x12 object reaches a best IoU of 0.141 against a standard anchor set, even when perfectly centred on it. Every common assignment rule labels that anchor background. The object generates no positive training sample at all, so the model never learns it exists. This is a strong argument for anchor-free heads on small-object work; see anchor-free detection.

48 tiles for one image. That is 48 forward passes for one photo. Tiling is not free, and batching the tiles is the main way to make it tolerable.

The last block is the trap that people ship. The object straddles a seam. The tile at y=0 sees all of it; the tile at y=512 sees the bottom 33 pixels. NMS ranks by score, and the truncated detection happened to score higher. NMS kept the wrong box and lost 12 pixels off the top.

A union merge recovers the true box exactly. This is why SAHI offers non-maximum merging alongside suppression. When two detections of the same object are both truncated, picking one is wrong. Merging is right.

A workable tiled inference loop

tiled_inference.py
import torch
from torchvision.ops import batched_nms

def tiled_detect(detect_fn, image, tile=640, overlap=0.2, batch=8):
    """detect_fn(batch_of_tiles) -> list of (boxes_xyxy, scores, labels) per tile."""
    _, H, W = image.shape
    step = int(tile * (1 - overlap))
    xs = sorted({*range(0, max(W - tile, 0) + 1, step), max(W - tile, 0)})
    ys = sorted({*range(0, max(H - tile, 0) + 1, step), max(H - tile, 0)})
    origins = [(x, y) for y in ys for x in xs]

    all_boxes, all_scores, all_labels = [], [], []
    for i in range(0, len(origins), batch):
        chunk = origins[i:i + batch]
        crops = torch.stack([image[:, y:y + tile, x:x + tile] for x, y in chunk])
        for (x, y), (b, s, l) in zip(chunk, detect_fn(crops)):
            if len(b) == 0:
                continue
            b = b + torch.tensor([x, y, x, y], dtype=b.dtype)   # tile -> image coords
            all_boxes.append(b); all_scores.append(s); all_labels.append(l)

    if not all_boxes:
        return torch.zeros(0, 4), torch.zeros(0), torch.zeros(0, dtype=torch.long)
    boxes = torch.cat(all_boxes); scores = torch.cat(all_scores); labels = torch.cat(all_labels)
    keep = batched_nms(boxes, scores, labels, 0.5)
    return boxes[keep], scores[keep], labels[keep]

# A fake detector so this runs with no weights: it "finds" one box per tile.
def fake_detect(crops):
    n = crops.shape[0]
    return [(torch.tensor([[100., 100., 140., 145.]]),
             torch.tensor([0.9]),
             torch.tensor([0])) for _ in range(n)]

img = torch.zeros(3, 1500, 2000)
b, s, l = tiled_detect(fake_detect, img)
print("tiles processed :", len(sorted({*range(0, 2000 - 640 + 1, 512), 1360})) *
                            len(sorted({*range(0, 1500 - 640 + 1, 512), 860})))
print("detections kept :", len(b))
print("first box       :", [round(v, 1) for v in b[0].tolist()])
print("last box        :", [round(v, 1) for v in b[-1].tolist()])
Output
tiles processed : 12
detections kept : 12
first box       : [100.0, 100.0, 140.0, 145.0]
last box        : [1460.0, 960.0, 1500.0, 1005.0]

The fake detector returns the identical tile-local box for every tile. The 12 detections should therefore land at 12 different places, with no overlap. All 12 survived NMS, and the first and last boxes are far apart. That is the correctness check for the coordinate shift: identical local boxes must become distinct global boxes.

Swap fake_detect for a real model and the structure is unchanged. The sorted({*range(...), max(W - tile, 0)}) trick guarantees the right and bottom edges are covered. That holds even when the image size is not a multiple of the step.

The fixes, ranked by what they cost you

FixEffectCost
Measure object sizes after resizingnone, but it tells you what to doten minutes
Raise input resolutiondoubles object size for double the side lengtharound 4x compute
Tile the image (SAHI-style)keeps native resolutionone pass per tile
Add a finer pyramid levelmore cells for small objectslarge memory increase
Go anchor-freeremoves the anchor-IoU wallarchitecture change
Copy-paste small objects in trainingmore positive samplesdata pipeline work
Fine-tune on tiles as wellinference and training distributions matcha training run

The SAHI paper reports gains from sliced inference alone of 6.8, 5.1 and 5.3 AP. Those are for FCOS, VFNet and TOOD. Adding slicing-aided fine-tuning took the cumulative gains to 12.7, 13.4 and 14.5 points. Those numbers are from their benchmark, not a promise about yours. The ordering they imply is still worth taking seriously. Tile at inference first, then fine-tune on tiles.

Common mistakes

Tiling at inference but training on whole images. The model then sees a different object-size distribution at test time than it trained on. Fine-tune on tiles as well.

Overlap too small. Overlap must exceed the largest object you expect, or some object is truncated in every tile. 20 percent of the tile size is a reasonable default for genuinely small objects.

Merging with plain NMS. As shown above, this keeps whichever truncated box scored higher. Use a merge that takes the union for boxes touching a tile edge.

Forgetting the coordinate shift. Add the tile origin to every box. This bug produces boxes clustered in the top-left corner, which is at least easy to spot.

Judging with overall mAP. COCO reports AP-small, AP-medium and AP-large separately. Track AP-small, or your improvement will be invisible inside a number dominated by large objects.

Skipping the diagnosis. Print a histogram of your object widths in pixels after resizing. Do this before anything else on this page.

Try it yourself

Take the tiling function and set overlap=0.0. Re-run the seam test with the object at [1290., 500., 1330., 545.]. It is now truncated in every tile that sees it, and no merge recovers it. Then raise overlap until the object is complete somewhere. The value you find is a lower bound on the overlap your data needs.

What to learn next

Researcher — Mathematics and papers.

Why small objects are hard, in four separable terms

Information. After resizing by factor $s$, an object of side $d$ occupies $sd$ pixels. Detail below the new Nyquist limit is unrecoverable. This term is upstream of the model entirely.

Sampling. At stride $r$, an object of side $d$ covers $(d/r)^2$ feature positions. Below one position, no location is fully responsible for the object. The classification and regression heads then see a feature vector dominated by context.

Assignment. Under IoU assignment, the achievable IoU between a $d \times d$ object and an anchor of side $a$, perfectly centred, is $\min(d,a)^2 / \max(d,a)^2$. For $d = 12$ and $a = 32$ this is $144/1024 \approx 0.14$, below every common positive threshold. The object contributes no positive sample. This is the term the output block above measures.

Evaluation. IoU is scale-invariant in value but not in sensitivity to pixel error. Shifting a $d \times d$ box by $\delta$ pixels along one axis gives $\mathrm{IoU} = (d - \delta)/(d + \delta)$. For $d = 100, \delta = 5$ that is 0.905. For $d = 10, \delta = 5$ it is 0.333. One absolute localisation error moves a small object across the match threshold. The same error leaves a large one comfortably inside it.

That last term means AP-small is partly a measure of sub-pixel regression precision, not only of detection. Xu et al. (2022) propose the Normalized Wasserstein Distance as an assignment metric. It degrades gracefully on small boxes, where IoU falls off a cliff.

Slicing-aided inference

Akyon et al. (2022), SAHI, formalise the tiling pipeline: slice into overlapping patches, run any detector per patch, map boxes back, and merge. The method is detector-agnostic and needs no retraining, which is why it is the standard first intervention.

Their reported results are worth reading precisely. Sliced inference alone gave AP increases of 6.8, 5.1 and 5.3 points for FCOS, VFNet and TOOD. Adding slicing-aided fine-tuning, where the model is also trained on slices, produced cumulative increases of 12.7, 13.4 and 14.5 points. The fine-tuning term is roughly as large as the inference term, and it is the half most people skip.

Two implementation details decide whether it works.

Overlap must exceed the largest expected object, or some object is truncated in every slice.

Merging must not be plain NMS. Two truncated detections of one object are both wrong; greedy selection keeps one wrong box. Non-maximum merging replaces the cluster with the union of its members, which is correct when both members are partial views.

Architectural responses

  • Higher-resolution levels. Adding P2 (stride 4) or P1 (stride 2) multiplies feature-map area by 4 or 16. This is the dominant memory cost in high-resolution detectors.
  • Anchor-free heads remove the assignment term entirely, since positives are defined by containment rather than IoU. See anchor-free detection.
  • Scale-aware assignment. ATSS's per-object adaptive threshold naturally admits smaller objects than a fixed 0.5.
  • Alternative assignment metrics. NWD and dot-distance measures replace IoU where IoU is numerically unstable at small scales.
  • Query-based detectors with deformable attention sample specific points rather than pooling a region. That suits objects covering a fraction of a feature cell. YOLO26's small-target-aware label assignment is a recent example of the same concern in a convolutional head.

Data-side responses

Copy-paste augmentation (Kisantal et al., 2019, for small objects; Ghiasi et al., 2021, in general) pastes additional instances of small objects into training images. On COCO, small objects appear in a minority of images and are a small fraction of instances. Oversampling those images and duplicating instances within them both help.

Scale jitter with a wide range forces the model to see the same object at many sizes. Large-scale jitter is one of the more reliable gains in modern detection recipes.

Mosaic augmentation, standard in YOLO training, composes four images into one. That shrinks objects and raises the density of small instances per training image.

Measuring it properly

Report AP-small, AP-medium and AP-large separately. COCO defines the boundaries at area $< 32^2$ and $> 96^2$ pixels, in original image coordinates.

Two cautions. Those boundaries are absolute pixel areas defined for COCO's typical resolution. On 4000-pixel-wide imagery almost every object is "small" by that definition. The split stops being informative. Define your own boundaries from your own size histogram.

And report the size distribution after your resize, not from the annotation file. The gap between those two histograms is, in most projects, the entire problem.

Papers

What to learn next