Object Detection in Depth

Bounding box formats

A bounding box is four numbers around an object, and the four different ways of writing those numbers cause more detection bugs than any model choice.

On this page 8
  1. The short answer
  2. The analogy
  3. Why formats exist at all
  4. How it works
  5. What goes wrong
  6. Where you have already seen boxes
  7. Remember this
  8. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

A bounding box is a rectangle drawn around an object, written down as four numbers.

The analogy

Think about telling a friend where you parked the scooter in a crowded market. You could say "third row, fourth slot". You could say "twenty steps past the tea stall, then left". You could hand them a map with a circle on it.

All three directions lead to the same scooter. They are the same place, described in different languages. A bounding box format is the language you pick for writing a rectangle down.

Why formats exist at all

Nobody sat down and decided there should be four of them. They grew separately.

The Pascal VOC dataset wrote the corners of the rectangle. The COCO dataset wrote the top-left corner plus a width and a height. YOLO wrote the centre point plus a width and a height. It then divided everything by the picture size, so the numbers sit between zero and one.

Each choice made sense for the people who made it. Then everyone started mixing datasets and tools, and the trouble began.

How it works

Here is one rectangle, said four ways.

   the picture                          the same box, four languages
   +-----------------------+
   |                       |            corners     : left, top, right, bottom
   |     +---------+       |            corner+size : left, top, width, height
   |     |  car    |       |            centre+size : centre-x, centre-y, width, height
   |     +---------+       |            normalised  : the same, divided by picture size
   |                       |
   +-----------------------+

Nothing is lost between them. Every format holds the same rectangle. You can move between them and back again, and land on the exact numbers you started with.

What goes wrong

The numbers do not carry a label saying which language they are in.

Hand a tool the corners of a box when it expects a width and a height. It will not complain. It will happily read the third number as a width. Your boxes end up in the wrong place. Your training loss looks strange. Nothing in the error messages points at the cause.

This is the most common bug in detection work. It is confusing for almost everyone the first time. Read that paragraph twice, because it will save you a weekend.

Where you have already seen boxes

  • The green rectangle around a face when your phone camera focuses.
  • The box a toll booth camera draws around a number plate.
  • The outlines an online shop puts around clothes in a photo.
  • The rectangle a video call app uses to keep your face centred.

Remember this

  • A bounding box is a rectangle around an object, stored as four numbers.
  • Four common formats exist, and the numbers alone do not tell you which one you have.
  • Converting between formats loses nothing, so pick one for your own code and convert at the edges.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install torch torchvision

Written against torch 2.5.1 and torchvision 0.20.1. torchvision.ops.box_convert has been stable for many releases, so older versions behave the same.

The four formats, converted and checked

box_formats.py
import torch
from torchvision.ops import box_convert, clip_boxes_to_image

# One box around a scooter in a 640x480 photo, written the four common ways.
xyxy = torch.tensor([[100., 150., 260., 390.]])          # left, top, right, bottom

xywh   = box_convert(xyxy, in_fmt="xyxy", out_fmt="xywh")    # left, top, width, height
cxcywh = box_convert(xyxy, in_fmt="xyxy", out_fmt="cxcywh")  # centre x, centre y, w, h

print("xyxy   ", xyxy.tolist())
print("xywh   ", xywh.tolist())
print("cxcywh ", cxcywh.tolist())

W, H = 640, 480
norm = cxcywh / torch.tensor([W, H, W, H])                 # YOLO label file numbers
print("yolo   ", [round(v, 6) for v in norm[0].tolist()])

# Round trip: every format holds the same information.
back = box_convert(cxcywh, in_fmt="cxcywh", out_fmt="xyxy")
print("round trip exact:", torch.equal(back, xyxy))

# The bug that eats a week: flipping an image without flipping the box.
flipped_wrong = xyxy.clone()
flipped_wrong[:, [0, 2]] = W - flipped_wrong[:, [0, 2]]    # forgot to swap left and right
print("\nafter a naive horizontal flip:", flipped_wrong.tolist())
print("width now:", (flipped_wrong[:, 2] - flipped_wrong[:, 0]).item())

flipped_right = xyxy.clone()
flipped_right[:, [0, 2]] = W - xyxy[:, [2, 0]]             # swap, then subtract
print("done properly:                ", flipped_right.tolist())
print("width now:", (flipped_right[:, 2] - flipped_right[:, 0]).item())

# Boxes that run off the edge of the image after a crop or a resize.
ragged = torch.tensor([[-20., 30., 700., 500.]])
print("\nclipped to the image:", clip_boxes_to_image(ragged, (H, W)).tolist())
Output
xyxy    [[100.0, 150.0, 260.0, 390.0]]
xywh    [[100.0, 150.0, 160.0, 240.0]]
cxcywh  [[180.0, 270.0, 160.0, 240.0]]
yolo    [0.28125, 0.5625, 0.25, 0.5]
round trip exact: True

after a naive horizontal flip: [[540.0, 150.0, 380.0, 390.0]]
width now: -160.0
done properly:                 [[380.0, 150.0, 540.0, 390.0]]
width now: 160.0

clipped to the image: [[0.0, 30.0, 640.0, 480.0]]

Reading that output

The round trip is exact. torch.equal returned True, not "close enough". Conversion is arithmetic on floats that divide cleanly here, so nothing drifts. With awkward sizes you can pick up tiny floating-point error, so compare with torch.allclose in tests rather than torch.equal.

The naive flip produced a width of minus 160. The right edge ended up to the left of the left edge. Almost no library raises an error for this. box_iou returns zero, your loss goes strange, and you spend the afternoon suspecting the model.

The fix is one line: swap the two columns while subtracting, W - xyxy[:, [2, 0]]. Anything that mirrors an image has this trap, including horizontal flip, vertical flip, transpose and rotation.

Clipping changed three of the four numbers. After a random crop or a resize, boxes routinely poke outside the picture. clip_boxes_to_image pulls them back. Do this before computing loss, and drop boxes that end up with zero area afterwards.

The dataset conventions you will actually meet

FormatFieldsCoordinatesUsed by
xyxyleft, top, right, bottomabsolute pixelsPascal VOC, torchvision, most PyTorch code
xywhleft, top, width, heightabsolute pixelsCOCO JSON annotations
cxcywhcentre x, centre y, width, heightabsolute pixelsDETR internals
YOLO txtclass, centre x, centre y, width, heightdivided by image sizeUltralytics label files

Two extra traps sit inside that table.

Pascal VOC XML files are 1-indexed: the top-left pixel is (1, 1), not (0, 0). COCO is 0-indexed. Mixing them shifts every box by one pixel, which is invisible on a car and significant on a 12-pixel bird.

A YOLO label file has no image size in it. The numbers are fractions of width and height. Resize the image, and the label file stays correct without being touched. That is the whole reason for the normalisation.

Converting a COCO annotation by hand

coco_to_xyxy.py
import torch
from torchvision.ops import box_convert

# A slice of a real COCO-style annotations block.
coco = {
    "images": [{"id": 7, "width": 640, "height": 426, "file_name": "market.jpg"}],
    "annotations": [
        {"image_id": 7, "category_id": 3, "bbox": [124.5, 200.1, 88.0, 132.4]},
        {"image_id": 7, "category_id": 1, "bbox": [401.0,  90.0, 60.5, 210.0]},
    ],
}

boxes = torch.tensor([a["bbox"] for a in coco["annotations"]])
labels = torch.tensor([a["category_id"] for a in coco["annotations"]])

xyxy = box_convert(boxes, in_fmt="xywh", out_fmt="xyxy")
print("torchvision target format:")
print(xyxy)
print("labels:", labels.tolist())

img = coco["images"][0]
wh = torch.tensor([img["width"], img["height"], img["width"], img["height"]])
yolo = box_convert(boxes, in_fmt="xywh", out_fmt="cxcywh") / wh
print("\nthe same boxes as YOLO label lines:")
for cls, row in zip(labels.tolist(), yolo.tolist()):
    print(f"{cls} " + " ".join(f"{v:.6f}" for v in row))
Output
torchvision target format:
tensor([[124.5000, 200.1000, 212.5000, 332.5000],
        [401.0000,  90.0000, 461.5000, 300.0000]])
labels: [3, 1]

the same boxes as YOLO label lines:
3 0.263281 0.625117 0.137500 0.310798
1 0.673828 0.457746 0.094531 0.492958

Common mistakes

Assuming bbox means corners. In COCO JSON it never does. It is left, top, width, height. Converting COCO to torchvision without box_convert produces boxes that are far too small and sit in the wrong place.

Losing the image size before normalising. Normalised YOLO numbers need the original width and height. If your pipeline resizes first and reads the size second, every label is silently wrong. Read the size from the annotation record, not from the tensor you happen to be holding.

Forgetting dtype. Integer pixel coordinates divided by an integer image size floor to zero in some code paths. Cast boxes to float32 when you load them, once, at the edge.

Silently keeping degenerate boxes. After cropping you can get a box with zero width. Filter with torchvision.ops.remove_small_boxes before it reaches the loss.

Mixing 0-indexed and 1-indexed sources. Pascal VOC XML starts at one. Subtract one from xmin and ymin when you convert.

Try it yourself

Write a function to_yolo(xyxy, width, height) and its inverse from_yolo. Then feed a box through both and assert the result matches the input with torch.allclose. Now break it on purpose: pass the height where the width belongs, and see whether your assertion catches it. If it does not, your test is not testing anything.

What to learn next

Researcher — Mathematics and papers.

The representation, formally

A box is a point in a four-dimensional space, but the parameterisation you choose changes the loss surface. Let a box be $b = (x_1, y_1, x_2, y_2)$ with $x_1 < x_2$ and $y_1 < y_2$.

The centre-size form is a smooth bijection on the interior of the valid region:

$$ c_x = \frac{x_1 + x_2}{2}, \quad c_y = \frac{y_1 + y_2}{2}, \quad w = x_2 - x_1, \quad h = y_2 - y_1 $$

The constraint $x_1 < x_2$ becomes $w > 0$. This matters because a network predicting $w$ directly can emit a negative number, whereas a network predicting $\log w$ cannot. Every anchor-based detector since Girshick et al. (2014) therefore regresses in log-space:

$$ t_x = \frac{c_x - c_x^a}{w^a}, \quad t_y = \frac{c_y - c_y^a}{h^a}, \quad t_w = \log \frac{w}{w^a}, \quad t_h = \log \frac{h}{h^a} $$

Where the superscript $a$ denotes the anchor. Dividing the translation terms by the anchor size makes the targets scale-invariant. A 10-pixel error on a 400-pixel truck and the same error on a 40-pixel bird stop being equally bad.

Why the parameterisation changes what the model learns

Regressing $(x_1, y_1, x_2, y_2)$ with an $L_1$ loss treats the four coordinates as independent. They are not. Two of them jointly determine width. An error in $x_1$ is a location error or a size error, depending on what $x_2$ did.

This is a large part of why IoU-family losses replaced coordinate losses; see intersection over union. Rezatofighi et al. (2019) make the point directly. Minimising an $L_n$ coordinate loss is not equivalent to maximising IoU. Two pairs of boxes with the same $L_2$ error can have very different IoU.

Discretisation and the half-pixel question

A pixel is an area, not a point. Where the centre of the top-left pixel sits decides where a box edge falls. Libraries disagree about it.

  • COCO uses 0-indexed continuous coordinates; box area is $w \cdot h$.
  • Pascal VOC XML is 1-indexed and treats coordinates as inclusive pixel indices, so the width is $x_2 - x_1 + 1$.
  • torchvision uses 0-indexed continuous coordinates with width $x_2 - x_1$.

The $+1$ convention was standard in early detection code and survives in some evaluation scripts. On a 500-pixel object the difference is a fifth of one percent of the area. On an 8-pixel object it is around a quarter of the area. That can move the IoU across a match threshold. The same half-pixel reasoning drives aligned=True in roi_align; see RoI pooling and RoI align.

Rotated and beyond-rectangle representations

The axis-aligned rectangle is a modelling assumption, and a poor one for aerial imagery, text and shelf products.

  • Oriented boxes add an angle, $(c_x, c_y, w, h, \theta)$. The angle is periodic, so a naive $L_1$ loss has a discontinuity at the wrap-around. Fixes include the Gaussian-Wasserstein-distance loss of Yang et al. (2021) and representing the box by its corner set.
  • Quadrilaterals store eight numbers and avoid the periodicity problem, at the cost of an ordering ambiguity among the corners.
  • Distance-to-edge, $(l, t, r, b)$ from a point, is what anchor-free detectors regress. It has no negative-size failure mode when passed through a non-negative activation. See anchor-free detection.
  • Masks discard the rectangle entirely. See image segmentation.

Papers

What to learn next

What to learn next

These follow on from what you just read.

  • Object Detection in Depth

    Intersection over union

    IoU is the overlap between two boxes divided by the area they cover together, and it is the number that decides whether a detection counts as correct.

  • Object Detection in Depth

    Anchor boxes

    Anchor boxes are a fixed grid of guessed rectangles that a detector nudges into place, which turns finding objects into correcting guesses.

  • Object Detection in Depth

    Non-maximum suppression

    NMS keeps the highest-scoring box in a cluster and deletes its overlapping neighbours, which is how thousands of raw detections become a handful of answers.