Vision Datasets and Annotation

Converting between dataset formats

Every format writes boxes differently, and the conversion bugs are silent, so always convert both ways and check you got the original back.

On this page 10
  1. The short answer
  2. The analogy
  3. The three you will meet
  4. Why converting is riskier than it looks
  5. The check that catches everything
  6. The other thing to check
  7. The counting trap
  8. Where you have seen this
  9. Remember this
  10. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

Different tools write the same box in different ways, and converting between them is where quiet mistakes live.

The analogy

Think about giving someone directions to your house.

One person says "third lane on the left, second gate". Another gives the plot number. A third sends a map pin.

All three describe the same house. Hand one to somebody expecting another and they end up somewhere else, confidently.

Box formats work the same way. Same box, three descriptions.

The three you will meet

COCO describes a box as left edge, top edge, then width and height, counted in pixels.

YOLO describes it as the centre point, then width and height, all as fractions of the picture size. So a box in the middle of any picture is the same numbers, whatever the picture's size.

Pascal VOC describes it as the two corners: left, top, right, bottom, in pixels.

   the same box, three ways

   COCO   :  left 100, top 80, width 40, height 35
   VOC    :  left 100, top 80, right 140, bottom 115
   YOLO   :  centre across 0.1875, centre down 0.2083,
             width 0.0625, height 0.0729

Why converting is riskier than it looks

Nothing crashes. That is the whole problem.

Convert wrongly and you get boxes. They are in roughly sensible places. Training runs. The score is bad and nobody knows why, because there is nothing to see in the logs.

So you have to check deliberately. Nobody will tell you.

The check that catches everything

Convert forwards. Convert back. Compare with what you started with.

If you get the same numbers, your two functions agree. If you do not, you have found the bug in five seconds instead of five days.

This is called a round trip, and it should be in your code, not in your head.

The other thing to check

Draw some boxes on some pictures and look at them.

Twenty images is enough. Your eyes catch an upside-down or mirrored coordinate system instantly, and no amount of number-checking substitutes for it.

The counting trap

There is one more thing that trips people, and it is not really about formats.

A shelf's leftmost position is called position one. The rightmost is position twenty. How many positions are there? Twenty. But twenty minus one is nineteen.

Pascal VOC counts positions starting from one and includes both ends. So its width is the difference plus one. Miss that and every single box shrinks by a pixel on each side.

On a large object, nobody notices. On a small one, it is a real loss, and the next section measures it.

Where you have seen this

  • Downloading a dataset that will not load in your chosen framework.
  • A model that trains fine and scores terribly for no visible reason.
  • Two tools reporting different object counts for the same file.

Remember this

  • The same box has different numbers in every format.
  • Convert both ways and compare with the original. Always.
  • Look at drawn boxes on real pictures before trusting any converter.

What to learn next

Developer — Code and libraries.

Setup

bash
python3 --version     # standard library only

The conversions, with a round-trip test and the off-by-one

convert.py
import xml.etree.ElementTree as ET

W, H = 640, 480
COCO_CATS = [{"id": 1, "name": "helmet"}, {"id": 2, "name": "head"}]   # COCO ids start anywhere
YOLO_NAMES = ["helmet", "head"]                                        # YOLO ids start at 0
coco_boxes = [(1, [100.0, 80.0, 40.0, 35.0]), (2, [0.0, 0.0, 640.0, 480.0]),
              (1, [612.5, 455.25, 27.5, 24.75])]

def coco_to_yolo(cat_id, bbox, w, h):
    x, y, bw, bh = bbox
    idx = YOLO_NAMES.index(next(c["name"] for c in COCO_CATS if c["id"] == cat_id))
    return idx, (x + bw / 2) / w, (y + bh / 2) / h, bw / w, bh / h

def yolo_to_coco(idx, cx, cy, nw, nh, w, h):
    bw, bh = nw * w, nh * h
    cat_id = next(c["id"] for c in COCO_CATS if c["name"] == YOLO_NAMES[idx])
    return cat_id, [cx * w - bw / 2, cy * h - bh / 2, bw, bh]

print("COCO -> YOLO -> COCO")
for cid, b in coco_boxes:
    y = coco_to_yolo(cid, b, W, H)
    back = yolo_to_coco(*y, W, H)
    line = f"{y[0]} " + " ".join(f"{v:.6f}" for v in y[1:])
    err = max(abs(p - q) for p, q in zip(b, back[1]))
    print(f"  {str(b):32s} -> {line:38s} -> max error {err:.2e} px")

print("\nsix decimal places is the usual YOLO text precision. What that costs:")
for cid, b in coco_boxes:
    y = coco_to_yolo(cid, b, W, H)
    rounded = (y[0],) + tuple(round(v, 6) for v in y[1:])
    back = yolo_to_coco(*rounded, W, H)
    print(f"  {str(b):32s} -> {max(abs(p-q) for p,q in zip(b, back[1])):.2e} px after rounding")

# ---- Pascal VOC, and the off-by-one that quietly shrinks every box ----
xml = """<annotation><size><width>640</width><height>480</height></size>
<object><name>helmet</name><bndbox><xmin>101</xmin><ymin>81</ymin>
<xmax>140</xmax><ymax>115</ymax></bndbox></object></annotation>"""
o = ET.fromstring(xml).find("object").find("bndbox")
xmin, ymin, xmax, ymax = (int(o.find(t).text) for t in ("xmin", "ymin", "xmax", "ymax"))
print(f"\nVOC box  xmin={xmin} ymin={ymin} xmax={xmax} ymax={ymax}")
print(f"  naive conversion : bbox = [{xmin}, {ymin}, {xmax-xmin}, {ymax-ymin}]  area {(xmax-xmin)*(ymax-ymin)}")
print(f"  VOC devkit rule  : bbox = [{xmin-1}, {ymin-1}, {xmax-xmin+1}, {ymax-ymin+1}]  "
      f"area {(xmax-xmin+1)*(ymax-ymin+1)}")
print("  The original VOC coordinates are 1-indexed and inclusive of both endpoints.")
print("  Getting this wrong shrinks every box by one pixel on each axis. On a 20 px")
print(f"  object that is a {1-(19*19)/(20*20):.0%} area error, which moves your mAP.")
Output
COCO -> YOLO -> COCO
  [100.0, 80.0, 40.0, 35.0]        -> 0 0.187500 0.203125 0.062500 0.072917  -> max error 0.00e+00 px
  [0.0, 0.0, 640.0, 480.0]         -> 1 0.500000 0.500000 1.000000 1.000000  -> max error 0.00e+00 px
  [612.5, 455.25, 27.5, 24.75]     -> 0 0.978516 0.974219 0.042969 0.051562  -> max error 0.00e+00 px

six decimal places is the usual YOLO text precision. What that costs:
  [100.0, 80.0, 40.0, 35.0]        -> 1.60e-04 px after rounding
  [0.0, 0.0, 640.0, 480.0]         -> 0.00e+00 px after rounding
  [612.5, 455.25, 27.5, 24.75]     -> 2.40e-04 px after rounding

VOC box  xmin=101 ymin=81 xmax=140 ymax=115
  naive conversion : bbox = [101, 81, 39, 34]  area 1326
  VOC devkit rule  : bbox = [100, 80, 40, 35]  area 1400
  The original VOC coordinates are 1-indexed and inclusive of both endpoints.
  Getting this wrong shrinks every box by one pixel on each axis. On a 20 px
  object that is a 10% area error, which moves your mAP.

Reading the output

The round trip is exact in memory. All three test boxes return 0.00e+00 error, including a full-image box and one with fractional coordinates. That test is four lines. It separates a converter you trust from one you hope about.

Writing to a text file costs about two ten-thousandths of a pixel. After rounding to six decimals, the worst case is 2.40e-04 px. Sub-thousandth-pixel error is irrelevant for any IoU threshold, so the usual worry about YOLO's normalised precision is misplaced. The real risks are elsewhere.

The VOC off-by-one is a 5 percent linear error and a 10 percent area error on a 20-pixel object. Naive subtraction gives area 1326; the devkit rule gives 1400.

The rule comes from the original Pascal VOC devkit. Its coordinates are 1-indexed with both endpoints inclusive, so the true extent is xmax - xmin + 1. Many datasets distributed as "VOC format" are 0-indexed and exclusive, following the COCO convention despite the file layout. There is no way to tell from the XML.

The only reliable resolution is to draw the boxes on the images and look. A one-pixel systematic offset is visible on a small object when you zoom in.

What conversion silently discards

Round-tripping coordinates is the easy half. The information loss is the half that hurts.

Going toYou lose
YOLOSegmentation, iscrowd, area, per-annotation attributes, licences, image metadata
YOLOImages with zero annotations, unless you write empty .txt files for them
VOCSegmentation (it lives in separate PNG files), crowd regions
COCO from VOCThe difficult and truncated flags, unless you invent custom fields

The zero-annotation case is the one that costs people real accuracy. In COCO, an image listed with no annotations is an explicit negative: it was reviewed and contains nothing. Most COCO-to-YOLO converters write no file for it, and most YOLO loaders then skip the image entirely. Your hard negatives vanish and your false-positive rate rises.

Write an empty .txt file for every reviewed-and-empty image. Check your loader treats that as a background image, not a missing label.

A conversion checklist

  1. Write both directions. Not one. The second one is your test.
  2. Assert the round trip over every annotation in the dataset, not a sample.
  3. Count. Number of images and annotations before and after must match, or the difference must be explained.
  4. Compare class histograms before and after. A shifted class map shows up here instantly.
  5. Render 20 random images with boxes drawn, and look at them.
  6. Check the extremes: a box touching each edge, a full-image box, the smallest box, and an empty image.
  7. Record the source dataset version in the output, so a later mismatch is traceable.

Steps 1 to 4 are automatic and belong in a test file. Steps 5 and 6 need eyes, once, per dataset.

Common mistakes

Class index off by one. COCO ids need not start at 0 or be contiguous; YOLO indices must. Build the mapping explicitly from the category list and never assume category_id - 1.

Using stale image dimensions. YOLO coordinates are normalised, so the width and height used for conversion must be the real ones. If images were resized after annotation, the COCO width/height fields are wrong and every conversion inherits it.

Clamping too late. Boxes that spill past the image edge produce normalised values above 1.0. Clamp during conversion, and count how many you clamped. A large count means an upstream annotation problem, not a conversion problem.

Converting through an intermediate you did not test. COCO to VOC to YOLO compounds two possible errors. Convert directly.

Assuming a folder called VOC2007 follows the devkit convention. Verify visually.

Try it yourself

Add a segmentation polygon to one of the COCO boxes. Convert to YOLO and back, then assert the round trip. It will fail, because YOLO detection format has nowhere to put it. Make the converter raise an error rather than silently dropping it. That single change converts a class of invisible data loss into a loud failure.

What to learn next

Researcher — Mathematics and papers.

Coordinate conventions, stated exactly

Three independent choices produce most of the confusion, and they are orthogonal, so eight combinations exist in the wild.

Index origin. Is the first pixel at index 0 or 1? COCO and YOLO use 0; the original Pascal VOC devkit uses 1.

Endpoint inclusivity. Does xmax name the last pixel inside the box, or the first pixel outside it? The VOC devkit is inclusive, giving extent $x_{\max} - x_{\min} + 1$. Most modern re-releases are exclusive, giving $x_{\max} - x_{\min}$.

Pixel model. Is a coordinate the centre of a pixel or a corner of the grid? Under the centre model, a box from $x=0$ to $x=1$ spans two pixel centres. Its extent is 2 in pixel counts, and 1 in continuous coordinates. Detection evaluation is continuous, so IoU arithmetic assumes the corner model, while annotation tools frequently report the centre model.

The combination that matters most is index origin with inclusivity, since they can cancel. A 1-indexed inclusive box $[101, 140]$ and a 0-indexed exclusive box $[100, 140]$ describe the identical region. That is why the naive conversion in the output above is off by exactly one on each axis. There is a second, equally common bug. Converting a 1-indexed inclusive box by subtracting 1 from the origin, without adding 1 to the extent.

The error is systematic, not random

A one-pixel shrink on every side is not noise. For a square object of side $s$, IoU between the true box and the shrunk one is:

$$ \mathrm{IoU} = \frac{(s-2)^2}{s^2} $$

giving 0.64 at $s = 10$, 0.81 at $s = 20$, and 0.96 at $s = 100$. COCO's primary AP averages over IoU thresholds from 0.50 to 0.95. A systematic shrink removes small objects from the high-threshold bins, and leaves large objects untouched. The visible symptom is a plausible-looking $\mathrm{AP}_L$ with a collapsed $\mathrm{AP}_S$. That is easily misread as a small-object modelling problem, rather than a coordinate bug.

Any systematic bias in coordinates has this signature: it appears as a size-dependent metric distortion. When $\mathrm{AP}_S$ is anomalously low, check the conversion before changing the architecture.

Federated and partial annotation

Format conversion assumes every file expresses the same thing, and beyond simple detection it does not.

Open Images records verified negatives explicitly. For a given image and class, the label may be present, absent, or unverified. Converting it to COCO collapses the last two. Unverified classes become implicit negatives, and any detection of them is a false positive.

LVIS annotates each category exhaustively only within that category's own image set. Evaluating LVIS with a COCO-style protocol counts detections in images where the class was never checked. Large phantom false-positive counts follow.

Panoptic formats carry one label per pixel, with thing instances and stuff regions in a shared id space. There is no lossless mapping to instance-only or semantic-only formats in either direction.

The generalisation is worth stating. Conversion is lossless only when the target format can express the source's annotation completeness semantics. Most cannot. A converter should assert this rather than assume it.

Practical resolution of ambiguity

When the convention is undocumented, three empirical tests distinguish the cases without guessing.

Look for coordinate 0. If no annotation in a large dataset has xmin = 0, the file is likely 1-indexed.

Look for xmax equal to image width. Under a 0-indexed exclusive convention, xmax == width is legal and common. Under 1-indexed inclusive, xmax == width is also legal, but xmax == width + 1 never appears in either, so this test alone is weak.

Render and inspect. The definitive test. A systematic half-pixel or one-pixel offset is visible on a magnified crop of a small object. No statistical test on the numbers replaces looking.

Record the answer in the dataset's datasheet, in the sense of Gebru et al. (2021), so the next person does not repeat the investigation.

References

What to learn next