Vision Datasets and Annotation
The COCO dataset format
COCO is one JSON file listing images, categories and annotations, and knowing its exact field meanings prevents most silent dataset bugs.
- 14 min read
- 3 reading levels
- Updated
Read these first
On this page 9
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
COCO is one text file listing your images, your class names, and every mark drawn on them.
The analogy
Think about a school register. It has three lists that refer to each other.
A list of pupils, each with a roll number. A list of subjects, each with a code. And a list of entries saying "roll number twelve, subject code four, marks obtained".
Nothing is written twice. The pupil's name appears once. The subject name appears once. Everything else points at them by number.
COCO is exactly that register, for pictures.
The three lists
Images. Every picture, with a number, a filename, and how wide and tall it is.
Categories. Every class you can label, with a number and a name.
Annotations. Every single mark. Each one says which image it is on, which category it is, and where.
images: 1 -> site_001.jpg (640 wide, 480 tall)
2 -> site_002.jpg
categories: 1 -> helmet
2 -> head
annotations: on image 1, category 1, box at ...
on image 1, category 2, box at ...
on image 2, category 1, box at ...The one detail that trips everybody
A box in COCO is written as four numbers. Most people assume they are the two corners: left, top, right, bottom.
They are not. They are left, top, width, height.
Getting this wrong does not crash anything. It produces boxes in roughly the right place, with roughly the wrong size. The model then trains and performs badly, for reasons nobody can find.
Read the four numbers as left, top, width, height. Every time.
Why the file is so widely used
COCO was a large public dataset released around 2014, and its file layout came with it. Enough tools supported it that supporting it became the default.
So it is not the best possible format. It is the one everything can read, which is usually more useful.
Outlines, not only boxes
COCO also stores outlines, and it stores them two ways.
For a normal object, it stores the corner points of a shape drawn around it, like joining dots.
For a big messy region, listing corner points is hopeless. Think of a crowd of people, or a pile of gravel. So there is a second way. Walk through the picture. Record how many pixels in a row sit outside the shape, then how many sit inside, then outside again. Long runs compress to short numbers.
That second method is why a mask covering half a picture takes a handful of numbers, not thousands.
Where you have seen this
- Almost every object-detection tutorial you will find.
- Model leaderboards, which are usually reported on the COCO dataset.
- Annotation tools, which nearly all offer "export as COCO".
Remember this
- Three cross-referenced lists: images, categories, annotations.
- A box is left, top, width, height. Not two corners.
- Outlines come as joined dots or as a run-length code for large regions.
What to learn next
- Converting between dataset formats — moving COCO to YOLO and back without losing anything.
- Object detection — the task this format was designed for.
- Image segmentation — what the polygons and RLE masks are for.
Developer — Code and libraries.
Setup
pip install numpy==1.26.4Building a valid COCO file, validating it, and encoding a mask
import numpy as np
# A complete, valid COCO detection file. Five top-level keys; only three are required
# by the tooling ("images", "annotations", "categories").
coco = {
"info": {"description": "helmet demo", "version": "1.0", "year": 2026},
"licenses": [{"id": 1, "name": "CC BY 4.0", "url": "https://creativecommons.org/licenses/by/4.0/"}],
"images": [
{"id": 1, "file_name": "site_001.jpg", "width": 640, "height": 480, "license": 1},
{"id": 2, "file_name": "site_002.jpg", "width": 640, "height": 480, "license": 1},
],
"categories": [{"id": 1, "name": "helmet", "supercategory": "ppe"},
{"id": 2, "name": "head", "supercategory": "person"}],
"annotations": [
{"id": 1, "image_id": 1, "category_id": 1,
"bbox": [100.0, 80.0, 40.0, 35.0], # [x, y, WIDTH, HEIGHT] - not [x1,y1,x2,y2]
"area": 1400.0, "iscrowd": 0,
"segmentation": [[100, 80, 140, 80, 140, 115, 100, 115]]}, # one flat [x1,y1,x2,y2,...] list
{"id": 2, "image_id": 1, "category_id": 2,
"bbox": [98.0, 78.0, 44.0, 44.0], "area": 1936.0, "iscrowd": 0, "segmentation": []},
],
}
for k, v in coco["annotations"][0].items():
print(f"{k:14s} {v}")
# Validation you should run before every training job.
img_ids = {i["id"] for i in coco["images"]}
cat_ids = {c["id"] for c in coco["categories"]}
for a in coco["annotations"]:
assert a["image_id"] in img_ids, a
assert a["category_id"] in cat_ids, a
x, y, w, h = a["bbox"]
im = next(i for i in coco["images"] if i["id"] == a["image_id"])
assert 0 <= x and 0 <= y and x + w <= im["width"] and y + h <= im["height"], a
assert abs(a["area"] - w * h) < 1e-6 or a["segmentation"], a
print("\nvalidation passed\n")
# ---- COCO's RLE, implemented from its own definition ----
# "given M=[0 0 1 1 1 0 1] the RLE counts would be [2 3 1 1]", and the mask is
# flattened in COLUMN-major (Fortran) order, always starting the count with zeros.
def rle_encode(mask):
flat = np.asfortranarray(mask).ravel(order="F")
counts, last, run = [], 0, 0
for v in flat:
if v == last: run += 1
else: counts.append(run); last, run = v, 1
counts.append(run)
return {"size": list(mask.shape), "counts": counts}
def rle_decode(rle):
h, w = rle["size"]
out, val, i = np.zeros(h * w, np.uint8), 0, 0
for c in rle["counts"]:
out[i:i + c] = val; i += c; val ^= 1
return out.reshape((h, w), order="F")
print("the example from the COCO source itself:")
print(" M = [0 0 1 1 1 0 1] -> counts", rle_encode(np.array([[0,0,1,1,1,0,1]]).T)["counts"])
mask = np.zeros((6, 8), np.uint8); mask[2:5, 3:6] = 1
r = rle_encode(mask)
print("\n3x3 square inside a 6x8 mask")
print(" size ", r["size"])
print(" counts", r["counts"], f" ({len(r['counts'])} numbers instead of {mask.size} pixels)")
print(" round-trip identical:", np.array_equal(rle_decode(r), mask))
print(" area from RLE:", sum(r["counts"][1::2]), " area from the mask:", int(mask.sum()))id 1 image_id 1 category_id 1 bbox [100.0, 80.0, 40.0, 35.0] area 1400.0 iscrowd 0 segmentation [[100, 80, 140, 80, 140, 115, 100, 115]] validation passed the example from the COCO source itself: M = [0 0 1 1 1 0 1] -> counts [2, 3, 1, 1] 3x3 square inside a 6x8 mask size [6, 8] counts [20, 3, 3, 3, 3, 3, 13] (7 numbers instead of 48 pixels) round-trip identical: True area from RLE: 9 area from the mask: 9
Every field, and what it actually means
bbox: [x, y, width, height], absolute pixels, floats, 0-indexed. The pycocotools source unpacks it as [bbox_x, bbox_y, bbox_w, bbox_h] = ann['bbox']. The mask utilities document boxes as [x y w h], with bbox=[0 0 1 1] enclosing the first pixel. Origin is the top-left corner.
area: the area in pixels. For a box-only annotation, width * height. For a segmented one, the mask area, which is smaller than the box area. This is not decoration. The COCO evaluator uses area to split results into small, medium and large objects. The thresholds are 32-squared and 96-squared pixels. Set it from the box when you have a mask and your small-object metrics will be wrong.
iscrowd: 0 or 1. A crowd region is a single annotation covering many instances that were not separated. The evaluator handles them differently. For a crowd ground truth, a detection may match any sub-region. The score becomes intersection over detection area, not intersection over union. That is the modified criterion documented in the COCO mask utilities. Crowd regions are always stored as RLE, never as polygons.
segmentation: two possible types, distinguished at runtime. A list means polygons, each a flat [x1, y1, x2, y2, ...], and one annotation may hold several for a disconnected object. A dict with size and counts means RLE. Code that reads COCO segmentation must branch on the Python type, which pycocotools does explicitly.
categories: ids need not start at 1 and need not be contiguous. The 80-class COCO detection set uses ids running to 90, with gaps. Every training framework maps them to contiguous indices internally. Getting that mapping wrong is a classic silent bug. The model trains, and every class name is shifted.
images: width and height must match the real files. They are used for normalisation in almost every conversion. Resize your images after labelling without updating these and every box moves.
The RLE, read carefully
The counts always start with a run of zeros. So [2, 3, 1, 1] for [0 0 1 1 1 0 1] reads as two zeros, three ones, one zero, one one. If the mask starts with a 1, the first count is 0. The COCO source gives exactly this example, and gives [0 6 1] for [1 1 1 1 1 1 0].
The flattening is column-major. This is the detail that breaks hand-written implementations, since NumPy defaults to row-major. In the output above, the 3-by-3 square inside a 6-by-8 mask gives [20, 3, 3, 3, 3, 3, 13]. That is twenty zeros to reach column 3. Then alternating runs of three, as each of columns 3, 4 and 5 crosses the square. Then thirteen zeros. Read down the columns and it is obvious. Read across the rows and it is nonsense.
Area comes free. Sum the odd-indexed counts, since those are the runs of ones. 9, matching the mask. Most set operations on RLE work directly on the counts, without decoding. That is why the format is fast as well as small.
The real pycocotools compresses further. It stores the counts with LEB128 variable-length encoding. So counts arrives as a byte string, not a list of integers. Both forms are valid; frPyObjects converts an uncompressed one.
Common mistakes
Treating bbox as [x1, y1, x2, y2]. The most common COCO bug there is. Symptom: boxes in the right place with wrong sizes, and mAP around 0.1 with no error message.
Copying area from the box when a mask exists. Corrupts the small, medium and large breakdown, which is often the most informative part of a detection report.
Ignoring iscrowd. Training on crowd regions as if they were single objects teaches the detector to emit enormous boxes. Most pipelines filter them out for training and keep them for evaluation.
Assuming category ids are contiguous from zero. They are not, in real COCO. Build the mapping explicitly.
Producing float ids. JSON has no integer type distinction, so a careless writer emits 1.0 and a strict reader rejects it.
Not recording images with no annotations. COCO can hold an image with zero annotations, and that is meaningful. It says the image was reviewed and is empty. Many converters drop such images entirely and your negatives disappear.
Try it yourself
Encode a mask with two separate blobs and check that the counts alternate correctly across both. Then write the polygon-to-mask direction. Fill a polygon into a boolean array. Confirm your RLE of it matches the area you expect. That pair of functions is the whole segmentation-format toolkit.
What to learn next
- Converting between dataset formats — moving COCO to YOLO and back without losing anything.
- Object detection — the task this format was designed for.
- Image segmentation — what the polygons and RLE masks are for.
Researcher — Mathematics and papers.
The format as a specification of the evaluation protocol
COCO's file layout is inseparable from COCOeval. Several fields exist only because the evaluator reads them. A dataset that fills them carelessly produces a metric measuring something other than it claims.
Area-based stratification. The evaluator splits results by the area field, at thresholds $32^2$ and $96^2$ pixels. It reports $\mathrm{AP}_S$, $\mathrm{AP}_M$ and $\mathrm{AP}_L$. When a mask is present, area is the segmentation area rather than the box area. Copying box area into the field moves objects into larger buckets. That inflates $\mathrm{AP}_S$, by removing genuinely small objects from it.
Crowd handling changes the matching rule. The mask utilities document it exactly: for a ground truth marked iscrowd, the modified criterion is
$$ \mathrm{IoU}(g, d, \text{iscrowd}) = \frac{\mathrm{area}(g \cap d)}{\mathrm{area}(d)} $$
rather than the usual $\mathrm{area}(g \cap d) / \mathrm{area}(g \cup d)$. A detection matching any sub-region of a crowd should not be penalised for the crowd's full extent. Crowd ground truths are also excluded from the false-positive count. A converter that drops iscrowd turns ignore regions into hard negatives. Detections landing on them become false positives.
Detections per image are capped. COCOeval uses maxDets = [1, 10, 100] by default, and the headline AP is computed at 100. A model emitting 300 boxes per image is truncated by confidence rank. A poorly calibrated confidence head then loses recall the model actually has.
The AP definition
COCO's primary metric averages the 101-point interpolated average precision. It averages over categories and over 10 IoU thresholds, ${0.50, 0.55, \ldots, 0.95}$:
$$ \mathrm{AP} = \frac{1}{101}\sum_{r \in {0, 0.01, \ldots, 1}} p_{\text{interp}}(r), \qquad p_{\text{interp}}(r) = \max_{\tilde{r} \ge r} p(\tilde{r}) $$
$p(r)$ is precision at recall $r$. The 101-point sampling replaces Pascal VOC's 11-point scheme. Averaging over IoU thresholds is what makes COCO AP sensitive to localisation quality, not only detection. It is also what makes annotation tightness matter. The previous lesson showed two annotators differing by five pixels sit near IoU 0.79. Thresholds from 0.80 upward are therefore measuring annotation noise.
Matching is greedy by descending detection confidence, one ground truth per detection. It runs independently per category and per IoU threshold.
Storage characteristics
RLE size is $O(\sqrt{n})$ in the object area $n$, for simple shapes. It is proportional to the number of vertical boundary crossings, not to the pixel count. The COCO mask documentation states this directly. Union, intersection, area and IoU can be computed on the encoded form, in time linear in the RLE length. That is why large-scale mask evaluation is tractable at all.
Polygons are smaller still for convex objects and are resolution-independent, which matters when images are resized after annotation. They cannot represent holes or disconnected regions without convention. They must also be rasterised before any set operation. That is why the evaluator converts everything to RLE internally, through frPyObjects.
Extensions and successors
| Format | Change | Motivation |
|---|---|---|
| COCO Panoptic | PNG-encoded segment ids plus a JSON sidecar | One label per pixel; things and stuff unified |
| LVIS | 1200+ categories, federated annotation | Long-tailed distribution; exhaustive annotation is infeasible |
| Open Images | CSV, hierarchical labels, explicit negatives | Class hierarchy and verified absence |
| WebDataset / Parquet | Sharded tar or columnar tables | Streaming from object storage at scale |
LVIS's federated design is the most interesting departure. With 1200 categories, no image is exhaustively annotated for all of them. Each category carries its own set of images, in which it was verified present or absent. Evaluation only counts an image against a category when that image is in the category's set. A single COCO-style file cannot express this. Treating LVIS as if it were COCO produces enormous phantom false-positive counts.
The general lesson is this. As label spaces grow, exhaustive annotation stops being affordable. The format must record what was checked, not only what was found. COCO's single iscrowd flag and its ability to hold zero-annotation images are the small, early version of that idea.
References
- Lin et al., Microsoft COCO: Common Objects in Context, 2014 — arxiv.org/abs/1405.0312
- COCO API source,
pycocotools/coco.pyandpycocotools/mask.py— github.com/cocodataset/cocoapi - Kirillov et al., Panoptic Segmentation, 2019 — arxiv.org/abs/1801.00868
- Gupta et al., LVIS: A Dataset for Large Vocabulary Instance Segmentation, 2019 — arxiv.org/abs/1908.03195
- Kuznetsova et al., The Open Images Dataset V4, 2020 — arxiv.org/abs/1811.00982
What to learn next
- Converting between dataset formats — moving COCO to YOLO and back without losing anything.
- Object detection — the task this format was designed for.
- Image segmentation — what the polygons and RLE masks are for.