Vision Datasets and Annotation
Choosing an annotation tool
The annotation tool decides your labelling speed, your export format and whether you can leave, so pick it on export quality and not on the demo video.
- 12 min read
- 3 reading levels
- Updated
Read these first
On this page 10
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
An annotation tool is the software your team uses to draw boxes on images and export them.
The analogy
Think about choosing a kitchen before catering a wedding. The showroom kitchen looks beautiful. What decides the wedding is whether four people can work in it at once for eleven hours.
Counter space. Whether the taps reach the big pots. Whether you can find anything.
Annotation tools are the same. The demo always looks fine. What matters is the twentieth hour, and whether your data can leave when you want it to.
What the tool has to do
Four jobs, in order of how much they will hurt you if done badly.
Export cleanly. Your labels must come out in a format your training code reads. If they only come out in the tool's own private format, you are stuck.
Be fast to use. A labeller draws thousands of boxes. One extra click each time costs days.
Handle many people. Who labelled what, who reviewed it, what is left to do.
Show you the work. Which classes are rare, which images nobody has touched, which labels look wrong.
The kinds of tool
Free and self-hosted. You run it on your own machine or server. Nothing leaves your control, and nobody sends you a bill. You do the setting up and the maintaining.
Paid and hosted. Someone else runs it. Faster to start. Your images live on their servers, which may not be allowed for your data.
Written by you. Sometimes right, and much more often a mistake. Annotation tools look easy and are not. Zoom, undo, keyboard shortcuts and multi-user handling are all more work than they appear.
The question to ask first
Where does the data live, and who is allowed to see it?
Medical images and faces of identifiable people often cannot go to somebody else's service. Nor can anything a customer gave you under contract. Answer that before you look at features. It removes most of the list immediately.
The test that actually predicts success
Label fifty images. Export them. Load the export in your training code. Time the whole thing.
Fifty images tells you more than a week of comparing feature tables. You will find the awkward keyboard shortcut and the missing export field. You will find the class list that cannot be edited afterwards.
Do that before you commit a team to a tool.
A warning about being locked in
Some tools make it easy to get data in and hard to get it out.
Check the export before you start, not after. Ask for the full annotation history, not only the final labels. If you cannot get your work out in a standard file, you do not own your work.
Where you have seen this
- Any "select all squares with a bus" verification puzzle. That is annotation, done by you, for free.
- Maps apps asking you to confirm a shop name from a photograph.
- Photo apps asking whether the person they found is a particular friend.
Remember this
- Choose on export quality and privacy rules first, features second.
- Run a fifty-image trial before committing anyone.
- If you cannot export in a standard format, you do not own your labels.
What to learn next
- Writing annotation guidelines — the document that decides whether your labels agree.
- The COCO dataset format — the export format everything else converts to.
- Data labelling — the wider practice this fits inside.
Developer — Code and libraries.
Setup
python3 --version # this uses the standard library onlyThe script to run on every export, before anything else
Whatever tool you pick, its export will disagree with your image folder. Not sometimes; every time. Run this before you train, not after a confusing evaluation.
import json, collections
# A tool's export, as it actually arrives: a label file plus a folder of images.
# These two lists never agree on the first try.
images_on_disk = ["a.jpg", "b.jpg", "c.jpg", "d.jpg", "e.jpg", "f.jpg", "g.jpg"]
export = {
"images": [{"id": 1, "file_name": "a.jpg", "width": 640, "height": 480},
{"id": 2, "file_name": "b.jpg", "width": 640, "height": 480},
{"id": 3, "file_name": "c.jpg", "width": 640, "height": 480},
{"id": 4, "file_name": "d.jpg", "width": 640, "height": 480},
{"id": 5, "file_name": "z.jpg", "width": 640, "height": 480}],
"annotations": [{"id": 1, "image_id": 1, "category_id": 1, "bbox": [10, 10, 50, 60]},
{"id": 2, "image_id": 1, "category_id": 1, "bbox": [10, 10, 50, 60]},
{"id": 3, "image_id": 2, "category_id": 2, "bbox": [0, 0, 640, 480]},
{"id": 4, "image_id": 3, "category_id": 1, "bbox": [5, 5, 0, 40]},
{"id": 5, "image_id": 4, "category_id": 1, "bbox": [600, 400, 100, 200]},
{"id": 6, "image_id": 5, "category_id": 3, "bbox": [1, 1, 20, 20]}],
"categories": [{"id": 1, "name": "helmet"}, {"id": 2, "name": "no-helmet"}],
}
by_id = {im["id"]: im for im in export["images"]}
named = {im["file_name"] for im in export["images"]}
cat_ids = {c["id"] for c in export["categories"]}
ann_per_image = collections.Counter(a["image_id"] for a in export["annotations"])
print(f"images listed in the export : {len(export['images'])}")
print(f"images present on disk : {len(images_on_disk)}")
print(f"annotations : {len(export['annotations'])}\n")
problems = []
for f in sorted(named - set(images_on_disk)): # sorted, so the report is reproducible
problems.append(f"label file references a missing image: {f}")
for f in sorted(set(images_on_disk) - named):
problems.append(f"image on disk has no entry in the labels: {f}")
for im in export["images"]:
if ann_per_image[im["id"]] == 0:
problems.append(f"image with zero annotations: {im['file_name']} (is it truly empty?)")
seen = {}
for a in export["annotations"]:
if a["category_id"] not in cat_ids:
problems.append(f"annotation {a['id']} uses category {a['category_id']}, which is not declared")
x, y, w, h = a["bbox"]
if w <= 0 or h <= 0:
problems.append(f"annotation {a['id']} has a zero-size box {a['bbox']}")
im = by_id.get(a["image_id"])
if im and (x + w > im["width"] or y + h > im["height"]):
problems.append(f"annotation {a['id']} spills outside {im['file_name']}: {a['bbox']}")
key = (a["image_id"], tuple(a["bbox"]))
if key in seen:
problems.append(f"annotation {a['id']} is an exact duplicate of {seen[key]}")
seen[key] = a["id"]
print("PROBLEMS FOUND")
for p in problems:
print(" -", p)
print("\nclass balance")
names = {c["id"]: c["name"] for c in export["categories"]}
for cid, n in collections.Counter(a["category_id"] for a in export["annotations"]).most_common():
print(f" {names.get(cid, f'UNDECLARED id {cid}'):20s} {n}")images listed in the export : 5 images present on disk : 7 annotations : 6 PROBLEMS FOUND - label file references a missing image: z.jpg - image on disk has no entry in the labels: e.jpg - image on disk has no entry in the labels: f.jpg - image on disk has no entry in the labels: g.jpg - annotation 2 is an exact duplicate of 1 - annotation 4 has a zero-size box [5, 5, 0, 40] - annotation 5 spills outside d.jpg: [600, 400, 100, 200] - annotation 6 uses category 3, which is not declared class balance helmet 4 no-helmet 1 UNDECLARED id 3 1
Reading the output, and why each check earns its place
A label referencing a missing image crashes some loaders and is silently skipped by others. Silent skipping is worse: your dataset is smaller than you think and you never find out.
Images with no labels are ambiguous and this ambiguity ruins detectors. Does e.jpg contain no helmets, or did nobody look at it? A detector trained on unlabelled positives learns that helmets are background. Every export needs an explicit list of reviewed-and-empty images.
Exact duplicate annotations come from double-clicking or from merging two exports. They inflate your class counts and, in detection, they teach the model to predict two boxes where one belongs.
Zero-size boxes come from a click that was meant to be a drag. Some training pipelines crash on them; others produce a not-a-number loss twenty minutes in.
Boxes spilling past the image edge come from panning during a drag. They break IoU arithmetic and the coordinate conversions in the next lesson.
Undeclared category ids mean the class list changed mid-project and the export was not regenerated. This one is common and quietly poisonous.
The class balance table is the first thing to look at, every time. The ratio here is four to one, which is already worth planning for. Ratios of a hundred to one are normal in real projects. They change what you train and how you evaluate.
The sorted() calls are deliberate. Set iteration order in Python varies between runs. An unsorted report differs each time, so you cannot diff two audits.
The landscape, as of 2026
Verify anything here before committing; licences and free tiers move.
| Tool | Model | Notes |
|---|---|---|
| CVAT | Open source, self-hosted | Boxes, polygons, semantic and instance masks, skeletons, cuboids, point clouds, video. Auto-annotation integrations including SAM-family and YOLO models. Self-hosted is the complete product. |
| Label Studio | Open source, self-hosted | Strongest for multi-modal projects mixing images with text or audio. Flexible labelling-config system. |
| FiftyOne | Open source library | Not a labelling tool. Dataset exploration and evaluation, with integrations to CVAT, Label Studio and others. Most teams pair it with a labelling tool rather than replacing one. |
| Roboflow | Hosted, commercial | Fast to start, detection-focused, strong export coverage. Free public tier publishes your data and models to their public catalogue, which decides it for many projects. |
| labelme | Open source, desktop | Single-user, minimal, no server. Good for a few hundred images and a solo project. |
FiftyOne deserves a second mention because it is the piece teams most often discover too late. Labelling and looking at your labels are different jobs, and no labelling tool is good at the second one.
Common mistakes
Choosing before the export test. Label fifty images, export, load in your training code. An afternoon spent here saves weeks.
Letting each labeller define the class list. Freeform class names produce helmet, Helmet, hard hat and hardhat as four classes. Fix the list before anyone starts.
No review stage. A single-pass labelling process has whatever error rate your labellers have. Sample and re-check at minimum.
Not versioning the export. Save each export with a date and a commit hash. Otherwise "the model got worse" becomes unanswerable, because you cannot reconstruct the data it was trained on.
Building your own tool. Reach for this only when your task genuinely has no existing support. Zoom, undo, shortcuts, review workflows and multi-user locking are all more work than they look.
Try it yourself
Add a check for boxes smaller than 4 by 4 pixels and count how many there are. Tiny boxes are usually accidental, and most detector architectures cannot predict them anyway. Then add a check that the image dimensions in the export match the real files on disk. That catches the case where someone resized the images after labelling.
What to learn next
- Writing annotation guidelines — the document that decides whether your labels agree.
- The COCO dataset format — the export format everything else converts to.
- Data labelling — the wider practice this fits inside.
Researcher — Mathematics and papers.
Why tooling is a measurement problem
Annotation quality bounds achievable accuracy. Suppose two competent annotators disagree at IoU 0.7 on a detection task. An evaluation reporting mAP at IoU 0.75 is then measuring annotator variance as much as model quality. The tool determines how much of that variance you can observe. It decides whether multiple independent annotations per item are even possible.
The design questions worth asking of any tool, ordered by how much they affect the resulting statistics:
Can it collect more than one annotation per item? Without this, inter-annotator agreement is unmeasurable, and every quality claim is an assertion.
Does it record time per item? Time correlates with ambiguity. Slow items are informative items. A time distribution is a cheap way to find the parts of your guideline that fail.
Does it preserve the full event history? Which annotations were created, edited, or accepted from a model pre-label. Without this, you cannot separate human labels from model labels after the fact. That makes the automation-bias analysis in a later lesson impossible.
Does it support explicit uncertainty? A forced binary choice on a genuinely ambiguous image manufactures noise. An "unclear" option preserves the information. The images so marked are usually the most valuable ones in the set.
Format as an interoperability constraint
Export format is not a detail; it constrains what can be expressed.
- COCO JSON carries polygons and RLE masks, crowd regions, and per-annotation attributes. Verbose, and the de facto standard.
- YOLO text carries a class and a normalised box per line. No masks, no attributes, no image-level metadata, no explicit list of empty images.
- Pascal VOC XML carries per-object
difficultandtruncatedflags that most converters silently discard. - Parquet or Arrow table formats scale to millions of rows and support columnar filtering. That is how large-scale pipelines actually store this.
The failure to plan for is not conversion difficulty; it is silent loss. Converting COCO to YOLO discards segmentation, crowd flags and empty-image markers, and the pipeline reports success. The next lesson measures exactly what these conversions cost.
Programmatic dataset curation
FiftyOne (Voxel51) is the reference open-source implementation of dataset-as-a-queryable-object. The operations it makes routine are the ones that change project outcomes:
- Filtering by prediction and ground-truth jointly, to enumerate false positives above a confidence threshold.
- Sorting by embedding similarity or by uniqueness, to find near-duplicates and outliers.
- Computing per-sample evaluation results so mAP can be decomposed by attribute rather than reported as one number.
- Round-tripping a filtered subset to an annotation tool for correction.
The general point: dataset work needs a query layer. Some teams treat the label file as an opaque blob, handed to a training script. They cannot ask the questions that matter.
Agreement, and its cost
Multiple annotation is expensive and partial coverage is the correct answer. A standard design labels 100 percent of items once, and 5 to 10 percent redundantly. The redundant subset estimates agreement, calibrates individual annotators, and detects drift as the project runs.
Report agreement alongside dataset size. A benchmark paper may state the number of images and not the agreement between annotators. It has not stated its own noise floor.
Further reading
- Voxel51, FiftyOne documentation — docs.voxel51.com
- CVAT documentation — docs.cvat.ai
- Label Studio documentation — labelstud.io/guide
- Northcutt et al., Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks, 2021 — arxiv.org/abs/2103.14749
What to learn next
- Writing annotation guidelines — the document that decides whether your labels agree.
- The COCO dataset format — the export format everything else converts to.
- Data labelling — the wider practice this fits inside.