Vision Datasets and Annotation

Building an image dataset from the web

Collecting images from the internet is mostly a question of what you are allowed to use and what you can prove about it, not of how fast you can download.

On this page 10
  1. The short answer
  2. The analogy
  3. Three different permissions, often confused
  4. The one thing to record
  5. What to throw away straight away
  6. The part about people
  7. What your dataset will be biased toward
  8. Where you have seen this
  9. Remember this
  10. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

Downloading images from the internet is easy. Being allowed to use them, and proving where each one came from, is the actual work.

The analogy

Think about picking fruit for a shop. Filling a basket takes an hour.

The rest of the job is longer. Whose orchard was it? Was there a sign? Did you throw away the bruised ones? Can you tell a customer where any particular mango came from?

Nobody buys from a shop that cannot answer those questions. Nobody should train on a dataset that cannot either.

Three different permissions, often confused

Can you fetch the page? Websites publish a file saying which robots may fetch which parts. It is a request, not a law, and good practice is to respect it.

Can you keep the image? That is copyright, and it belongs to whoever made the picture. A picture being visible in public does not make it free to reuse.

Can you train on it? A separate question again, answered differently in different countries, and the answer is still moving.

Those three are not the same permission. People merge them constantly and get into trouble.

The one thing to record

For every image you keep, write down where it came from. Add the date you fetched it, and what the licence said.

Do this while collecting. Never afterwards, because afterwards is impossible.

Somebody may ask you in two years to remove one person's photographs. This record is the only thing that lets you say yes.

   for each image, store:
     the address it came from
     the date you fetched it
     the licence stated on the page
     a fingerprint of the file itself

The fingerprint is a short code computed from the file. Two identical files give the same code, which is how you spot duplicates.

What to throw away straight away

Files that are not images. Files that will not open. Pictures too small to be useful. Exact copies of things you already have.

Doing this at collection time is cheap. Doing it after labelling means you paid to label rubbish.

The part about people

If your pictures contain identifiable people, you have more than a copyright question.

Several large public image sets have been withdrawn because they contained photos of people who never agreed. One large face dataset was taken down entirely.

Ask three things before you collect. Did the people know? Is there a way to remove someone on request? Would you be comfortable explaining this collection publicly?

If the last answer is no, stop.

What your dataset will be biased toward

Whatever the internet has a lot of.

Search for "wedding" and you will get one country's weddings far more than others. Search for "doctor" and see what comes back. Your model learns whatever pattern your collection contains.

The fix is not clever code. It is deliberately searching for what is missing, and counting what you have.

Where you have seen this

  • Image search results, which are the same collection process with a different purpose.
  • Large public datasets that news stories say were built from scraped photos.
  • Websites that now state whether AI crawlers are permitted.

Remember this

  • Fetching, keeping and training on an image are three separate permissions.
  • Record source, date, licence and a fingerprint for every file, while collecting.
  • If the pictures show identifiable people, consent and removal are the first questions.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install pillow==11.0.0     # everything else is the standard library

The pipeline, with the checks that belong in it

Every downloaded file passes through validation, deduplication and provenance recording before it is allowed into the dataset. This runs entirely offline with synthesised files so you can study the structure without touching anyone's server.

collect.py
import hashlib, io, json
from urllib.robotparser import RobotFileParser
from PIL import Image

# ---- 1. read robots.txt BEFORE you write the downloader, not after ----
ROBOTS = """
User-agent: *
Disallow: /private/
Disallow: /admin
Crawl-delay: 5

User-agent: GreedyBot
Disallow: /
"""
rp = RobotFileParser(); rp.parse(ROBOTS.splitlines())
for agent in ("my-dataset-bot", "GreedyBot"):
    for path in ("/photos/cat.jpg", "/private/cat.jpg", "/admin/x"):
        print(f"{agent:16s} {path:20s} allowed: {rp.can_fetch(agent, 'https://example.com' + path)}")
print("crawl delay for us:", rp.crawl_delay("my-dataset-bot"), "seconds between requests")
print("robots.txt is a request, not a licence. It says what you may fetch, never what you may train on.\n")

# ---- 2. every downloaded file gets checked, hashed and recorded ----
def make_photo(seed, size=(64, 48)):
    im = Image.new("RGB", size)
    im.putdata([((x * 7 + seed) % 256, (y * 5) % 256, 90) for y in range(size[1]) for x in range(size[0])])
    buf = io.BytesIO(); im.save(buf, "JPEG", quality=90, exif=Image.Exif())
    return buf.getvalue()

MIN_SIDE, MAX_BYTES = 32, 5_000_000
manifest, seen = [], {}
downloads = [("https://example.com/photos/a.jpg", make_photo(0)),
             ("https://example.com/photos/b.jpg", make_photo(0)),        # same bytes, different URL
             ("https://example.com/photos/c.jpg", make_photo(40)),
             ("https://example.com/photos/tiny.jpg", make_photo(3, (20, 20))),
             ("https://example.com/photos/broken.jpg", b"not an image at all")]

for url, blob in downloads:
    digest = hashlib.sha256(blob).hexdigest()
    if len(blob) > MAX_BYTES:
        print(f"reject {url:40s} too large"); continue
    try:
        im = Image.open(io.BytesIO(blob)); im.verify()
        im = Image.open(io.BytesIO(blob))
    except Exception as e:
        print(f"reject {url:40s} not a decodable image ({type(e).__name__})"); continue
    if min(im.size) < MIN_SIDE:
        print(f"reject {url:40s} {im.size[0]}x{im.size[1]} is below the {MIN_SIDE}px floor"); continue
    if digest in seen:
        print(f"reject {url:40s} byte-identical to {seen[digest]}"); continue
    seen[digest] = url
    clean = Image.new(im.mode, im.size); clean.putdata(list(im.getdata()))   # drops EXIF and ICC
    manifest.append({"sha256": digest[:16], "source_url": url, "size": list(im.size),
                     "bytes": len(blob), "licence": "UNKNOWN - resolve before use",
                     "retrieved": "2026-08-31"})
    print(f"accept {url:40s} {im.size[0]}x{im.size[1]}  sha256 {digest[:16]}")

print("\nmanifest row:")
print(json.dumps(manifest[0], indent=2))
print(f"\n{len(manifest)} of {len(downloads)} downloads survived the checks.")
Output
my-dataset-bot   /photos/cat.jpg      allowed: True
my-dataset-bot   /private/cat.jpg     allowed: False
my-dataset-bot   /admin/x             allowed: False
GreedyBot        /photos/cat.jpg      allowed: False
GreedyBot        /private/cat.jpg     allowed: False
GreedyBot        /admin/x             allowed: False
crawl delay for us: 5 seconds between requests
robots.txt is a request, not a licence. It says what you may fetch, never what you may train on.

accept https://example.com/photos/a.jpg         64x48  sha256 d8b18c58f04c6485
reject https://example.com/photos/b.jpg         byte-identical to https://example.com/photos/a.jpg
accept https://example.com/photos/c.jpg         64x48  sha256 c2a382f3ce561985
reject https://example.com/photos/tiny.jpg      20x20 is below the 32px floor
reject https://example.com/photos/broken.jpg    not a decodable image (UnidentifiedImageError)

manifest row:
{
  "sha256": "d8b18c58f04c6485",
  "source_url": "https://example.com/photos/a.jpg",
  "size": [
    64,
    48
  ],
  "bytes": 1196,
  "licence": "UNKNOWN - resolve before use",
  "retrieved": "2026-08-31"
}

The hash digits and byte count depend on your Pillow and libjpeg build, so yours will differ. Everything else reproduces.

Reading the output

RobotFileParser is in the standard library. No dependency, three lines. It also returns crawl_delay, which most scrapers ignore. That value separates polite collection from a denial-of-service attack on a small site.

Two of five downloads were rejected before ever reaching the dataset. A byte-identical duplicate arriving under a different URL, and a file below the size floor. Both are extremely common in real collection. Content delivery networks serve the same image at many addresses, and thumbnails masquerade as photos.

Image.verify() catches truncated and non-image files. Note the reopen afterwards, which is required. verify() consumes the file object, so the image must be opened again. That is a documented Pillow behaviour and a frequent bug.

The EXIF strip is one line and matters more than it looks. Image.new plus putdata copies pixels and nothing else. Camera EXIF routinely carries GPS coordinates, the device serial number and the date. Shipping a dataset with those intact publishes personal information you did not intend to publish.

The licence field defaults to a string that fails loudly. "UNKNOWN - resolve before use" is deliberately not null and not "". A downstream filter for permissive licences will not silently accept it.

What the manifest is for

Write one row per image, once, at collection time. It is the only thing that makes these operations possible later:

  • Removal on request. Someone asks for their photographs to be taken out. Without the source URL you cannot find them.
  • Licence re-audit. A source changes its terms, or you discover a source was misattributed. You need to know which images came from it.
  • Duplicate tracing. Two images have the same hash; the manifest tells you which URLs served them.
  • Reproducibility. A collaborator asks how the dataset was built. The manifest is the answer.
  • Split hygiene. Grouping by source domain before splitting prevents a whole category of leakage, covered later in this section.

The three permissions, concretely

robots.txt governs fetching. It is a voluntary protocol. Ignoring it is not usually illegal. It is a good way to get your address blocked and your organisation named.

Copyright governs copying and redistribution. Visible on the open web means nothing. Look for an explicit licence: Creative Commons variants, public domain marks, or a site's stated terms. CC BY requires attribution, which means your manifest needs an author field. CC BY-NC prohibits commercial use, which probably includes your product.

Text and data mining exemptions govern training, and they vary by jurisdiction. The EU has an exemption with an opt-out mechanism for rights holders. Other jurisdictions differ and active litigation is ongoing. This page is not legal advice; if the dataset matters commercially, get advice that is.

Common mistakes

Hammering one host. Respect crawl_delay, set a real User-Agent identifying you with a contact address, and back off on errors. A polite scraper is allowed to keep running.

Collecting first and asking about licences later. The metadata is on the page you already fetched. Capture it in the same request or it is gone.

Trusting the file extension. Check the decoded content, as above. A .jpg that is actually HTML is a normal occurrence.

Keeping EXIF. GPS coordinates in a public dataset is a privacy incident.

Only removing exact duplicates. Byte-identical duplicates are the easy tenth. Re-compressed and resized copies need perceptual hashing, which has its own lesson later in this section.

Assuming your search queries produced a balanced dataset. Count what you got, by class and by every attribute you can measure, before you label anything.

Try it yourself

Add a licence extraction step that looks for a Creative Commons URL in an HTML snippet. Images without one should land in a quarantine list, not in the dataset. Then count how large the quarantine is. On most real collections it is the majority, and knowing that number early changes the project plan.

What to learn next

Researcher — Mathematics and papers.

What went wrong with the large public datasets

The failure modes are documented and specific. They are worth knowing, because they are all reachable from a small project.

Consent and dignity. The MS-Celeb-1M face dataset held roughly 10 million images of about 100,000 individuals, collected from the web. Microsoft withdrew it in 2019, after reporting on its use. Prabhu and Birhane (2020) audited ImageNet and 80 Million Tiny Images. They found non-consensual imagery, including of minors, and slurs used as category labels. Their paper is Large Image Datasets: A Pyrrhic Win for Computer Vision?. The 80 Million Tiny Images dataset was withdrawn by its authors, who requested that the community stop using it.

Label taxonomy. ImageNet's person subtree inherited WordNet's categories, which include offensive and unverifiable descriptors. Yang et al. (2020), Towards Fairer Datasets, filtered 1,593 of 2,832 person categories as unsafe or non-imageable. They documented substantial demographic imbalance in the remainder.

Provenance loss at scale. LAION-5B is distributed as URLs and captions rather than images. That shifts both the copyright position and the audit burden onto the user. Birhane et al. (2021) audited LAION-400M and documented explicit and hateful content. In 2023 the dataset was temporarily withdrawn. A Stanford Internet Observatory report had identified child sexual abuse material within it. A re-released version followed after filtering.

The pattern across all four is the same. Scale made manual review infeasible, and automated filters were treated as sufficient. The harms were found by external auditors, not by the dataset authors.

Provenance as an engineering requirement

Treat the manifest as a first-class artefact with the same versioning discipline as code.

  • Content hash as the primary key, not the filename. Filenames collide, change and lie.
  • Source URL, fetch timestamp, and the HTTP response headers that stated the licence or the content type.
  • A licence field with a controlled vocabulary. Include an explicit unknown value that downstream filters treat as excluded, not permitted.
  • An exclusion list keyed by hash. A removal request is then applied at build time, and stays applied across every rebuild.

Distributing URL lists rather than images, as LAION does, is not a full answer. Link rot makes the dataset unreproducible within a few years. The audit burden moves to every user, rather than being discharged once.

Sampling bias, and why it is not fixable downstream

Search engines rank by engagement, recency and language. A query-driven collection samples from the ranking function, not from the world. Shankar et al. (2017), No Classification without Representation, studied geo-tagged Open Images and ImageNet images. Roughly 60 percent came from six countries in Europe and North America. They demonstrated the resulting accuracy gap on images from under-represented regions. de Vries et al. (2019) asked Does Object Recognition Work for Everyone?. Object classifiers performed markedly worse on household items photographed in low-income countries.

Reweighting at training time cannot fix a category that has zero examples. The intervention has to happen during collection. Use stratified queries, multiple languages and region-specific sources. Count explicitly against a target distribution you write down before starting.

Torralba and Efros (2011), Unbiased Look at Dataset Bias, remains the cleanest demonstration that this is measurable. Their "name that dataset" experiment trains a classifier to predict which dataset an image came from. It succeeds far above chance. Suppose a classifier can identify your dataset from an image. That signature is available to any model you train on it.

Deduplication and split hygiene

Web collections contain heavy near-duplication: content delivery networks, thumbnails, re-uploads, watermarked variants, and crops. Exact hashing catches only byte-identical files.

Deduplicate before splitting, never after. The consequence of getting this wrong is quantified by Barz and Denzler in the ciFAIR analysis, covered later here. It is large enough to reverse benchmark conclusions.

Group by source domain as well as by content. Two images from the same photo gallery are not independent samples, and splitting them across train and test leaks.

Documentation

Gebru et al. (2021), Datasheets for Datasets, defines the reporting structure that makes a collection auditable. The sections are motivation, composition, collection process, preprocessing, uses, distribution and maintenance. The questions that most often go unanswered, and most often should not:

  • Does the dataset contain data that might be considered confidential, or that relates to identifiable people?
  • Were the individuals notified, and did they consent?
  • Is there a mechanism to revoke consent, and how is a revocation propagated to trained models?
  • Was the collection reviewed by an ethics board or an equivalent?

That fourth question has an uncomfortable implication. Suppose a model was trained on data that must now be removed. Removing it from the dataset does not remove it from the weights. Machine unlearning is an active research area for that reason. Retraining is often the only reliable remedy, and it is expensive.

References

What to learn next