Datasets and DataLoaders

Surviving corrupt and missing files

Real datasets contain broken files, and one of them can kill hour six of training — the fix is a wrapper that catches the failure, a collate that drops the gap, and a count that keeps everyone honest.

Read these first

On this page 5
  1. Why it exists
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Corrupt-sample handling means deciding, in advance, what happens when one file out of a million refuses to open — instead of letting it crash your training.

A cook prepping fifty kilos of tomatoes will meet rotten ones. No cook shuts the kitchen over a bad tomato. It goes in a separate crate, a tally mark goes on the wall, and the chopping continues.

Now the important part: at closing time, the cook counts the crate. Three rotten tomatoes — normal. Twenty kilos rotten — someone must call the supplier, because the problem is not tomatoes anymore.

Data pipelines need exactly this: keep going past one bad file, count every skip, and treat a rising count as an alarm rather than a nuisance.

Why it exists

Any dataset that touched the real world contains damage. Downloads truncated mid-file. Photos with the wrong extension. A disk hiccup at hour nine of scraping. In a million files, some are broken — this is a law, not bad luck.

The default behaviour is the worst one: the first broken file raises an error, and training dies — possibly six hours in, on file 800,000. Restart, hit it again, die again.

But silent skipping is a trap of its own. If skips are quiet and thousands happen, your model trained on far less data than you believe — or worse, on a biased slice, if one camera or one day produced most of the damage. Hence the two rules: never crash on one file, never skip without counting.

How it works

 "give me sample i"
        |
   try to load it
     ok?  → hand it over
     broken? → note it in the tally  → hand over "nothing"
        |
 batch assembly: drop the "nothing"s, pack the rest
        |
 end of run: read the tally.  3 skips? fine.  30,000? investigate.

A real example you have seen

Every photo-recognition model trained on internet images swam through broken downloads and mislabelled formats. The pipelines that produced working models were the ones that skipped, counted, and audited the counts.

Remember this

  • Broken files are guaranteed at scale; crashing on them is a design choice — a bad one.
  • Skip past damage, but count every skip.
  • A rising count means a supplier problem, not a tomato problem.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install torch

Written and tested against torch 2.5 on CPU. Damage is simulated, so this runs anywhere, byte-identical.

Skip, count, and keep the batch flowing

skip_broken.py
import torch
from torch.utils.data import Dataset, DataLoader, default_collate

BROKEN = {13, 27, 41}                # pretend these files fail to decode

class FlakyDataset(Dataset):
    def __len__(self):
        return 64

    def __getitem__(self, idx):
        if idx in BROKEN:
            raise OSError(f"file {idx:05d}.jpg is corrupt")
        g = torch.Generator().manual_seed(idx)
        return torch.rand(3, 8, 8, generator=g), idx % 2

class SkipBroken(Dataset):
    """Wraps any dataset. A broken sample becomes None instead of a crash."""
    def __init__(self, inner):
        self.inner = inner

    def __len__(self):
        return len(self.inner)

    def __getitem__(self, idx):
        try:
            return self.inner[idx]
        except OSError as e:
            print("skipping:", e)
            return None

def drop_none_collate(batch):
    batch = [s for s in batch if s is not None]
    if not batch:                      # a whole batch of corrupt files: give up loudly
        raise RuntimeError("every sample in this batch was corrupt")
    return default_collate(batch)

loader = DataLoader(SkipBroken(FlakyDataset()), batch_size=16,
                    collate_fn=drop_none_collate)
sizes = [len(y) for _, y in loader]
print("batch sizes:", sizes, "- total", sum(sizes), "of 64")
Output
skipping: file 00013.jpg is corrupt
skipping: file 00027.jpg is corrupt
skipping: file 00041.jpg is corrupt
batch sizes: [15, 15, 15, 16] - total 61 of 64

The walkthrough

The wrapper pattern. SkipBroken wraps any dataset without touching its code — damage policy stays separate from loading logic, so one wrapper serves every project. Catch the exceptions loading genuinely raises: OSError covers most file damage; add ValueError for decode libraries that throw it. Do not catch bare Exception — that would swallow your own bugs, and a typo in the dataset becomes ten thousand "corrupt" files.

None as the tombstone. The wrapper returns None; the custom collate filters it out. The result is visible in the output — batches of 15 where a broken file fell out. Downstream code must not assume constant batch size (one more reason BatchNorm dislikes stray tiny batches).

The all-corrupt batch raises. If sixteen consecutive files are broken, something systemic is wrong — a mount lost, a directory deleted. That is precisely when you want a crash. Skipping is for retail damage; wholesale damage should stop the line.

In real projects, count properly. The print is for this demo; production wants a counter (per worker) logged at epoch end, with a threshold that fails the run: assert skipped / total < 0.01. The counting is not decoration — it is the difference between resilience and self-deception.

The alternative: substitute instead of skip

When constant batch size matters, replace a broken sample with a working neighbour:

python
def __getitem__(self, idx):
    for attempt in range(10):
        try:
            return self.inner[(idx + attempt) % len(self.inner)]
        except OSError:
            continue
    raise RuntimeError(f"10 consecutive failures starting at index {idx}")

Trade-off: batch size stays fixed, but neighbours of broken files get sampled slightly more often — a tiny bias, acceptable in almost all cases, and the ten-strike limit again turns systemic damage into a loud stop.

The best fix happens before training: audit once

python
# one-off sweep, run before the first epoch ever starts
ds = FlakyDataset()          # or whatever raw Dataset you are about to train on
bad = []
for i in range(len(ds)):
    try:
        ds[i]
    except OSError:
        bad.append(i)
print(len(bad), "broken files:", bad[:10])

Run the sweep (with workers, for speed, on big sets), quarantine the list, and train on a clean index. Runtime skipping then remains as the safety net it should be — catching new damage, not known damage, every epoch.

Common mistakes

Silent swallowing. except: return None with no count. Months later: "why is this model worse than the paper?" Because it trained on 60% of the data, and nobody knew.

Catching too much. Bare except Exception hides real bugs as fake corruption. Catch what loading throws; let your own errors crash honestly.

Retrying the same file forever. A retry loop without an attempt cap plus one permanently-broken file equals an infinite loop inside __getitem__ — training hangs with no error and no progress bar movement.

Auditing once and trusting forever. Datasets grow and get re-synced; new damage arrives. Keep the runtime net even after a clean audit.

Try it yourself

Extend SkipBroken with a self.skipped counter and print it after the loop — then grow BROKEN to 30 indices and watch batch sizes shrink. Decide, before you look, what threshold should abort a real run.

What to learn next

Researcher — Mathematics and papers.

The statistical cost of skipping

Skipping is data censoring. If corruption is independent of content — missing completely at random — dropping damaged samples costs sample size and nothing else. The dangerous regime is informative missingness: damage correlated with acquisition conditions (one faulty sensor, one truncated scrape day, one camera model whose JPEGs decode badly), which selectively removes a region of the input distribution. The empirical risk then optimises a reweighted population, and the deployed model meets the missing region cold. Detection is cheap and worth institutionalising: log skipped identifiers, not counts alone, and test the skip set for structure (per-source rates, per-class rates) — a two-way frequency table catches most real incidents. This is the missing-data taxonomy of Rubin (1976) — MCAR/MAR/MNAR — transplanted from statistics into pipeline engineering.

The substitute-neighbour pattern induces a related, smaller distortion: sampling weight shifts from broken indices onto their successors — bounded by the corruption rate, and worth noting in reproducibility docs since it couples sample exposure to file ordering.

Failure semantics in the loader machinery

An exception raised in a worker's __getitem__ propagates: the worker catches it, ships it to the parent, and the parent re-raises at the next(loader) boundary — with a traceback naming the worker. Uncaught, one bad file therefore kills the epoch regardless of num_workers; the wrapper moves the decision into user space. A worker dying (segfault in a decode library — libjpeg on truncated data has history here) is a different class: the parent raises the DataLoader worker exited unexpectedly error, and no Python-level catch inside the dataset can help — process-level supervision or pre-audit are the only defences. This distinction — exception versus process death — decides where your safety net can live.

Timeouts complete the taxonomy: a hung read (network filesystem stall) blocks a worker indefinitely; DataLoader(timeout=...) bounds the parent's wait, converting hangs into catchable errors.

Engineering doctrine

Large-scale practice converges on defence-in-depth: manifest with checksums at ingestion (immutable list of verified files — the audit, formalised); quarantine rather than delete (damaged files carry forensic value — the truncation pattern identifies the failing component); runtime skip-with-metrics as the last line, alarmed at a rate threshold, per the developer block. Dataset cards (Gebru et al., 2021, Datasheets for Datasets) provide the documentation frame: known damage rates and exclusion criteria belong in the datasheet, since two teams "training on the same dataset" with different skip policies are not training on the same dataset. Webdataset-style pipelines expose the same choice as a handler parameter (continue/warn/stop callbacks) — evidence that the pattern in this lesson is the field's consensus shape, not one site's habit.

Reading

  • Rubin (1976), Inference and Missing Data — the MCAR/MAR/MNAR frame.
  • Gebru et al. (2021), Datasheets for Datasets.
  • Sambasivan et al. (2021), "Everyone wants to do the model work, not the data work" — field evidence on where ML actually fails.

What to learn next