Datasets and DataLoaders

Datasets that do not fit in memory

When data outgrows RAM, stop loading it and start reading it — memory-mapped files give random access to hundred-gigabyte arrays while RAM holds only the rows in use.

Read these first

On this page 5
  1. Why it exists
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

When a dataset is bigger than your computer's memory, the fix is to leave it on disk and fetch only the piece you need, the moment you need it.

Nobody brings the whole library home. You would need a truck, and your room would fit a fraction of the shelves anyway. You borrow the one book you are reading, return it, borrow the next.

Your computer has the same two rooms. RAM — the desk — is fast and small. Disk — the library — is slow and enormous. A dataset of 200 GB does not fit on a 16 GB desk, and no amount of clever code changes that arithmetic.

What changes everything is the borrowing habit: keep the collection in the library, bring one book at a time to the desk.

Why it exists

The natural first instinct — load everything, then train — dies the day your data outgrows RAM. The crash even has a famous name: the out-of-memory error, the OOM.

But look at what training actually needs. Each step touches one batch — a few hundred samples. The other hundreds of gigabytes are furniture. Holding them on the desk serves nothing.

So the whole trick is a reading discipline, and the Dataset contract is already shaped for it: "give me sample i" can be answered by reading sample i from disk right then. The desk holds one batch; the library holds the dataset.

How it works

 the doomed way:
   disk (200 GB) → load ALL into RAM (16 GB) → crash

 the borrowing way:
   disk (200 GB) ← stays put
        |
   "sample 4,081,337, please"
        |
   read ONE row from disk  →  desk holds one batch  →  train  →  repeat

One refinement makes it fast: a memory map — the operating system pretends the whole file is in memory, and quietly fetches only the pages you touch. Your code indexes a giant array; the machinery borrows books behind your back.

A real example you have seen

Models trained on internet-scale text and photos face this at absurd scale — no machine holds those collections in RAM. Streaming and mapped reads are the only physics that work.

Remember this

  • RAM is the desk, disk is the library; big data stays in the library.
  • Fetch per sample, inside "give me sample i" — the Dataset contract already fits.
  • Memory maps make disk look like memory and fetch only what you touch.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install torch numpy

Written and tested against torch 2.5 and NumPy 1.26 on Windows. The file created is a modest 32 MB — big enough to be honest, small enough to be polite.

A memory-mapped Dataset

memmap_dataset.py
import numpy as np
import torch
from torch.utils.data import Dataset, DataLoader

# One-time step: write 200,000 rows to disk (~30 MB). Real projects do this once.
path = "big_features.npy"
rows, cols = 200_000, 40
np.save(path, np.random.default_rng(0).standard_normal((rows, cols), dtype=np.float32))

class MemmapDataset(Dataset):
    """Reads rows straight from disk. RAM holds one batch, not the file."""
    def __init__(self, path):
        self.data = np.load(path, mmap_mode="r")     # opens the file, copies nothing

    def __len__(self):
        return len(self.data)

    def __getitem__(self, idx):
        row = np.array(self.data[idx])               # this line reads ONE row from disk
        return torch.from_numpy(row)

ds = MemmapDataset(path)
print("rows on disk:", len(ds))
print("one sample:", tuple(ds[123].shape), ds[123].dtype)

loader = DataLoader(ds, batch_size=256, shuffle=True,
                    generator=torch.Generator().manual_seed(0))
x = next(iter(loader))
print("one batch:", tuple(x.shape))
import os
print(f"file size: {os.path.getsize(path) / 1e6:.0f} MB")
Output
rows on disk: 200000
one sample: (40,) torch.float32
one batch: (256, 40)
file size: 32 MB

At 32 MB this file fits in RAM easily — the point is that nothing in the code would change at 320 GB. That is the property you are buying.

The walkthrough

mmap_mode="r" is the entire magic. np.load without it reads the whole file into RAM; with it, the call returns instantly having read only a small header, and the array is a window onto the file. Indexing a row makes the operating system fetch the pages containing that row — a few kilobytes — and nothing else.

np.array(self.data[idx]) forces a real copy of the one row. Skip it and you return a live view into the map; a batch of views keeps the map pinned across process boundaries and interacts messily with workers. Copy the row; it is tiny.

shuffle=True works untouched — this is the quiet advantage over streaming. The map preserves random access, so the shelf semantics — shuffling, samplers, exact epochs — all survive. Streams give none of that back.

Building the file in the first place: convert once, in chunks — read 10,000 rows from the source (CSVs, a database), write them into a pre-allocated np.lib.format.open_memmap(...) array, repeat. The conversion script also fits in constant RAM.

Storage formats, honestly ranked for this job

Raw .npy + memmap: unbeatable for fixed-size numeric rows — zero decode cost, perfect random access. Fixed size is the catch: variable-length samples need either padding on disk or an offsets file alongside. Images usually stay as individual encoded files (JPEG is its own compression; a folder is an out-of-core store, as the lazy-loading pattern showed). Parquet/Arrow suit tabular work with column selection. HDF5 handles groups of arrays but carries a sharp multiprocessing edge, noted below.

Common mistakes

Benchmarking on warm cache. The second epoch reads from the operating system's page cache — RAM — and flies. The first epoch off a cold disk tells the truth. Judge speed on epoch one, or after a reboot.

Random access on a spinning hard drive. Memmap plus shuffle=True on an HDD means a physical head seek per row — brutal. On SSDs this pattern is fine; on HDDs, coarsen randomness: shuffle chunks of a few thousand contiguous rows, then shuffle within the loaded chunk.

HDF5 handles crossing into workers. An h5py.File opened in __init__ breaks when workers fork or spawn — corrupt reads or crashes. The fix is lazy per-worker opening: store the path, open on first __getitem__ call in each worker (if self.f is None: self.f = h5py.File(...)).

dtype surprises doubling the file. Saving float64 doubles size and halves throughput for no modelling gain. Decide dtype at conversion — float32, or even float16/uint8 where the data allows — and write it once.

Try it yourself

Time one full pass at batch_size=256 twice in the same run — cold-ish, then warm. The gap you measure is the page cache at work, and it is exactly why the "benchmark warm" mistake fools people.

What to learn next

Researcher — Mathematics and papers.

What mmap actually does

mmap maps file extents into virtual address space; page faults trigger 4 KB-page reads on first touch, and the page cache retains them under LRU-ish eviction. Consequences worth engineering around: (1) the second epoch is a RAM benchmark — steady-state throughput equals page-cache hit rate times RAM bandwidth plus miss rate times storage bandwidth; (2) memory pressure from the cache is reclaimable and does not OOM, but competes with your model's allocations, visible as unexplained slowdowns rather than errors; (3) access granularity is the page, so rows smaller than 4 KB cost amplification on random reads — batching contiguous index ranges (chunked shuffling) recovers sequential bandwidth. The random-versus-sequential gulf is hardware-tiered: NVMe sustains high IOPS making per-row random access viable; object storage inverts the economics entirely, favouring the streaming designs — WebDataset shards, Mosaic StreamingDataset — that convert access into large sequential GETs with local caching.

np.memmap versus np.load(mmap_mode=...): the latter reads the .npy header to recover shape/dtype and then memmaps the payload — same mechanism, self-describing file. Writable modes (r+, w+) make memmap a constant-RAM writer too, which is the standard conversion-script tool (np.lib.format.open_memmap).

Interaction with the loader machinery

Fork inheritance (Linux) shares the map read-only across workers at zero cost — page cache is shared by design. Spawn (Windows) re-opens per worker via pickling of the path-holding dataset; the map itself is not (and must not be) serialised — the lazy-open pattern is the portable idiom, identical in shape to the HDF5 fix. With pin_memory=True, the copy chain is disk page → cache → pinned staging → device; eliminating the intermediate copy motivates frameworks that decode straight into pinned buffers (DALI) or store pre-tensorised samples (FFCV's custom format, Leclerc et al., 2023). For LLM-scale token stores, the memmapped-token-array pattern (uint16 tokens, offsets implicit in fixed block length) is the de facto standard in open pretraining codebases — the exact MemmapDataset above with cols = context length.

Little's-law sizing from the workers lesson applies with storage as producer: required read bandwidth is batch bytes per step over step time; when storage cannot meet it, options are compression (trading CPU decode for bandwidth), caching tiers, or data echoing (Choi et al., 2020). Measure before choosing — the profiling recipe closes this section in the starvation lesson.

Reading

  • Leclerc et al. (2023), FFCV — storage layout co-designed with training loops.
  • Aizman, Maltby, Breuel (2019) — WebDataset and sequential-read economics.
  • Operating systems texts on demand paging (any edition of Arpaci-Dusseau, OSTEP, ch. "Beyond Physical Memory") — the substrate this lesson stands on.

What to learn next