The Datasets library
A Dataset is a table backed by a memory-mapped Arrow file, so it behaves like a Python list of dictionaries while staying larger than your RAM.
- 9 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
A Dataset is a table of examples you can index like a list. The rows stay on disk until you touch them.
Think of a library reading room with a card catalogue. The catalogue tells you exactly which shelf and which page holds what you want. You do not carry every book to your desk — you fetch one page at a time, and the reading room stays usable.
A Dataset works that way with your training data. The table lives in a file on disk. Ask for row 4,000,000 and that row alone comes to you.
Why it exists
The old habit was loading a CSV into memory with pandas. That works beautifully until the file is bigger than your RAM. Then it stops working at all — with a crash, after an hour of loading.
There was a second problem. Everyone wrote their own loading code, their own train/test split, their own filtering. The same bugs got reinvented in every project.
The library fixes both. One table format, on disk, with the splitting and filtering already written and tested.
How it works
reviews.csv → Arrow table on disk → ds[0] one row
(stays there) → ds["stars"] one column
→ ds.filter(...) a smaller view
→ ds.train_test_split() two tablesThree ideas do most of the work:
- Rows are your examples — one review, one image, one recording.
- Features are the column descriptions: this one is text, that one is a whole number.
- Splits are named slices:
train,validation,test.
A real example you have seen
Your phone's contact list. Thousands of entries, and scrolling is instant, because the phone loads the few names on screen and leaves the rest in storage. Nobody loads all contacts into memory to show you five.
Remember this
- A Dataset is a table on disk that you use like a Python list.
- Features name and type the columns; splits are the named train/test slices.
- Rows are fetched when touched, so the data can be far bigger than your RAM.
What to learn next
- map, batched and the cache — turning raw text into model inputs, once.
- Data collators — the step between a Dataset and a batch.
- Datasets too big for RAM — the same problem, solved from the PyTorch side.
Developer — Code and libraries.
Setup
pip install datasetsTested with datasets 5.0. Nothing here downloads a model, and the data is five rows we write ourselves — the whole lesson runs offline after install.
Building a dataset from a file you own
import csv
from datasets import load_dataset
rows = [
{"text": "The biryani arrived cold", "stars": 1},
{"text": "Perfect filter coffee, came back twice", "stars": 5},
{"text": "Service was slow but the dosa was good", "stars": 3},
{"text": "Overpriced and the AC was broken", "stars": 2},
{"text": "Best vada pav in the neighbourhood", "stars": 5},
]
with open("reviews.csv", "w", newline="", encoding="utf-8") as f:
w = csv.DictWriter(f, fieldnames=["text", "stars"]); w.writeheader(); w.writerows(rows)
ds = load_dataset("csv", data_files="reviews.csv", split="train")
print(ds)
print("features :", ds.features)
print("one row :", ds[0])
print("one column:", ds["stars"])
print("a slice :", ds[:2]["text"])
good = ds.filter(lambda r: r["stars"] >= 4)
print("kept :", good.num_rows, good["text"])
print(ds.train_test_split(test_size=0.4, seed=0))Dataset({
features: ['text', 'stars'],
num_rows: 5
})
features : {'text': Value('string'), 'stars': Value('int64')}
one row : {'text': 'The biryani arrived cold', 'stars': 1}
one column: Column([1, 5, 3, 2, 5])
a slice : ['The biryani arrived cold', 'Perfect filter coffee, came back twice']
kept : 2 Column(['Perfect filter coffee, came back twice', 'Best vada pav in the neighbourhood'])
DatasetDict({
train: Dataset({
features: ['text', 'stars'],
num_rows: 3
})
test: Dataset({
features: ['text', 'stars'],
num_rows: 2
})
})Progress bars for the generating and filtering steps also appear on your terminal. They are written to the error stream, so they are left out above.
Naming your labels
from datasets import Dataset, ClassLabel
ds = Dataset.from_dict({
"text": ["cold biryani", "great coffee", "slow service", "best vada pav"],
"label": [0, 1, 0, 1],
})
ds = ds.cast_column("label", ClassLabel(names=["bad", "good"]))
print("features:", ds.features)
print("row 0 :", ds[0], "->", ds.features["label"].int2str(ds[0]["label"]))
print("in-memory (no cache files):", ds.cache_files)
print("arrow bytes:", ds.data.nbytes)
torch_ds = ds.with_format("torch")
print("torch format:", torch_ds[0])features: {'text': Value('string'), 'label': ClassLabel(names=['bad', 'good'])}
row 0 : {'text': 'cold biryani', 'label': 0} -> bad
in-memory (no cache files): []
arrow bytes: 97
torch format: {'text': 'cold biryani', 'label': tensor(0)}The walkthrough
Indexing by row gives a dict; indexing by name gives a column. ds[0] is one example. ds["stars"] is every value in that column. ds[:2]["text"] slices rows first, then picks a column. Reading that pattern correctly saves a lot of confused debugging.
A Column is not a Python list. In datasets 5 the column accessor returns a lazy Column object. Iterate it, index it, or wrap it in list(...) when a plain list is what you need.
ClassLabel turns integers into names that travel with the data. Now int2str and str2int exist, train_test_split(stratify_by_column="label") becomes legal, and the Trainer can print readable metric names. Do this early — retrofitting label names after a confusing evaluation is the harder path.
from_dict builds in memory; load_dataset builds on disk. Notice cache_files is empty above. Data created from a Python dict lives in RAM. Data loaded from a file is written to an Arrow file and memory-mapped, which is where the bigger-than-RAM property comes from.
with_format("torch") changes what indexing returns, not what is stored. The label came back as a tensor while text stayed a string, because strings have no tensor form. with_format("torch", columns=["input_ids", "label"]) is the usual call before feeding a DataLoader.
Splitting is seeded, so it is reproducible. train_test_split(test_size=0.4, seed=0) gives the same two tables on every machine. Leaving the seed out gives a different split each run, and metrics that wobble for no visible reason — see train/test split.
Common mistakes
Reaching for pandas out of habit. ds.to_pandas() works and is the right move for a plot or a quick group-by. It also loads everything into memory, which defeats the point on a large dataset. Convert a sample, not the table.
Expecting filter and map to modify in place. They return a new dataset and leave the original untouched. ds.filter(...) on its own line changes nothing. Assign the result.
Splitting after shuffling badly, or not at all. train_test_split shuffles by default. If your rows are grouped by class, an unshuffled split hands you a train set of one class. If rows are duplicated across the split, your test score is a memory test — see leakage.
Loading community datasets without reading them. A Hub dataset can be tens of gigabytes and carry a license that forbids your use. Check the size on the dataset page and the license field before load_dataset starts filling your disk.
Assuming loading scripts still run. Datasets 4 removed Python loading scripts from the Hub. Old tutorials that pass trust_remote_code=True now fail; those datasets have mostly been re-published as plain Parquet, which loads with no argument at all.
Try it yourself
Add a length column with ds.add_column("length", [len(t) for t in ds["text"]]), then sort by it with ds.sort("length"). Print ds.cache_files before and after. Which operations wrote a new Arrow file, and which returned a view?
What to learn next
- map, batched and the cache — turning raw text into model inputs, once.
- Data collators — the step between a Dataset and a batch.
- Datasets too big for RAM — the same problem, solved from the PyTorch side.
Researcher — Mathematics and papers.
Arrow, and why the table is not in your process
The backing store is Apache Arrow IPC: a columnar, fixed-layout format read through mmap. The operating system pages in only the bytes touched, so resident memory tracks your access pattern rather than dataset size, and two processes opening the same file share those pages. ds.data.nbytes reports the logical table size — 97 bytes above — not process RSS.
Three properties follow directly, and they are the library's whole argument:
- Zero-copy slicing.
ds[:2]builds a view over existing buffers rather than copying rows. - Cheap concatenation and interleaving. Multi-dataset training mixes tables without materialising the union.
- Language-neutral files. The same Arrow file is readable from Rust, R, DuckDB and Polars, which is why
ds.datacan be handed to other engines.
The cost is that Arrow is immutable. Every transformation writes a new table, which is exactly why caching gets its own lesson: map, batched and the cache.
The Features schema
Features is a typed schema mapping column names to Value, ClassLabel, List, Image, Audio and nested structs. It serves two jobs beyond documentation. It fixes the Arrow physical type, which determines file size and read speed — int8 labels against int64 is a real difference across a billion rows. And it makes decoding lazy: an Image or Audio column stores a path or bytes and decodes on access, so a dataset of a million JPEGs never decodes the 999,999 you skipped.
ClassLabel additionally carries the integer-to-name mapping into every downstream tool, which is how stratified splitting and readable metric reporting work without a side-channel dictionary.
Map-style versus iterable
Dataset is map-style: it knows its length and supports random access, so shuffling is a permutation of indices and splitting is exact. IterableDataset (from streaming=True, or ds.to_iterable_dataset()) yields examples in order from a remote or sharded source. It has no len, shuffles through a fixed-size reservoir buffer plus shard-order permutation, and splits only by taking prefixes. The trade is precision for the ability to start training on a 10 TB corpus before any of it lands on disk. PyTorch's own split of the same idea is covered in iterable datasets and streaming.
Version note: datasets 4 removed Hub loading scripts (trust_remote_code) in favour of Parquet-only data repos, closing an arbitrary-code-execution path structurally identical to the pickle problem in the Hub cache. Datasets 5 returns Column objects from column access and prints the compact Value('string') schema repr shown above; both are cosmetic changes over the same Arrow core.
What to learn next
- map, batched and the cache — turning raw text into model inputs, once.
- Data collators — the step between a Dataset and a batch.
- Datasets too big for RAM — the same problem, solved from the PyTorch side.