MemoryError: Unable to allocate N GiB for an array
The array you asked for does not fit in RAM. Read the shape in the message first — a huge dimension usually means an accidental one-hot or pivot explosion, not a genuinely huge dataset.
Updated
The error
MemoryError: Unable to allocate 3.51 GiB for an array with shape (471, 1000000) and data type float64
What it means
NumPy (or pandas on top of it) asked the operating system for one contiguous block of memory and was refused. The message does the maths for you: rows × columns × 8 bytes for float64. Read the shape before anything else — it often reveals the real bug.
A shape like (471, 1000000) rarely comes from honest data. A million columns usually means you one-hot encoded a high-cardinality column (like user IDs), or pivoted on a column with a million distinct values.
Why it happens
Three distinct situations produce this message:
- An accidental explosion.
pd.get_dummieson an ID column, a pivot on timestamps, a cross join. The fix is to not build that matrix at all. - Wasteful dtypes. Everything is
float64andint64by default, twice the size offloat32/int32, and text columns stored as objects are far worse. - Data genuinely bigger than RAM. Then the answer is chunking, not a bigger array.
Note that Windows fails fast on big allocations, while Linux may grant the memory and kill the process later — same disease, different symptom.
How to fix it
1. Sanity-check the shape. If one dimension is absurd, trace where it came from.
print(df.shape)
print(df["user_id"].nunique()) # about to one-hot this? that many columnsFor high-cardinality categoricals, use target encoding, hashing, or an embedding — not one-hot. If you must one-hot, keep it sparse:
from sklearn.preprocessing import OneHotEncoder
enc = OneHotEncoder() # sparse output by default
X = enc.fit_transform(df[["city"]])A sparse matrix stores only the non-zero entries, which for one-hot is one per row.
2. Shrink dtypes. Usually a 2-4x saving for two lines of code.
df["price"] = df["price"].astype("float32")
df["qty"] = df["qty"].astype("int32")
df["city"] = df["city"].astype("category")
print(df.memory_usage(deep=True).sum() / 1e9, "GB")3. Load only what you need.
df = pd.read_csv("big.csv", usecols=["date", "price", "qty"],
dtype={"price": "float32"})4. Process in chunks when the data truly is big.
totals = {}
for chunk in pd.read_csv("big.csv", chunksize=500_000):
for city, s in chunk.groupby("city")["price"].sum().items():
totals[city] = totals.get(city, 0) + sFor heavier work, Polars or DuckDB handle bigger-than-RAM data with far less ceremony.
5. Check you are on 64-bit Python. 32-bit Python caps out near 2 GB regardless of your RAM.
import sys; print(sys.maxsize > 2**32) # True means 64-bitHow to prevent it
Print df.memory_usage(deep=True).sum() after loading anything sizeable. Set dtypes at read time. Before any get_dummies, pivot or merge, estimate the output shape — rows × columns × bytes — and refuse to build anything you cannot justify.
Related errors
- DataLoader worker killed / shared memory — the same shortage during training
- CUDA out of memory — the GPU version
- Killed while loading checkpoint shards