Error database

MemoryError: Unable to allocate N GiB for an array

The array you asked for does not fit in RAM. Read the shape in the message first — a huge dimension usually means an accidental one-hot or pivot explosion, not a genuinely huge dataset.

The message you saw
MemoryError: Unable to allocate N GiB for an array

By Updated

The error

Output
MemoryError: Unable to allocate 3.51 GiB for an array with shape (471, 1000000) and data type float64

What it means

NumPy (or pandas on top of it) asked the operating system for one contiguous block of memory and was refused. The message does the maths for you: rows × columns × 8 bytes for float64. Read the shape before anything else — it often reveals the real bug.

A shape like (471, 1000000) rarely comes from honest data. A million columns usually means you one-hot encoded a high-cardinality column (like user IDs), or pivoted on a column with a million distinct values.

Why it happens

Three distinct situations produce this message:

  • An accidental explosion. pd.get_dummies on an ID column, a pivot on timestamps, a cross join. The fix is to not build that matrix at all.
  • Wasteful dtypes. Everything is float64 and int64 by default, twice the size of float32/int32, and text columns stored as objects are far worse.
  • Data genuinely bigger than RAM. Then the answer is chunking, not a bigger array.

Note that Windows fails fast on big allocations, while Linux may grant the memory and kill the process later — same disease, different symptom.

How to fix it

1. Sanity-check the shape. If one dimension is absurd, trace where it came from.

python
print(df.shape)
print(df["user_id"].nunique())     # about to one-hot this? that many columns

For high-cardinality categoricals, use target encoding, hashing, or an embedding — not one-hot. If you must one-hot, keep it sparse:

python
from sklearn.preprocessing import OneHotEncoder
enc = OneHotEncoder()              # sparse output by default
X = enc.fit_transform(df[["city"]])

A sparse matrix stores only the non-zero entries, which for one-hot is one per row.

2. Shrink dtypes. Usually a 2-4x saving for two lines of code.

python
df["price"] = df["price"].astype("float32")
df["qty"] = df["qty"].astype("int32")
df["city"] = df["city"].astype("category")
print(df.memory_usage(deep=True).sum() / 1e9, "GB")

3. Load only what you need.

python
df = pd.read_csv("big.csv", usecols=["date", "price", "qty"],
                 dtype={"price": "float32"})

4. Process in chunks when the data truly is big.

python
totals = {}
for chunk in pd.read_csv("big.csv", chunksize=500_000):
    for city, s in chunk.groupby("city")["price"].sum().items():
        totals[city] = totals.get(city, 0) + s

For heavier work, Polars or DuckDB handle bigger-than-RAM data with far less ceremony.

5. Check you are on 64-bit Python. 32-bit Python caps out near 2 GB regardless of your RAM.

python
import sys; print(sys.maxsize > 2**32)   # True means 64-bit

How to prevent it

Print df.memory_usage(deep=True).sum() after loading anything sizeable. Set dtypes at read time. Before any get_dummies, pivot or merge, estimate the output shape — rows × columns × bytes — and refuse to build anything you cannot justify.

The lessons behind this error.

  • Python for AI

    NumPy

    NumPy lets you do one operation to millions of numbers at once instead of one at a time. It is the foundation every AI library in Python is built on.

  • Python for AI

    Pandas

    Pandas is a table with named columns that you can filter, group and summarise in one line. It is where almost every AI project starts, because real data arrives as a table.

Back to all errors