Reading and fixing CUDA out of memory
The CUDA out of memory error tells you exactly what did not fit and where the space went — once you learn to read it, the fix is usually one of five known moves.
- 7 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
"CUDA out of memory" means the GPU's storage is full, and the error message itself tells you what would not fit.
Think of packing a fridge for a wedding. The fridge has fixed space. One more giant pot arrives, it does not fit, and it goes back to the kitchen. The pot being refused does not mean the pot caused the problem — the fridge was already nearly full of everything packed before it.
This is the most common crash in deep learning. Everyone who trains models meets it, usually in their first week.
Why this error exists
A GPU has its own separate memory, and it is small compared to your computer's main storage. A laptop easily has hundreds of gigabytes of disk. A typical free Colab GPU has around 15 gigabytes of working memory, total.
Training fills that space fast. The model's knowledge lives there. So do its ongoing corrections, and all the in-between results for the batch being processed.
PyTorch refuses loudly instead of slowing down quietly. That is a kindness, even when it does not feel like one.
How to read the message
The error always follows the same shape:
Tried to allocate 2.50 GiB <- the pot that did not fit
GPU 0 has a total capacity of 15 GiB <- the size of the fridge
of which 1.20 GiB is free <- the space left when it arrivedThe "tried to allocate" number is rarely the villain. The real question is: what filled the other thirteen gigabytes? The next lesson answers that properly.
A real example you have seen
Every Colab user knows the ritual. Training runs fine for days. You double the batch size to go faster, and the very first step crashes with this error. Nothing is broken. The fridge is the same size — you packed a bigger pot.
Remember this
- The GPU's memory is separate and small; the error means it is full.
- The message names the failed request, the total space, and the free space.
- The failed request is usually not the cause — the space was eaten earlier.
What to learn next
- What is using your GPU memory — the full breakdown, measured live.
- Gradient accumulation — big-batch results on a small card.
- Mixed precision training — the first structural memory saving.
Developer — Code and libraries.
Setup
pip install torchTriggering the error needs a GPU. The reading and the fixes below work everywhere. Output captured with torch 2.5.1 on an NVIDIA RTX A6000 (48 GB); your capacity numbers will differ.
Trigger it on purpose, safely
Meeting the error deliberately removes its fear. This asks for far more than any card has:
import torch
if not torch.cuda.is_available():
print("no GPU on this machine - the message below is from a GPU run")
else:
try:
# 64 * 1024^3 float32 values = 256 GiB, more than any single card has
x = torch.empty(64, 1024, 1024, 1024, device="cuda")
except torch.OutOfMemoryError as e:
print(type(e).__name__)
print(str(e).split(" If reserved")[0])OutOfMemoryError CUDA out of memory. Tried to allocate 256.00 GiB. GPU 0 has a total capacity of 44.99 GiB of which 43.47 GiB is free. Of the allocated memory 0 bytes is allocated by PyTorch, and 0 bytes is reserved by PyTorch but unallocated.
Note that it is a catchable Python exception. A serving system can catch it and retry with a smaller batch instead of dying.
The three numbers that matter
Tried to allocate — the one request that failed. Total capacity — your card's size. Free — what was left. When free is nearly zero and the request is modest, your training state fills the card. When the request itself is huge, you built an accidentally enormous tensor — usually a wrong broadcast or a missing reshape.
One more suspect pair appears in the full message: allocated by PyTorch versus reserved but unallocated. A large reserved-but-unallocated number means fragmentation — free space exists, but in scattered pieces too small for the request. The message suggests expandable_segments:True for exactly that case.
The fixes, in the order to try them
- Halve the batch size. Activation memory scales with batch size. This is the honest first move, and gradient accumulation rebuilds the effective batch for free.
- Wrap evaluation in
torch.no_grad(). Without it, evaluation builds the autograd graph and stores activations nobody will use. See detach vs no_grad. - Stop hoarding history.
losses.append(loss)keeps every computation graph alive. Storeloss.item()instead — a plain number. - Turn on mixed precision. Activations halve in size.
- Turn on gradient checkpointing. Trades some compute for a large activation saving.
Common mistakes
Calling torch.cuda.empty_cache() and hoping. It returns cached blocks to the driver. It does not free a single live tensor — your model, gradients and stored activations remain. It helps another process sharing the card; it almost never fixes your own OOM.
Believing the failed line is the guilty line. The allocation that fails is whichever one arrived after the card filled. Fix the biggest consumer, not the last arrival.
Shrinking the model when the batch is the problem. Do the arithmetic in the next lesson first. In training, activation memory usually dwarfs parameter memory.
A leak across epochs. Memory that climbs every epoch points to references kept alive — often a list of tensors, or a logger storing whole outputs.
Try it yourself
Run the script on Colab with a GPU runtime. Then change the shape to (2, 1024, 1024, 1024) — 8 GiB — and predict all three numbers in the message before running.
What to learn next
- What is using your GPU memory — the full breakdown, measured live.
- Gradient accumulation — big-batch results on a small card.
- Mixed precision training — the first structural memory saving.
Researcher — Mathematics and papers.
The allocator behind the message
PyTorch does not call cudaMalloc per tensor — raw device allocation synchronises and is slow. The caching allocator requests large segments from the driver, then serves tensor allocations from them, splitting and coalescing blocks internally. Freed tensors return blocks to the cache, not to the driver, which is why nvidia-smi reports more than the sum of live tensors.
Consequences:
- The OOM condition is "no suitable cached block and the driver refused a new segment", so fragmentation can fail a request smaller than total free memory.
- A stream of slightly-growing tensors (variable sequence lengths, for one) defeats block reuse and fragments the pool. Bucketing batch shapes measurably reduces peak reserved memory.
PYTORCH_CUDA_ALLOC_CONF=expandable_segments:Trueswitches to virtual-memory-backed segments that grow in place, removing most fragmentation failures. The older knobmax_split_size_mbforbids splitting the largest blocks.
Introspection: torch.cuda.memory_allocated(), memory_reserved(), max_memory_allocated(), torch.cuda.memory_summary(), and torch.cuda.memory._record_memory_history(), whose snapshots render at pytorch.org/memory_viz — the fastest route to "what exactly filled the card".
The model that predicts the crash
Peak training memory is approximately:
$$ M \approx P \cdot (b_p + b_g + b_o) + A(B) $$
- $P$ — parameter count.
- $b_p, b_g, b_o$ — bytes per parameter for weights, gradients, optimizer state. Float32 Adam: $4 + 4 + 8$.
- $A(B)$ — activation memory, linear in batch size $B$, and linear in sequence length for transformers once FlashAttention removes the quadratic term (see scaled_dot_product_attention).
The batch-independent term explains why a 7-billion-parameter model refuses to train on a 24 GB card at any batch size: $7\text{B} \times 16$ bytes = 112 GB before the first sample arrives. That arithmetic, not trial and error, should pick your strategy — see the breakdown and FSDP.
References
- Paszke et al. (2019), PyTorch: An Imperative Style, High-Performance Deep Learning Library, NeurIPS — §5.3 documents the caching allocator design.
- Rajbhandari et al. (2020), ZeRO: Memory Optimizations Toward Training Trillion Parameter Models — the canonical accounting of where training memory goes, and how to shard it away.
What to learn next
- What is using your GPU memory — the full breakdown, measured live.
- Gradient accumulation — big-batch results on a small card.
- Mixed precision training — the first structural memory saving.