CUDA out of memory (PyTorch)
PyTorch asked the GPU for more memory than was free. Lower the batch size first, then look for gradients you are keeping by accident.
Updated
The error
torch.cuda.OutOfMemoryError: CUDA out of memory. Tried to allocate 2.00 GiB. GPU 0 has a total capacity of 15.77 GiB of which 1.32 GiB is free. Process 12345 has 14.45 GiB memory in use. Of the allocated memory 13.61 GiB is allocated by PyTorch, and 412.53 MiB is reserved by PyTorch but unallocated. If reserved but unallocated memory is large try setting PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True to avoid fragmentation.
On older PyTorch versions the same problem reads:
RuntimeError: CUDA out of memory. Tried to allocate 2.00 GiB (GPU 0; 15.77 GiB total capacity; 13.61 GiB already allocated; 1.32 GiB free; 14.05 GiB reserved in total by PyTorch)
What it means
Your GPU has a fixed amount of memory, and PyTorch asked for a block that would not fit in what was left. That is the entire message. It is not a bug in your code, not a broken driver, and not a sign that your GPU is too old.
Read the numbers in the message before doing anything — they tell you which fix you need. "Tried to allocate 2.00 GiB" with "1.32 GiB is free" means you are close and a smaller batch size will clear it. "Tried to allocate 2.00 GiB" with "12 GiB is free" means something else is wrong, usually a shape far larger than you intended.
Why it happens
Three things live in GPU memory during training, and only the first is obvious.
The model weights are a fixed cost you can calculate: parameters × bytes per parameter. A 7B model in float16 is around 14 GB before anything else happens.
The activations are the big variable. Every intermediate value in the forward pass is kept, because backpropagation needs it to compute gradients. This scales with batch size and with sequence length or image size, and for transformers it can dwarf the weights.
The optimizer state is the quiet one. Adam keeps two extra values per trainable parameter, so it needs roughly twice the size of the model on top of the model and its gradients.
There is also a fourth cause that is a genuine mistake rather than a budget problem: holding on to tensors that are still attached to the computation graph. total_loss += loss keeps the entire graph of every batch so far alive, and memory climbs steadily until it dies — often at epoch 1, step 400, which looks mysterious until you know.
How to fix it
1. Halve the batch size. This is the fix in most cases, and it is one line.
train_loader = DataLoader(dataset, batch_size=16, shuffle=True) # was 32If your results depend on a larger effective batch, use gradient accumulation (step 5) to keep it without the memory.
2. Restart the process, then check the GPU is actually empty. A crashed run, a dead notebook kernel or a second script can still be holding the card.
nvidia-smiIf a python process is listed using several GB and you know it is finished, that memory is not available to you. Restarting the kernel or the process is the reliable way to release it. Calling torch.cuda.empty_cache() does not help here, because it only returns PyTorch's unused cached blocks — memory held by live tensors stays held.
3. Turn off gradient tracking for evaluation and inference. model.eval() does not do this; it changes dropout and batch-norm behaviour only. You need both lines.
model.eval()
with torch.no_grad(): # or torch.inference_mode()
for x, y in val_loader:
preds = model(x.to(device))Running validation without no_grad() stores activations for the whole validation set. This is a very common cause of "training worked, then it died during evaluation".
4. Detach anything you accumulate for logging.
total_loss = 0.0
for x, y in train_loader:
loss = criterion(model(x), y)
loss.backward()
optimizer.step()
optimizer.zero_grad(set_to_none=True)
total_loss += loss.item() # .item() gives a plain float, not a graphStoring predictions works the same way: all_preds.append(preds.detach().cpu()).
5. Use gradient accumulation to keep the effective batch size.
accum_steps = 4 # effective batch = 8 x 4 = 32
optimizer.zero_grad(set_to_none=True)
for i, (x, y) in enumerate(train_loader): # batch_size=8
loss = criterion(model(x.to(device)), y.to(device)) / accum_steps
loss.backward()
if (i + 1) % accum_steps == 0:
optimizer.step()
optimizer.zero_grad(set_to_none=True)Dividing the loss keeps the gradient magnitude the same as one large batch.
6. Train in mixed precision. Activations move to 16-bit, which usually saves a third to a half of activation memory and runs faster on any GPU from the RTX 20 series or newer.
scaler = torch.amp.GradScaler("cuda")
for x, y in train_loader:
optimizer.zero_grad(set_to_none=True)
with torch.amp.autocast("cuda", dtype=torch.float16):
loss = criterion(model(x.to(device)), y.to(device))
scaler.scale(loss).backward()
scaler.step(optimizer)
scaler.update()On an A100, H100 or RTX 30-series and newer, dtype=torch.bfloat16 is often a better choice — it has the same range as float32, so it does not need the loss scaler at all.
7. Turn on gradient checkpointing. This throws away most activations during the forward pass and recomputes them during the backward pass. It trades roughly 20-30% extra training time for a large memory saving, and it is the standard move when a model almost fits.
model.gradient_checkpointing_enable() # Hugging Face models
model.config.use_cache = False # required, cache conflicts with checkpointing8. Shorten what you feed in. Sequence length affects transformer memory more than batch size does, because attention cost grows with the square of the length. Cutting max_length from 2048 to 512 is often a larger saving than halving the batch. For images, the same applies to resolution.
9. If the message mentions fragmentation, act on it. When "reserved but unallocated" is large — say more than a gigabyte — the free memory exists but is broken into pieces too small for the request.
PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True python train.py10. For large-model inference, load in fewer bits. Quantization cuts the weights themselves.
from transformers import AutoModelForCausalLM, BitsAndBytesConfig
model = AutoModelForCausalLM.from_pretrained(
"mistralai/Mistral-7B-Instruct-v0.3",
quantization_config=BitsAndBytesConfig(load_in_4bit=True),
device_map="auto",
)How to prevent it
Measure instead of guessing. PyTorch will tell you the peak, so you can size the batch once and stop worrying.
torch.cuda.reset_peak_memory_stats()
train_one_epoch()
peak_gb = torch.cuda.max_memory_allocated() / 1024**3
print(f"peak allocated: {peak_gb:.2f} GB")A few habits keep this error rare. Aim for a peak around 80% of the card, leaving headroom for the occasional longer batch. Sort or bucket sequences by length so one 4,000-token outlier does not decide your batch size. Write .item() or .detach() reflexively when you store anything for logging. And when you move to a bigger dataset or a longer sequence length, run one epoch on a small subset first — the failure will show up in two minutes instead of two hours.
Related errors
- DataLoader worker is killed by signal / shared memory error — the same shortage, but in system RAM instead of GPU memory
- CUDA error: device-side assert triggered — another GPU error that also requires a process restart
- Expected all tensors to be on the same device
- torch.cuda.is_available() returns False