CUDA error: device-side assert triggered
An index went outside the valid range inside a GPU kernel — usually a label outside 0..classes-1, or a token id larger than the embedding table. Restart the process and reproduce it on the CPU to find the line.
Updated
The error
../aten/src/ATen/native/cuda/Loss.cu:250: nll_loss_forward_reduce_cuda_kernel_2d: block: [0,0,0], thread: [3,0,0] Assertion `t >= 0 && t < n_classes` failed. RuntimeError: CUDA error: device-side assert triggered CUDA kernel errors might be asynchronously reported at some other API call, so the stacktrace below might be incorrect. For debugging consider passing CUDA_LAUNCH_BLOCKING=1. Compile with `TORCH_USE_CUDA_DSA` to enable device-side assertions.
The embedding version reads:
../aten/src/ATen/native/cuda/Indexing.cu:1146: indexSelectLargeIndex: block: [12,0,0], thread: [64,0,0] Assertion `srcIndex < srcSelectDimSize` failed. RuntimeError: CUDA error: device-side assert triggered
What it means
A safety check written inside a GPU kernel failed. The GPU cannot raise a Python exception, so it stops the kernel and PyTorch reports it back as this error.
The useful information is in the first line, not the last. Assertion 't >= 0 && t < n_classes' failed means a target label was negative or was greater than or equal to the number of classes. Assertion 'srcIndex < srcSelectDimSize' failed means an index used to look up a row was past the end of the table.
Two properties make this error harder than it needs to be. CUDA kernels run asynchronously, so the Python traceback usually points at a line after the real one — often at an unrelated .item() or .cpu() call, which is where the failure finally surfaced. And once an assert fires, the CUDA context is unusable: every following CUDA call in that process fails, including the ones you add to debug it.
Why it happens
Two causes account for the great majority of cases.
Labels outside the valid range, in a classification loss. nn.CrossEntropyLoss expects targets in 0 to num_classes - 1. Data labelled 1 to 10 for ten classes fails on every label 10. So does a dataset that uses -1 for "unknown", or a final layer built with the wrong number of outputs — 10 outputs for a dataset that actually contains 12 classes. Label encoders that were fitted on the training split and then applied to a test split containing an unseen category do it too.
A token id larger than the embedding table. nn.Embedding(num_embeddings, dim) has exactly num_embeddings rows. Adding special tokens to a tokenizer without calling model.resize_token_embeddings(len(tokenizer)) produces ids beyond the last row. Loading a checkpoint trained with one tokenizer and running it with a different one does the same.
Less often: a manual index into a tensor computed from data, torch.gather with an out-of-range index, or a mask built at the wrong size.
How to fix it
1. Restart the process. This is not optional and it is not a superstition. After a device-side assert the CUDA context is poisoned, so anything you try next fails with the same message regardless of whether you fixed the bug. Restart the script, or restart the notebook kernel — "Restart" rather than "Interrupt".
2. Reproduce it on the CPU. The CPU versions of these operations raise a normal Python exception with a correct traceback and a readable message.
device = "cpu" # temporarily
model.to(device)IndexError: Target 10 is out of bounds.
That one line names the offending value. Running one batch on the CPU is the fastest route to the answer in almost every case.
3. If the CPU route is impractical, make the GPU report the right line.
CUDA_LAUNCH_BLOCKING=1 python train.pyThis forces every kernel to finish before the next call is issued, so the traceback points at the operation that actually failed. It slows training considerably, so use it for debugging and remove it afterwards.
4. Check the label range against the model's output size.
print("labels:", y.min().item(), "to", y.max().item())
print("model outputs:", model.fc.out_features)
num_classes = model.fc.out_features
assert y.min() >= 0, f"negative label: {y.min().item()}"
assert y.max() < num_classes, f"label {y.max().item()} needs at least {y.max().item() + 1} outputs"If your labels run 1 to 10, shift them once, in the dataset, rather than patching them in the training loop:
y = y - 1 # 1..10 → 0..9For string labels, let scikit-learn produce a clean mapping and use the same encoder everywhere:
from sklearn.preprocessing import LabelEncoder
encoder = LabelEncoder().fit(all_labels) # fit on every label, not the train split alone
y = encoder.transform(labels)
print(len(encoder.classes_)) # this is your num_classes5. Check token ids against the embedding table.
vocab_rows = model.get_input_embeddings().weight.shape[0]
print("max id in batch:", input_ids.max().item(), "| embedding rows:", vocab_rows)
print("len(tokenizer):", len(tokenizer))If len(tokenizer) is larger than the number of embedding rows, that is the bug, and there is one fix:
model.resize_token_embeddings(len(tokenizer))Any time you call tokenizer.add_special_tokens(...) or add_tokens(...), the resize belongs on the very next line. See the padding token error for where this usually starts.
6. Note the values that are legitimately negative. -100 in a label tensor is not a bug in cross-entropy — it is the default ignore_index, used to mark positions the loss should skip. It is valid there and invalid everywhere else, so a -100 reaching an embedding lookup or a different loss function will trigger this assert.
7. Scan the whole dataset once, not only the batch that failed. One bad row in fifty thousand is enough.
import torch
ys = torch.cat([y for _, y in DataLoader(dataset, batch_size=512)])
print("unique labels:", torch.unique(ys))
print("min:", ys.min().item(), "max:", ys.max().item(), "count:", len(ys))Unique values are more informative than min and max alone — they expose a stray -1 or a gap in the encoding immediately.
How to prevent it
Derive the number of classes from the data instead of typing it in:
num_classes = int(labels.max()) + 1 # or len(encoder.classes_)
model = Net(num_classes=num_classes)Hardcoding 10 is what turns a new dataset into a two-hour debugging session.
Validate the label range once, when the dataset is built, and let the assertion fire there with a clear message rather than inside a CUDA kernel at 2am. The same applies to token ids: after loading a model and tokenizer, assert that len(tokenizer) matches the embedding rows.
Run the first epoch on the CPU with a small subset when you start a new project. It costs a couple of minutes and every index bug in the pipeline surfaces with a readable traceback instead of an asynchronous kernel assert.
Related errors
- mat1 and mat2 shapes cannot be multiplied — a shape problem that sometimes appears as this assert instead
- Asking to pad but the tokenizer does not have a padding token — adding tokens without resizing leads directly here
- CUDA out of memory — the other GPU error that needs a process restart
- Loss is NaN during training