Error database

CUDA error: device-side assert triggered

An index went outside the valid range inside a GPU kernel — usually a label outside 0..classes-1, or a token id larger than the embedding table. Restart the process and reproduce it on the CPU to find the line.

The message you saw
CUDA error: device-side assert triggered

By Updated

The error

Output
../aten/src/ATen/native/cuda/Loss.cu:250: nll_loss_forward_reduce_cuda_kernel_2d: block: [0,0,0], thread: [3,0,0] Assertion `t >= 0 && t < n_classes` failed.
RuntimeError: CUDA error: device-side assert triggered
CUDA kernel errors might be asynchronously reported at some other API call, so the stacktrace below might be incorrect. For debugging consider passing CUDA_LAUNCH_BLOCKING=1.
Compile with `TORCH_USE_CUDA_DSA` to enable device-side assertions.

The embedding version reads:

Output
../aten/src/ATen/native/cuda/Indexing.cu:1146: indexSelectLargeIndex: block: [12,0,0], thread: [64,0,0] Assertion `srcIndex < srcSelectDimSize` failed.
RuntimeError: CUDA error: device-side assert triggered

What it means

A safety check written inside a GPU kernel failed. The GPU cannot raise a Python exception, so it stops the kernel and PyTorch reports it back as this error.

The useful information is in the first line, not the last. Assertion 't >= 0 && t < n_classes' failed means a target label was negative or was greater than or equal to the number of classes. Assertion 'srcIndex < srcSelectDimSize' failed means an index used to look up a row was past the end of the table.

Two properties make this error harder than it needs to be. CUDA kernels run asynchronously, so the Python traceback usually points at a line after the real one — often at an unrelated .item() or .cpu() call, which is where the failure finally surfaced. And once an assert fires, the CUDA context is unusable: every following CUDA call in that process fails, including the ones you add to debug it.

Why it happens

Two causes account for the great majority of cases.

Labels outside the valid range, in a classification loss. nn.CrossEntropyLoss expects targets in 0 to num_classes - 1. Data labelled 1 to 10 for ten classes fails on every label 10. So does a dataset that uses -1 for "unknown", or a final layer built with the wrong number of outputs — 10 outputs for a dataset that actually contains 12 classes. Label encoders that were fitted on the training split and then applied to a test split containing an unseen category do it too.

A token id larger than the embedding table. nn.Embedding(num_embeddings, dim) has exactly num_embeddings rows. Adding special tokens to a tokenizer without calling model.resize_token_embeddings(len(tokenizer)) produces ids beyond the last row. Loading a checkpoint trained with one tokenizer and running it with a different one does the same.

Less often: a manual index into a tensor computed from data, torch.gather with an out-of-range index, or a mask built at the wrong size.

How to fix it

1. Restart the process. This is not optional and it is not a superstition. After a device-side assert the CUDA context is poisoned, so anything you try next fails with the same message regardless of whether you fixed the bug. Restart the script, or restart the notebook kernel — "Restart" rather than "Interrupt".

2. Reproduce it on the CPU. The CPU versions of these operations raise a normal Python exception with a correct traceback and a readable message.

python
device = "cpu"          # temporarily
model.to(device)
Output
IndexError: Target 10 is out of bounds.

That one line names the offending value. Running one batch on the CPU is the fastest route to the answer in almost every case.

3. If the CPU route is impractical, make the GPU report the right line.

bash
CUDA_LAUNCH_BLOCKING=1 python train.py

This forces every kernel to finish before the next call is issued, so the traceback points at the operation that actually failed. It slows training considerably, so use it for debugging and remove it afterwards.

4. Check the label range against the model's output size.

python
print("labels:", y.min().item(), "to", y.max().item())
print("model outputs:", model.fc.out_features)

num_classes = model.fc.out_features
assert y.min() >= 0, f"negative label: {y.min().item()}"
assert y.max() < num_classes, f"label {y.max().item()} needs at least {y.max().item() + 1} outputs"

If your labels run 1 to 10, shift them once, in the dataset, rather than patching them in the training loop:

python
y = y - 1                                        # 1..10 → 0..9

For string labels, let scikit-learn produce a clean mapping and use the same encoder everywhere:

python
from sklearn.preprocessing import LabelEncoder
encoder = LabelEncoder().fit(all_labels)         # fit on every label, not the train split alone
y = encoder.transform(labels)
print(len(encoder.classes_))                     # this is your num_classes

5. Check token ids against the embedding table.

python
vocab_rows = model.get_input_embeddings().weight.shape[0]
print("max id in batch:", input_ids.max().item(), "| embedding rows:", vocab_rows)
print("len(tokenizer):", len(tokenizer))

If len(tokenizer) is larger than the number of embedding rows, that is the bug, and there is one fix:

python
model.resize_token_embeddings(len(tokenizer))

Any time you call tokenizer.add_special_tokens(...) or add_tokens(...), the resize belongs on the very next line. See the padding token error for where this usually starts.

6. Note the values that are legitimately negative. -100 in a label tensor is not a bug in cross-entropy — it is the default ignore_index, used to mark positions the loss should skip. It is valid there and invalid everywhere else, so a -100 reaching an embedding lookup or a different loss function will trigger this assert.

7. Scan the whole dataset once, not only the batch that failed. One bad row in fifty thousand is enough.

python
import torch

ys = torch.cat([y for _, y in DataLoader(dataset, batch_size=512)])
print("unique labels:", torch.unique(ys))
print("min:", ys.min().item(), "max:", ys.max().item(), "count:", len(ys))

Unique values are more informative than min and max alone — they expose a stray -1 or a gap in the encoding immediately.

How to prevent it

Derive the number of classes from the data instead of typing it in:

python
num_classes = int(labels.max()) + 1              # or len(encoder.classes_)
model = Net(num_classes=num_classes)

Hardcoding 10 is what turns a new dataset into a two-hour debugging session.

Validate the label range once, when the dataset is built, and let the assertion fire there with a clear message rather than inside a CUDA kernel at 2am. The same applies to token ids: after loading a model and tokenizer, assert that len(tokenizer) matches the embedding rows.

Run the first epoch on the CPU with a small subset when you start a new project. It costs a couple of minutes and every index bug in the pipeline surfaces with a readable traceback instead of an asynchronous kernel assert.