Error database

TypeError: can't convert cuda:0 device type tensor to numpy

NumPy works in CPU memory, and your tensor lives on the GPU. Chain .detach().cpu().numpy() — and do the move once, not inside every metric call.

The message you saw
TypeError: can't convert cuda:0 device type tensor to numpy

By Updated

The error

Output
TypeError: can't convert cuda:0 device type tensor to numpy. Use Tensor.cpu() to copy the tensor to host memory first.

Its close sibling appears for tensors still attached to the graph:

Output
RuntimeError: Can't call numpy() on Tensor that requires grad. Use tensor.detach().numpy() instead.

What it means

NumPy arrays live in ordinary CPU memory. Your tensor lives in GPU memory, and possibly carries gradient history. Converting to NumPy therefore needs up to two explicit steps: drop the gradient link, copy to CPU. PyTorch makes both explicit because both have a cost you should see.

Why it happens

The conversion is often invisible. plt.plot(preds), np.mean(preds), pd.DataFrame(preds) and sklearn.metrics.accuracy_score(y, preds) all call NumPy under the hood. Training on GPU means every tensor coming out of the model is a CUDA tensor, and any of those innocent lines triggers the error.

How to fix it

1. The full incantation, safe in all cases.

python
arr = preds.detach().cpu().numpy()

.detach() drops the autograd link, .cpu() copies off the GPU, .numpy() converts. Extra calls are free when not needed, so this exact chain is the reflex worth memorising.

2. In evaluation loops, wrap with no_grad and move once per batch.

python
model.eval()
all_preds, all_targets = [], []
with torch.no_grad():
    for x, y in val_loader:
        out = model(x.to(device))
        all_preds.append(out.argmax(dim=1).cpu())
        all_targets.append(y)
preds = torch.cat(all_preds).numpy()
targets = torch.cat(all_targets).numpy()

Inside no_grad() nothing requires grad, so .detach() becomes unnecessary — and you save the memory of the graph entirely.

3. For a single number, .item() is shorter.

python
acc = (out.argmax(1) == y).float().mean().item()

4. Do not sprinkle .cpu() inside the training step. Each transfer synchronises the GPU and stalls the pipeline. Compute on GPU, transfer at the end — once per batch for metrics, once per epoch for anything heavier.

How to prevent it

Keep a clean boundary in your code: tensors on GPU inside the model and loss; NumPy only at the edges (logging, plotting, metrics from other libraries). When a function needs NumPy, convert at its call site with the full .detach().cpu().numpy() chain, and resist converting back and forth.

The lessons behind this error.

  • Python for AI

    NumPy

    NumPy lets you do one operation to millions of numbers at once instead of one at a time. It is the foundation every AI library in Python is built on.

Back to all errors