TypeError: can't convert cuda:0 device type tensor to numpy
NumPy works in CPU memory, and your tensor lives on the GPU. Chain .detach().cpu().numpy() — and do the move once, not inside every metric call.
Updated
The error
TypeError: can't convert cuda:0 device type tensor to numpy. Use Tensor.cpu() to copy the tensor to host memory first.
Its close sibling appears for tensors still attached to the graph:
RuntimeError: Can't call numpy() on Tensor that requires grad. Use tensor.detach().numpy() instead.
What it means
NumPy arrays live in ordinary CPU memory. Your tensor lives in GPU memory, and possibly carries gradient history. Converting to NumPy therefore needs up to two explicit steps: drop the gradient link, copy to CPU. PyTorch makes both explicit because both have a cost you should see.
Why it happens
The conversion is often invisible. plt.plot(preds), np.mean(preds), pd.DataFrame(preds) and sklearn.metrics.accuracy_score(y, preds) all call NumPy under the hood. Training on GPU means every tensor coming out of the model is a CUDA tensor, and any of those innocent lines triggers the error.
How to fix it
1. The full incantation, safe in all cases.
arr = preds.detach().cpu().numpy().detach() drops the autograd link, .cpu() copies off the GPU, .numpy() converts. Extra calls are free when not needed, so this exact chain is the reflex worth memorising.
2. In evaluation loops, wrap with no_grad and move once per batch.
model.eval()
all_preds, all_targets = [], []
with torch.no_grad():
for x, y in val_loader:
out = model(x.to(device))
all_preds.append(out.argmax(dim=1).cpu())
all_targets.append(y)
preds = torch.cat(all_preds).numpy()
targets = torch.cat(all_targets).numpy()Inside no_grad() nothing requires grad, so .detach() becomes unnecessary — and you save the memory of the graph entirely.
3. For a single number, .item() is shorter.
acc = (out.argmax(1) == y).float().mean().item()4. Do not sprinkle .cpu() inside the training step. Each transfer synchronises the GPU and stalls the pipeline. Compute on GPU, transfer at the end — once per batch for metrics, once per epoch for anything heavier.
How to prevent it
Keep a clean boundary in your code: tensors on GPU inside the model and loss; NumPy only at the edges (logging, plotting, metrics from other libraries). When a function needs NumPy, convert at its call site with the full .detach().cpu().numpy() chain, and resist converting back and forth.
Related errors
- Expected all tensors to be on the same device — the same CPU/GPU boundary, hit from the other side
- expected scalar type Float but found Double — what NumPy's float64 does on the way back in
- CUDA out of memory — evaluation without no_grad is a cause of both errors