Debugging PyTorch

Expected all tensors to be on the same device

The most famous GPU error names both devices and often the guilty argument — read it properly and the one tensor that missed the transfer identifies itself.

On this page 5
  1. Why it exists
  2. How it works
  3. Where you have seen this
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

This error means one operation received tensors living on different processors — some on the graphics card, some in ordinary memory — and PyTorch refuses to mix them.

Picture a group photo where everyone gathered on the terrace, except one cousin still downstairs. The photographer will not shoot until everybody stands in the same place. The photo is not ruined; someone is in the wrong room.

The error message is the photographer shouting who is missing. It names both locations, and often names the missing person too.

Why it exists

A computer's main processor and its graphics card keep separate memories — two rooms with no shared shelves. Numbers must be carried from one room to the other, and carrying takes time.

PyTorch could carry things secretly whenever they are in the wrong room. It refuses on principle: secret carrying would hide real slowness. Instead it stops and tells you, so the carrying stays visible and deliberate. The background story lives in moving tensors between CPU and GPU.

How it works

model:   graphics card ✓
inputs:  graphics card ✓          →  photo refused:
labels:  main memory   ✗             "cuda:0 and cpu!"

Almost always, the fix is one line: carry the one forgotten tensor to where the rest already are.

Where you have seen this

This is, by a wide margin, the first error every Google Colab learner meets. Switch the free notebook to GPU, move the model over, forget the labels — and this exact sentence appears. Meeting it is practically a rite of passage.

Remember this

  • Two memories, two rooms: everything in one operation must share a room.
  • The error usually happens because one tensor missed the transfer — often the labels.
  • The message names both rooms, and frequently the guilty argument too.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install torch

Reproducing this error needs a machine with a GPU (or Colab with a GPU runtime); on CPU-only machines the code below runs without crashing, because everything lands on the CPU together. Captured with torch 2.5.1 and CUDA.

The crash, and how to read every word

device_crash.py
import torch
import torch.nn as nn

device = "cuda" if torch.cuda.is_available() else "cpu"
model = nn.Linear(8, 2).to(device)
loss_fn = nn.CrossEntropyLoss()

inputs = torch.randn(4, 8).to(device)        # inputs made the trip
labels = torch.tensor([0, 1, 1, 0])          # labels forgotten on the CPU

out = model(inputs)                          # fine: model and inputs agree
loss = loss_fn(out, labels)                  # crash: out is on cuda, labels on cpu
Output
RuntimeError: Expected all tensors to be on the same device, but found at least two devices, cuda:0 and cpu! (when checking argument for argument target in method wrapper_CUDA_nll_loss_forward)

Most people read the first half and stop. The second half is the treasure: argument target — the loss function's name for its second input, the labels. The message did not only report a mismatch; it fingered the tensor. nll_loss_forward further tells you the crash happened inside the loss, not the model.

When the parenthetical is missing or cryptic, audit by hand:

device_audit.py
for name, t in [("inputs", inputs), ("labels", labels)]:
    print(f"{name:7s} on {t.device}")
print(f"model   on {next(model.parameters()).device}")
Output
inputs  on cuda:0
labels  on cpu
model   on cuda:0

Three lines of printing beats ten minutes of staring. The fix, as usual, is one line: labels = labels.to(device).

The usual suspects

SymptomCulprit
Crash in the lossLabels not moved (inputs get remembered, labels get forgotten)
Crash in the model's forwardA tensor created inside forward with torch.zeros(...) and no device=
Crash after loading a checkpointtorch.load restored GPU tensors onto a machine or process expecting otherwise
Crash only in one exotic branchA rarely-taken code path building a CPU tensor

For tensors born inside forward, create them on the input's device: torch.zeros(n, device=x.device), or x.new_zeros(n) which copies device and dtype from x.

Common mistakes

Fixing the symptom at the crash site only. Adding .to("cuda") where it crashed leaves the next forgotten tensor for tomorrow. The durable fix is structural: one device variable at the top, every tensor and model routed through it at creation or loading time.

Hardcoding "cuda" anywhere. The file then crashes on every CPU-only machine, including colleagues' laptops and CI. Always the "cuda" if torch.cuda.is_available() else "cpu" guard.

Storing tensors in plain Python lists on a module. model.to(device) moves parameters, buffers and submodules — not an ordinary attribute holding a list of tensors. Register such state properly; buffers vs parameters explains the machinery.

Checkpoints across machines. A state dict saved from GPU tensors tries to restore to GPU. Loading on a CPU box needs torch.load(path, map_location="cpu").

Try it yourself

Break the code the other way: move the labels to device but leave the inputs on CPU. Predict which line crashes now — the model call or the loss — and what the message's parenthetical will say. Then run it and check both predictions.

What to learn next

Researcher — Mathematics and papers.

Why no implicit transfer, formally

Each tensor carries a device in its TensorOptions; the dispatcher selects kernels keyed on device (CPU, CUDA, MPS, ...), and multi-input ops require a single coherent key — the mismatch error is raised at dispatch. NumPy-style silent conversion was rejected deliberately: a hidden host-device copy costs PCIe latency plus a synchronisation, and making that cost invisible would sabotage the performance transparency the explicit-device model buys. The asymmetric exception: CPU scalars (0-dim tensors and Python numbers) are allowed into CUDA ops and are folded into kernel launches, which is why gpu_tensor + 2 works but gpu_tensor + cpu_vector does not.

The error's anatomy across backends

The wrapper_CUDA_* fragment names the dispatched kernel wrapper, and thereby the op; on Apple silicon the same class of error names mps, with the added wrinkle of float64 being unsupported there entirely. Under torch.compile, device guards are baked at trace time, so a device change between calls surfaces as a guard-failure recompile rather than this RuntimeError.

Multi-GPU: same error, one level up

With several GPUs, cuda:0 vs cuda:1 mismatches join the family. nn.DataParallel replicates modules per device and scatters inputs, so module-level attributes created on the default device crash on replicas — one of several reasons DDP (one process per device, each with a coherent single device) replaced it. In DDP, the rule collapses back to the simple one: each process behaves like a single-device program with device = f"cuda:{local_rank}".

Debugging aids

  • CUDA_VISIBLE_DEVICES=0 pins visibility when multi-GPU ambiguity is suspected.
  • Because CUDA execution is asynchronous, unrelated later lines can report device-adjacent failures; CUDA_LAUNCH_BLOCKING=1 serialises launches for an honest stack trace — the same tool used for the asserts in finding label and target bugs.
  • A blanket model.to(device) is in-place for modules but not for tensors — t.to(device) returns a new tensor. The asymmetry (documented in the tensors lesson) is behind a large share of "I moved it but it did not move" reports.

What to learn next