Expected all tensors to be on the same device
The most famous GPU error names both devices and often the guilty argument — read it properly and the one tensor that missed the transfer identifies itself.
- 7 min read
- 3 reading levels
- Published
Read these first
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
This error means one operation received tensors living on different processors — some on the graphics card, some in ordinary memory — and PyTorch refuses to mix them.
Picture a group photo where everyone gathered on the terrace, except one cousin still downstairs. The photographer will not shoot until everybody stands in the same place. The photo is not ruined; someone is in the wrong room.
The error message is the photographer shouting who is missing. It names both locations, and often names the missing person too.
Why it exists
A computer's main processor and its graphics card keep separate memories — two rooms with no shared shelves. Numbers must be carried from one room to the other, and carrying takes time.
PyTorch could carry things secretly whenever they are in the wrong room. It refuses on principle: secret carrying would hide real slowness. Instead it stops and tells you, so the carrying stays visible and deliberate. The background story lives in moving tensors between CPU and GPU.
How it works
model: graphics card ✓
inputs: graphics card ✓ → photo refused:
labels: main memory ✗ "cuda:0 and cpu!"Almost always, the fix is one line: carry the one forgotten tensor to where the rest already are.
Where you have seen this
This is, by a wide margin, the first error every Google Colab learner meets. Switch the free notebook to GPU, move the model over, forget the labels — and this exact sentence appears. Meeting it is practically a rite of passage.
Remember this
- Two memories, two rooms: everything in one operation must share a room.
- The error usually happens because one tensor missed the transfer — often the labels.
- The message names both rooms, and frequently the guilty argument too.
What to learn next
- Expected scalar type Float but found Double — the same family of error, for types instead of places.
- Moving tensors between CPU and GPU — the transfer mechanics and their costs.
- Buffers vs parameters — making
model.to(device)move everything it should.
Developer — Code and libraries.
Setup
pip install torchReproducing this error needs a machine with a GPU (or Colab with a GPU runtime); on CPU-only machines the code below runs without crashing, because everything lands on the CPU together. Captured with torch 2.5.1 and CUDA.
The crash, and how to read every word
import torch
import torch.nn as nn
device = "cuda" if torch.cuda.is_available() else "cpu"
model = nn.Linear(8, 2).to(device)
loss_fn = nn.CrossEntropyLoss()
inputs = torch.randn(4, 8).to(device) # inputs made the trip
labels = torch.tensor([0, 1, 1, 0]) # labels forgotten on the CPU
out = model(inputs) # fine: model and inputs agree
loss = loss_fn(out, labels) # crash: out is on cuda, labels on cpuRuntimeError: Expected all tensors to be on the same device, but found at least two devices, cuda:0 and cpu! (when checking argument for argument target in method wrapper_CUDA_nll_loss_forward)
Most people read the first half and stop. The second half is the treasure: argument target — the loss function's name for its second input, the labels. The message did not only report a mismatch; it fingered the tensor. nll_loss_forward further tells you the crash happened inside the loss, not the model.
When the parenthetical is missing or cryptic, audit by hand:
for name, t in [("inputs", inputs), ("labels", labels)]:
print(f"{name:7s} on {t.device}")
print(f"model on {next(model.parameters()).device}")inputs on cuda:0 labels on cpu model on cuda:0
Three lines of printing beats ten minutes of staring. The fix, as usual, is one line: labels = labels.to(device).
The usual suspects
| Symptom | Culprit |
|---|---|
| Crash in the loss | Labels not moved (inputs get remembered, labels get forgotten) |
| Crash in the model's forward | A tensor created inside forward with torch.zeros(...) and no device= |
| Crash after loading a checkpoint | torch.load restored GPU tensors onto a machine or process expecting otherwise |
| Crash only in one exotic branch | A rarely-taken code path building a CPU tensor |
For tensors born inside forward, create them on the input's device: torch.zeros(n, device=x.device), or x.new_zeros(n) which copies device and dtype from x.
Common mistakes
Fixing the symptom at the crash site only. Adding .to("cuda") where it crashed leaves the next forgotten tensor for tomorrow. The durable fix is structural: one device variable at the top, every tensor and model routed through it at creation or loading time.
Hardcoding "cuda" anywhere. The file then crashes on every CPU-only machine, including colleagues' laptops and CI. Always the "cuda" if torch.cuda.is_available() else "cpu" guard.
Storing tensors in plain Python lists on a module. model.to(device) moves parameters, buffers and submodules — not an ordinary attribute holding a list of tensors. Register such state properly; buffers vs parameters explains the machinery.
Checkpoints across machines. A state dict saved from GPU tensors tries to restore to GPU. Loading on a CPU box needs torch.load(path, map_location="cpu").
Try it yourself
Break the code the other way: move the labels to device but leave the inputs on CPU. Predict which line crashes now — the model call or the loss — and what the message's parenthetical will say. Then run it and check both predictions.
What to learn next
- Expected scalar type Float but found Double — the same family of error, for types instead of places.
- Moving tensors between CPU and GPU — the transfer mechanics and their costs.
- Buffers vs parameters — making
model.to(device)move everything it should.
Researcher — Mathematics and papers.
Why no implicit transfer, formally
Each tensor carries a device in its TensorOptions; the dispatcher selects kernels keyed on device (CPU, CUDA, MPS, ...), and multi-input ops require a single coherent key — the mismatch error is raised at dispatch. NumPy-style silent conversion was rejected deliberately: a hidden host-device copy costs PCIe latency plus a synchronisation, and making that cost invisible would sabotage the performance transparency the explicit-device model buys. The asymmetric exception: CPU scalars (0-dim tensors and Python numbers) are allowed into CUDA ops and are folded into kernel launches, which is why gpu_tensor + 2 works but gpu_tensor + cpu_vector does not.
The error's anatomy across backends
The wrapper_CUDA_* fragment names the dispatched kernel wrapper, and thereby the op; on Apple silicon the same class of error names mps, with the added wrinkle of float64 being unsupported there entirely. Under torch.compile, device guards are baked at trace time, so a device change between calls surfaces as a guard-failure recompile rather than this RuntimeError.
Multi-GPU: same error, one level up
With several GPUs, cuda:0 vs cuda:1 mismatches join the family. nn.DataParallel replicates modules per device and scatters inputs, so module-level attributes created on the default device crash on replicas — one of several reasons DDP (one process per device, each with a coherent single device) replaced it. In DDP, the rule collapses back to the simple one: each process behaves like a single-device program with device = f"cuda:{local_rank}".
Debugging aids
CUDA_VISIBLE_DEVICES=0pins visibility when multi-GPU ambiguity is suspected.- Because CUDA execution is asynchronous, unrelated later lines can report device-adjacent failures;
CUDA_LAUNCH_BLOCKING=1serialises launches for an honest stack trace — the same tool used for the asserts in finding label and target bugs. - A blanket
model.to(device)is in-place for modules but not for tensors —t.to(device)returns a new tensor. The asymmetry (documented in the tensors lesson) is behind a large share of "I moved it but it did not move" reports.
What to learn next
- Expected scalar type Float but found Double — the same family of error, for types instead of places.
- Moving tensors between CPU and GPU — the transfer mechanics and their costs.
- Buffers vs parameters — making
model.to(device)move everything it should.