Expected all tensors to be on the same device
One tensor is on the GPU and another is on the CPU. Find the one you forgot to move, and remember that .to() returns a new tensor rather than changing it in place.
Updated
The error
RuntimeError: Expected all tensors to be on the same device, but found at least two devices, cuda:0 and cpu! (when checking argument for argument mat1 in method wrapper_CUDA_addmm)
A close relative, which means the same thing:
RuntimeError: Input type (torch.FloatTensor) and weight type (torch.cuda.FloatTensor) should be the same
In that second message, torch.FloatTensor is a CPU tensor and torch.cuda.FloatTensor is a GPU one.
What it means
A GPU can only do arithmetic on numbers that are in GPU memory. Your operation received one tensor sitting in GPU memory and another sitting in ordinary system RAM, and PyTorch refuses to guess which one should move.
The message names the operation, which narrows the search. addmm is a matrix multiply with a bias added, so it is almost always an nn.Linear layer. conv2d points at a convolution. embedding points at an embedding lookup, where the ids are usually the ones left behind on the CPU.
Why it happens
The single biggest cause is a difference in how PyTorch treats models and tensors, and it catches nearly everyone once.
model.to(device) # works — nn.Module.to() moves the module in place
x.to(device) # does nothing useful — Tensor.to() RETURNS a new tensor
x = x.to(device) # correctnn.Module.to() modifies the module in place and returning it is a convenience. Tensor.to() cannot modify in place, so it hands you a new tensor and leaves the original where it was. A line that reads x.to(device) with no assignment is silently a no-op.
The second common cause is tensors created inside the model. Anything you build during forward defaults to the CPU, no matter where the model lives:
def forward(self, x):
mask = torch.zeros(x.size(0), x.size(1)) # created on the CPU
return self.layer(x * mask) # boomThe third is registration. Attributes holding tensors are not moved by model.to(device), and neither are submodules stored in a plain Python list. Only registered buffers, parameters and properly registered submodules travel.
The fourth appears when you load a checkpoint saved on a GPU machine onto a CPU machine, or the reverse.
How to fix it
1. Set one device variable and move both the model and every input through it.
device = "cuda" if torch.cuda.is_available() else "cpu"
model = MyModel().to(device)
for x, y in train_loader:
x, y = x.to(device), y.to(device) # note the assignment on both
loss = criterion(model(x), y)Labels count as inputs. A loss function receiving GPU predictions and CPU labels raises exactly this error.
2. Create new tensors on the device of something that is already there. Inside a model, x.device is the reliable reference.
def forward(self, x):
mask = torch.zeros(x.shape[0], x.shape[1], device=x.device, dtype=x.dtype)
scale = torch.tensor(0.5, device=x.device)
return self.layer(x * mask) * scaleHelpers such as torch.zeros_like(x), torch.ones_like(x) and torch.empty_like(x) copy device and dtype for you, which is why they are worth preferring.
3. Register constants as buffers so they move with the model. A buffer is a tensor that is part of the module's state but is not trained.
class Model(nn.Module):
def __init__(self):
super().__init__()
self.register_buffer("positions", torch.arange(512)) # moves with .to()
self.scale = torch.tensor(0.5) # does NOT moveBuffers are also saved in state_dict(), which is usually what you want for things like positional indices or normalisation constants.
4. Use nn.ModuleList instead of a Python list. A list of layers is invisible to PyTorch, so .to(device) skips it and so does the optimizer.
self.layers = nn.ModuleList([nn.Linear(64, 64) for _ in range(4)]) # correct
self.layers = [nn.Linear(64, 64) for _ in range(4)] # brokenThe dictionary equivalent is nn.ModuleDict.
5. When loading a checkpoint, say where it should land.
state = torch.load("model.pt", map_location=device, weights_only=True)
model.load_state_dict(state)
model.to(device)weights_only=True is worth including — it refuses to execute arbitrary code while unpickling, which matters for any checkpoint you did not create yourself.
6. Find the stray tensor when the traceback is not enough. Print the devices of everything involved:
print("input:", x.device, "target:", y.device)
for name, p in model.named_parameters():
print(name, p.device)
for name, b in model.named_buffers():
print(name, b.device)Any name that says cpu while the rest say cuda:0 is your answer.
How to prevent it
Decide the device once, at the top of the script, and never write a bare device string anywhere else. Pass it in rather than reaching for a global.
Inside model code, use x.device rather than a captured variable, so the model works wherever it is placed. Reach for *_like constructors by default. Put every layer in nn.Module containers rather than plain Python collections.
Two extra notes for larger setups. With Hugging Face device_map="auto", different layers of one model can live on different GPUs, so move your inputs to model.device rather than a hardcoded cuda:0. And on Apple Silicon the device string is "mps", not "cuda" — device = "mps" if torch.backends.mps.is_available() else "cpu" keeps the same code working on a Mac.
Related errors
- mat1 and mat2 shapes cannot be multiplied — the other half of PyTorch's daily errors, where devices agree but shapes do not
- torch.cuda.is_available() returns False
- CUDA out of memory
- element 0 of tensors does not require grad