Error database

Expected all tensors to be on the same device

One tensor is on the GPU and another is on the CPU. Find the one you forgot to move, and remember that .to() returns a new tensor rather than changing it in place.

The message you saw
Expected all tensors to be on the same device

By Updated

The error

Output
RuntimeError: Expected all tensors to be on the same device, but found at least two devices, cuda:0 and cpu! (when checking argument for argument mat1 in method wrapper_CUDA_addmm)

A close relative, which means the same thing:

Output
RuntimeError: Input type (torch.FloatTensor) and weight type (torch.cuda.FloatTensor) should be the same

In that second message, torch.FloatTensor is a CPU tensor and torch.cuda.FloatTensor is a GPU one.

What it means

A GPU can only do arithmetic on numbers that are in GPU memory. Your operation received one tensor sitting in GPU memory and another sitting in ordinary system RAM, and PyTorch refuses to guess which one should move.

The message names the operation, which narrows the search. addmm is a matrix multiply with a bias added, so it is almost always an nn.Linear layer. conv2d points at a convolution. embedding points at an embedding lookup, where the ids are usually the ones left behind on the CPU.

Why it happens

The single biggest cause is a difference in how PyTorch treats models and tensors, and it catches nearly everyone once.

python
model.to(device)      # works — nn.Module.to() moves the module in place
x.to(device)          # does nothing useful — Tensor.to() RETURNS a new tensor
x = x.to(device)      # correct

nn.Module.to() modifies the module in place and returning it is a convenience. Tensor.to() cannot modify in place, so it hands you a new tensor and leaves the original where it was. A line that reads x.to(device) with no assignment is silently a no-op.

The second common cause is tensors created inside the model. Anything you build during forward defaults to the CPU, no matter where the model lives:

python
def forward(self, x):
    mask = torch.zeros(x.size(0), x.size(1))     # created on the CPU
    return self.layer(x * mask)                  # boom

The third is registration. Attributes holding tensors are not moved by model.to(device), and neither are submodules stored in a plain Python list. Only registered buffers, parameters and properly registered submodules travel.

The fourth appears when you load a checkpoint saved on a GPU machine onto a CPU machine, or the reverse.

How to fix it

1. Set one device variable and move both the model and every input through it.

python
device = "cuda" if torch.cuda.is_available() else "cpu"
model = MyModel().to(device)

for x, y in train_loader:
    x, y = x.to(device), y.to(device)     # note the assignment on both
    loss = criterion(model(x), y)

Labels count as inputs. A loss function receiving GPU predictions and CPU labels raises exactly this error.

2. Create new tensors on the device of something that is already there. Inside a model, x.device is the reliable reference.

python
def forward(self, x):
    mask = torch.zeros(x.shape[0], x.shape[1], device=x.device, dtype=x.dtype)
    scale = torch.tensor(0.5, device=x.device)
    return self.layer(x * mask) * scale

Helpers such as torch.zeros_like(x), torch.ones_like(x) and torch.empty_like(x) copy device and dtype for you, which is why they are worth preferring.

3. Register constants as buffers so they move with the model. A buffer is a tensor that is part of the module's state but is not trained.

python
class Model(nn.Module):
    def __init__(self):
        super().__init__()
        self.register_buffer("positions", torch.arange(512))   # moves with .to()
        self.scale = torch.tensor(0.5)                          # does NOT move

Buffers are also saved in state_dict(), which is usually what you want for things like positional indices or normalisation constants.

4. Use nn.ModuleList instead of a Python list. A list of layers is invisible to PyTorch, so .to(device) skips it and so does the optimizer.

python
self.layers = nn.ModuleList([nn.Linear(64, 64) for _ in range(4)])   # correct
self.layers = [nn.Linear(64, 64) for _ in range(4)]                  # broken

The dictionary equivalent is nn.ModuleDict.

5. When loading a checkpoint, say where it should land.

python
state = torch.load("model.pt", map_location=device, weights_only=True)
model.load_state_dict(state)
model.to(device)

weights_only=True is worth including — it refuses to execute arbitrary code while unpickling, which matters for any checkpoint you did not create yourself.

6. Find the stray tensor when the traceback is not enough. Print the devices of everything involved:

python
print("input:", x.device, "target:", y.device)
for name, p in model.named_parameters():
    print(name, p.device)
for name, b in model.named_buffers():
    print(name, b.device)

Any name that says cpu while the rest say cuda:0 is your answer.

How to prevent it

Decide the device once, at the top of the script, and never write a bare device string anywhere else. Pass it in rather than reaching for a global.

Inside model code, use x.device rather than a captured variable, so the model works wherever it is placed. Reach for *_like constructors by default. Put every layer in nn.Module containers rather than plain Python collections.

Two extra notes for larger setups. With Hugging Face device_map="auto", different layers of one model can live on different GPUs, so move your inputs to model.device rather than a hardcoded cuda:0. And on Apple Silicon the device string is "mps", not "cuda" — device = "mps" if torch.backends.mps.is_available() else "cpu" keeps the same code working on a Mac.