mat1 and mat2 shapes cannot be multiplied (shape mismatch)
A layer expected a different number of input features than it received. Print the shape going in, and set in_features to match.
Updated
The error
RuntimeError: mat1 and mat2 shapes cannot be multiplied (64x784 and 512x10)
Older PyTorch versions phrase it differently:
RuntimeError: size mismatch, m1: [64 x 784], m2: [512 x 10]
And the convolution version of the same class of problem:
RuntimeError: Given groups=1, weight of size [64, 3, 3, 3], expected input[32, 1, 28, 28] to have 3 channels, but got 1 channels instead
What it means
Matrix multiplication has one rule: the width of the first matrix must equal the height of the second. Your data arrived as 64 rows of 784 numbers, and the layer's weight matrix expects rows of 512 numbers. 784 and 512 do not match, so the operation is not defined.
Read the numbers this way:
(64 x 784) @ (512 x 10)
▲ ▲
└────────────┘
these must be equal — they are notThe 64 is your batch size and the 10 is the output size; neither is the problem. The mismatch is always the two inner numbers.
One detail that saves confusion: for nn.Linear(in_features, out_features), the second matrix prints as (in_features x out_features). So (512 x 10) means the layer was built as nn.Linear(512, 10). Your data has 784 features and the layer wants 512.
Why it happens
Almost always, a shape changed somewhere earlier and the layer definition was not updated to follow.
The classic case is a missing flatten. A convolution stack outputs a 4-dimensional tensor of shape (batch, channels, height, width), and a linear layer needs a 2-dimensional (batch, features). Without a flatten step, or with the wrong flattened size hardcoded, the numbers cannot line up.
Equally common: the flattened size was correct for one input resolution and you changed the resolution. A 28×28 image through two convolutions and two poolings gives a different feature count than a 32×32 one, so nn.Linear(1600, 10) becomes wrong the moment you switch datasets.
Then there are the smaller ones. Two tensors passed to torch.matmul in the wrong order. An RNN or transformer where batch_first is not what you assumed, so the batch and sequence dimensions are swapped. An embedding dimension that no longer matches the layer after you changed models. A pretrained backbone whose final feature size is not what you remembered — ResNet-18 ends at 512, ResNet-50 at 2048, and swapping one for the other produces exactly this error.
How to fix it
1. Print the shape immediately before the failing layer. Do this before theorising. It takes fifteen seconds and it usually ends the investigation.
def forward(self, x):
x = self.features(x)
print("after features:", x.shape) # e.g. torch.Size([64, 64, 7, 7])
x = x.flatten(1)
print("after flatten :", x.shape) # e.g. torch.Size([64, 3136])
return self.classifier(x)The second number after flattening is what in_features must be.
2. Compute the right in_features instead of guessing. Run a dummy tensor of your real input size through the feature part once, at construction time:
import torch
import torch.nn as nn
class Net(nn.Module):
def __init__(self, input_shape=(3, 32, 32), num_classes=10):
super().__init__()
self.features = nn.Sequential(
nn.Conv2d(input_shape[0], 32, 3, padding=1), nn.ReLU(), nn.MaxPool2d(2),
nn.Conv2d(32, 64, 3, padding=1), nn.ReLU(), nn.MaxPool2d(2),
)
# ask the network how big its own output is, rather than working it out on paper
with torch.no_grad():
n_features = self.features(torch.zeros(1, *input_shape)).flatten(1).shape[1]
self.classifier = nn.Linear(n_features, num_classes)
def forward(self, x):
return self.classifier(self.features(x).flatten(1))
print(Net()(torch.randn(8, 3, 32, 32)).shape)torch.Size([8, 10])
This model now survives a change of input resolution without any edit.
3. Flatten correctly. x.flatten(1) keeps dimension 0 (the batch) and merges everything after it. Inside an nn.Sequential, use nn.Flatten().
x = x.flatten(1) # (64, 64, 7, 7) → (64, 3136)
x = x.view(x.size(0), -1) # same result; view needs contiguous memory, flatten does notAvoid x.view(-1, 3136) with a hardcoded number. When the size is wrong it silently changes your batch size instead of failing, which produces a much harder bug later.
4. Make the model resolution-independent with adaptive pooling. This is what torchvision's own models do, and it removes the problem at the source.
self.pool = nn.AdaptiveAvgPool2d((1, 1)) # any H x W in, always 1 x 1 out
self.classifier = nn.Linear(64, num_classes) # 64 = channel count, fixed5. If the input really is transposed, transpose it — after confirming.
print(a.shape, b.shape) # (64, 512) (10, 512)
out = a @ b.T # (64, 512) @ (512, 10) → (64, 10)Only reach for this once the printed shapes show a genuine transpose. Adding a .T to make an error disappear, when the real problem is a wrong layer size, produces a model that trains on nonsense.
6. Check batch_first for recurrent and transformer layers. PyTorch defaults to (seq, batch, feature), while most people prepare data as (batch, seq, feature).
self.lstm = nn.LSTM(input_size=64, hidden_size=128, batch_first=True)7. For the convolution channel version of this error, the number after expected input[...] is what your data has and the second number in weight of size [...] is what the layer wants. Grayscale images have 1 channel, colour has 3, so nn.Conv2d(1, 32, 3) for MNIST and nn.Conv2d(3, 32, 3) for CIFAR-10 or photographs.
How to prevent it
Write a shape test and run it before training. It costs one second and catches this at the moment you edit the model rather than twenty minutes into an epoch:
def test_shapes():
model = Net(input_shape=(3, 32, 32), num_classes=10)
out = model(torch.randn(2, 3, 32, 32))
assert out.shape == (2, 10), out.shape
print("shapes fine")
test_shapes()Two more habits help. Pass the input shape and class count into the model as arguments rather than burying constants in the layer definitions. And when you build a new architecture, add a temporary print(x.shape) after every stage, run one batch, then delete them — five prints beat five guesses.
nn.LazyLinear(out_features) will infer in_features from the first batch it sees, which is convenient for quick experiments. Be aware that the layer has no weights until that first forward pass, so it must run once before you create the optimizer or save the model.
Related errors
- Expected all tensors to be on the same device
- CUDA error: device-side assert triggered — often a shape or index problem that surfaces on the GPU with a useless traceback
- ValueError: Found input variables with inconsistent numbers of samples — the scikit-learn equivalent