Error database

mat1 and mat2 shapes cannot be multiplied (shape mismatch)

A layer expected a different number of input features than it received. Print the shape going in, and set in_features to match.

The message you saw
mat1 and mat2 shapes cannot be multiplied (shape mismatch)

By Updated

The error

Output
RuntimeError: mat1 and mat2 shapes cannot be multiplied (64x784 and 512x10)

Older PyTorch versions phrase it differently:

Output
RuntimeError: size mismatch, m1: [64 x 784], m2: [512 x 10]

And the convolution version of the same class of problem:

Output
RuntimeError: Given groups=1, weight of size [64, 3, 3, 3], expected input[32, 1, 28, 28] to have 3 channels, but got 1 channels instead

What it means

Matrix multiplication has one rule: the width of the first matrix must equal the height of the second. Your data arrived as 64 rows of 784 numbers, and the layer's weight matrix expects rows of 512 numbers. 784 and 512 do not match, so the operation is not defined.

Read the numbers this way:

(64 x 784)  @  (512 x 10)
     ▲            ▲
     └────────────┘
   these must be equal — they are not

The 64 is your batch size and the 10 is the output size; neither is the problem. The mismatch is always the two inner numbers.

One detail that saves confusion: for nn.Linear(in_features, out_features), the second matrix prints as (in_features x out_features). So (512 x 10) means the layer was built as nn.Linear(512, 10). Your data has 784 features and the layer wants 512.

Why it happens

Almost always, a shape changed somewhere earlier and the layer definition was not updated to follow.

The classic case is a missing flatten. A convolution stack outputs a 4-dimensional tensor of shape (batch, channels, height, width), and a linear layer needs a 2-dimensional (batch, features). Without a flatten step, or with the wrong flattened size hardcoded, the numbers cannot line up.

Equally common: the flattened size was correct for one input resolution and you changed the resolution. A 28×28 image through two convolutions and two poolings gives a different feature count than a 32×32 one, so nn.Linear(1600, 10) becomes wrong the moment you switch datasets.

Then there are the smaller ones. Two tensors passed to torch.matmul in the wrong order. An RNN or transformer where batch_first is not what you assumed, so the batch and sequence dimensions are swapped. An embedding dimension that no longer matches the layer after you changed models. A pretrained backbone whose final feature size is not what you remembered — ResNet-18 ends at 512, ResNet-50 at 2048, and swapping one for the other produces exactly this error.

How to fix it

1. Print the shape immediately before the failing layer. Do this before theorising. It takes fifteen seconds and it usually ends the investigation.

python
def forward(self, x):
    x = self.features(x)
    print("after features:", x.shape)     # e.g. torch.Size([64, 64, 7, 7])
    x = x.flatten(1)
    print("after flatten :", x.shape)     # e.g. torch.Size([64, 3136])
    return self.classifier(x)

The second number after flattening is what in_features must be.

2. Compute the right in_features instead of guessing. Run a dummy tensor of your real input size through the feature part once, at construction time:

python
import torch
import torch.nn as nn

class Net(nn.Module):
    def __init__(self, input_shape=(3, 32, 32), num_classes=10):
        super().__init__()
        self.features = nn.Sequential(
            nn.Conv2d(input_shape[0], 32, 3, padding=1), nn.ReLU(), nn.MaxPool2d(2),
            nn.Conv2d(32, 64, 3, padding=1), nn.ReLU(), nn.MaxPool2d(2),
        )
        # ask the network how big its own output is, rather than working it out on paper
        with torch.no_grad():
            n_features = self.features(torch.zeros(1, *input_shape)).flatten(1).shape[1]
        self.classifier = nn.Linear(n_features, num_classes)

    def forward(self, x):
        return self.classifier(self.features(x).flatten(1))

print(Net()(torch.randn(8, 3, 32, 32)).shape)
Output
torch.Size([8, 10])

This model now survives a change of input resolution without any edit.

3. Flatten correctly. x.flatten(1) keeps dimension 0 (the batch) and merges everything after it. Inside an nn.Sequential, use nn.Flatten().

python
x = x.flatten(1)          # (64, 64, 7, 7) → (64, 3136)
x = x.view(x.size(0), -1) # same result; view needs contiguous memory, flatten does not

Avoid x.view(-1, 3136) with a hardcoded number. When the size is wrong it silently changes your batch size instead of failing, which produces a much harder bug later.

4. Make the model resolution-independent with adaptive pooling. This is what torchvision's own models do, and it removes the problem at the source.

python
self.pool = nn.AdaptiveAvgPool2d((1, 1))    # any H x W in, always 1 x 1 out
self.classifier = nn.Linear(64, num_classes)   # 64 = channel count, fixed

5. If the input really is transposed, transpose it — after confirming.

python
print(a.shape, b.shape)            # (64, 512) (10, 512)
out = a @ b.T                      # (64, 512) @ (512, 10) → (64, 10)

Only reach for this once the printed shapes show a genuine transpose. Adding a .T to make an error disappear, when the real problem is a wrong layer size, produces a model that trains on nonsense.

6. Check batch_first for recurrent and transformer layers. PyTorch defaults to (seq, batch, feature), while most people prepare data as (batch, seq, feature).

python
self.lstm = nn.LSTM(input_size=64, hidden_size=128, batch_first=True)

7. For the convolution channel version of this error, the number after expected input[...] is what your data has and the second number in weight of size [...] is what the layer wants. Grayscale images have 1 channel, colour has 3, so nn.Conv2d(1, 32, 3) for MNIST and nn.Conv2d(3, 32, 3) for CIFAR-10 or photographs.

How to prevent it

Write a shape test and run it before training. It costs one second and catches this at the moment you edit the model rather than twenty minutes into an epoch:

python
def test_shapes():
    model = Net(input_shape=(3, 32, 32), num_classes=10)
    out = model(torch.randn(2, 3, 32, 32))
    assert out.shape == (2, 10), out.shape
    print("shapes fine")

test_shapes()

Two more habits help. Pass the input shape and class count into the model as arguments rather than burying constants in the layer definitions. And when you build a new architecture, add a temporary print(x.shape) after every stage, run one batch, then delete them — five prints beat five guesses.

nn.LazyLinear(out_features) will infer in_features from the first batch it sees, which is convenient for quick experiments. Be aware that the layer has no weights until that first forward pass, so it must run once before you create the optimizer or save the model.