RuntimeError: Expected 4-dimensional input for 4-dimensional weight (conv2d)
Conv2d wants batches of channels-first images — shape [N, C, H, W]. A single image is missing its batch dimension; add it with unsqueeze(0), and permute HWC images to CHW.
Updated
The error
RuntimeError: Expected 4-dimensional input for 4-dimensional weight [64, 3, 7, 7], but got 3-dimensional input of size [3, 224, 224] instead
Newer PyTorch versions phrase it as:
RuntimeError: Expected 3D (unbatched) or 4D (batched) input to conv2d, but got input of size: [224, 224]
What it means
A CNN convolution layer processes images in a fixed layout: [N, C, H, W] — batch size, channels, height, width. The weight shape in the message decodes the layer's expectation: [64, 3, 7, 7] means 64 filters, 3 input channels, 7×7 kernels. Your input [3, 224, 224] is one image without a batch dimension. The second variant, [224, 224], is missing both batch and channel.
Why it happens
Training code gets its batch dimension free from the DataLoader. Inference code passes one image directly, and the batch dimension vanishes. Grayscale images lose the channel dimension the same way. And images loaded with OpenCV or PIL arrive as [H, W, C] — channels last — which has the right number of dimensions in the wrong order, producing either this error or a channel-count complaint from the same family.
How to fix it
1. Add the batch dimension for single-image inference.
x = img_tensor.unsqueeze(0) # [3,224,224] -> [1,3,224,224]
out = model(x)
result = out.squeeze(0) # back to per-image if you want2. Give grayscale images their channel dimension.
x = gray.unsqueeze(0).unsqueeze(0) # [H,W] -> [1,1,H,W]And make sure the first conv layer declares in_channels=1 to match.
3. Convert HWC images from OpenCV/PIL to CHW.
x = torch.from_numpy(img).permute(2, 0, 1).float() / 255.0 # [H,W,C] -> [C,H,W]
x = x.unsqueeze(0)Or let torchvision do the whole job — transforms.ToTensor() converts HWC uint8 to CHW float in [0,1] in one step.
4. Reuse the exact same transform pipeline at inference that training used.
x = train_transforms(pil_img).unsqueeze(0)Most single-image inference bugs — shape, scaling, normalisation — disappear when the transforms are shared instead of re-implemented.
How to prevent it
Write one preprocess(image) -> tensor[1,C,H,W] function and route every inference call through it. Print x.shape at the model boundary when developing; the expected [N, C, H, W] is quick to eyeball. Remember the two converters: unsqueeze adds dimensions, permute reorders them — most image-shape bugs need one of each.
Related errors
- mat1 and mat2 shapes cannot be multiplied — the equivalent error at linear layers
- stack expects each tensor to be equal size
- ValueError: cannot reshape array of size
- cv2.error: (-215:Assertion failed) — when the image never loaded at all