RuntimeError: stack expects each tensor to be equal size (DataLoader)
The DataLoader tried to stack samples of different sizes into one batch. Resize images in transforms, or pad variable-length sequences with a collate function.
Updated
The error
RuntimeError: stack expects each tensor to be equal size, but got [3, 224, 224] at entry 0 and [3, 375, 500] at entry 1
The Hugging Face equivalent, from a tokenized dataset without padding:
ValueError: expected sequence of length 12 at dim 1 (got 19)
What it means
A batch is one tensor holding several samples, and a tensor is rectangular — every sample must have the same shape. The DataLoader's default collate function builds batches with torch.stack, which refuses mixed sizes. Entry 0 is a 224×224 image, entry 1 is 375×500; no rectangle fits both.
The message shows you the two disagreeing shapes. That pair tells you whether the fix is resizing (images) or padding (sequences).
Why it happens
Real datasets are not uniform. Photos come in every resolution. Sentences have different token counts. The error appears the first time a batch happens to mix two sizes — which can be several batches in, making it look intermittent. batch_size=1 makes it vanish, which confirms the diagnosis but is not a fix.
How to fix it
1. Images: make the transform pipeline produce one fixed size.
from torchvision import transforms
tfm = transforms.Compose([
transforms.Resize(256),
transforms.CenterCrop(224), # every image exits as 3x224x224
transforms.ToTensor(),
])Check the order — an augmentation like RandomCrop placed before a conditional resize can still emit odd sizes. The last spatial operation should fix the size unconditionally.
2. Sequences: pad within each batch using a collate function.
from torch.nn.utils.rnn import pad_sequence
def collate(batch):
seqs, labels = zip(*batch)
seqs = pad_sequence(seqs, batch_first=True, padding_value=0)
return seqs, torch.tensor(labels)
loader = DataLoader(ds, batch_size=32, collate_fn=collate)Padding per batch wastes less memory than padding the whole dataset to its global maximum.
3. Hugging Face: let the tokenizer or collator pad.
enc = tokenizer(texts, padding=True, truncation=True, return_tensors="pt")Or for Trainer, pass DataCollatorWithPadding(tokenizer) — the same per-batch padding, handled for you. This pairs with an attention mask so the model ignores the padding.
4. Mixed grayscale/colour image sets: normalise channels too. [1, H, W] versus [3, H, W] triggers the same error. Add transforms.Grayscale(num_output_channels=3) or convert with img.convert("RGB") at load time.
How to prevent it
After writing any Dataset, check ten random samples: print({ds[i][0].shape for i in random.sample(range(len(ds)), 10)}). One distinct shape means safe batching. For sequence data, decide the padding strategy when you design the Dataset, not when the error arrives.
Related errors
- Asking to pad but the tokenizer does not have a padding token — padding blocked one level earlier
- ValueError: setting an array element with a sequence — the NumPy face of ragged data
- Expected 4-dimensional input
- DataLoader worker is killed by signal