The HuggingFace Stack

Padding, truncation and attention masks

Batches must be rectangular, so short sentences get padded, long ones get trimmed, and the attention mask tells the model which positions are real.

On this page 5
  1. Why it exists
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Padding fills short sentences with blanks, truncation trims long ones, and the attention mask marks which positions the model should actually read.

Think of an egg tray with twelve moulded slots. Whether you buy five eggs or eleven, the tray shape never changes — empty slots stay empty. And if you somehow had thirteen, one would not fit.

Batches of sentences work the same way. The computer wants every row the same length, like the tray wants its fixed grid. Padding is the empty slots. Truncation is the thirteenth egg that gets left out.

Why it exists

GPUs are fast because they do the same operation on a whole rectangle of numbers at once. Sentences refuse to cooperate — every one is a different length. So we force the rectangle: stretch the short rows with a special blank token, cut the long rows at a limit.

But now there are fake positions in the batch. Without a warning, the model would read the blanks as if they meant something. The attention mask is that warning: a row of 1s and 0s saying "read this, ignore that".

How it works

"Chai is life"            →  [CLS] chai  is   life [SEP] [PAD] [PAD]
mask                          1     1    1    1    1     0     0

"The Mumbai local at..."  →  [CLS] the  mumbai local at  rush [SEP]
mask                          1     1    1     1     1   1    1

Both rows are now the same width. The first earned two blanks; the second was cut at the limit. The mask rides along with the data, and the model multiplies the blanks' influence down to nothing.

A real example you have seen

Government forms with fixed boxes for your name. A short name leaves boxes empty, and you strike through the leftovers so nobody misreads them. A too-long name gets cut off at the last box. Fixed grid, blanks marked, overflow trimmed — all three ideas on one paper form.

Remember this

  • Batches must be rectangular; sentences are not. Padding and truncation force the fit.
  • The attention mask is 1 for real tokens and 0 for padding.
  • Trimming loses text silently — know your length limit and what falls off.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install transformers torch

Tested with transformers 5.6. Only the small tokenizer files of distilbert-base-uncased are downloaded here.

Two sentences forced into one rectangle

padding.py
from transformers import AutoTokenizer

tok = AutoTokenizer.from_pretrained("distilbert-base-uncased")

reviews = ["Chai is life",
           "The Mumbai local at rush hour is an experience nobody ever forgets"]

batch = tok(reviews, padding=True, truncation=True, max_length=8,
            return_tensors="pt")

print("input_ids:")
print(batch["input_ids"])
print("attention_mask:")
print(batch["attention_mask"])
print("row 1 decoded:", tok.decode(batch["input_ids"][1]))
print("pad token:", tok.pad_token, "=", tok.pad_token_id)
Output
input_ids:
tensor([[  101, 15775,  2072,  2003,  2166,   102,     0,     0],
        [  101,  1996,  8955,  2334,  2012,  5481,  3178,   102]])
attention_mask:
tensor([[1, 1, 1, 1, 1, 1, 0, 0],
        [1, 1, 1, 1, 1, 1, 1, 1]])
row 1 decoded: [CLS] the mumbai local at rush hour [SEP]
pad token: [PAD] = 0

The walkthrough

padding=True means "pad to the longest in this batch". Row 0 needed only 6 positions, so it got two 0s — and its mask ends in two 0s to match. This per-batch padding is the efficient default; padding="max_length" pads every batch to the full limit and mostly wastes compute.

truncation=True, max_length=8 cut the second review. Decode row 1: everything after "rush hour" is gone, and [SEP] was re-attached at the cut. Nothing warns you. The sentence's ending — "nobody ever forgets" — was quietly forgotten.

The mask is not optional decoration. When you call model(**batch), the ** unpacking delivers attention_mask alongside input_ids. Inside, masked positions are frozen out of attention so pad tokens influence nothing. Pass input_ids alone and pads are attended — outputs change, subtly and wrongly.

Why row 0 without padding would crash. Try tok(reviews, return_tensors="pt") with no padding: you get an error about converting lists of unequal length to a tensor. The rectangle is not a preference; tensors physically require it.

Common mistakes

Truncating the part you needed. In question answering, the answer often sits late in the passage — exactly what truncation removes first. Check length statistics of your data against max_length before trusting any metric. Long documents need chunking strategies, not bigger limits.

Dropping the mask in hand-rolled code. Code that does model(input_ids=ids) runs fine and scores a little worse forever. Always pass the whole tokenizer output: model(**batch).

Padding decoder models on the right before generating. Causal models continue from the last position. Right-padding puts [PAD]s there, and generation continues from garbage. Generation with batches needs tok.padding_side = "left" — the trap resurfaces in controlling generate.

Assuming every model has a pad token. GPT-2 has none; batching immediately errors with "Asking to pad but the tokenizer does not have a padding token". The standard fix is tok.pad_token = tok.eos_token, plus passing the mask so the fake pads stay ignored.

Try it yourself

Rerun with max_length=16 and padding="max_length". Predict the shape and the two mask rows before printing. Then count wasted positions in the batch — real tokens versus rectangle size — and compute the waste as a percentage.

What to learn next

Researcher — Mathematics and papers.

The mask in the attention equation

Padding is implemented as an additive mask inside scaled dot-product attention: A = softmax(QKᵀ/√d + M), with M_{ij} = 0 for allowed keys and −∞ (in practice, a large negative constant appropriate to the dtype) where key j is padding. Post-softmax, masked keys carry exactly zero weight, so padded positions contribute nothing to any real position's representation. This padding mask composes with the causal mask (upper-triangular −∞) in decoder models; they are independent constraints and both appear in Vaswani et al. (2017), Attention is all you need.

Two subtleties survive the mask. Padded positions still produce hidden states of their own — downstream pooling must exclude them (mean-pooling implementations divide by mask sums, not sequence length). And in fp16, careless use of −inf can produce NaN through softmax of an all-masked row; implementations clamp or use finite negatives.

The economics of the rectangle

For a batch with lengths L_1..L_B padded to L_max, the wasted fraction is 1 − ΣL_i/(B·L_max), and attention cost scales with L_max² — so one outlier sequence taxes the entire batch quadratically. Standard mitigations, in increasing order of sophistication:

  • Dynamic padding: pad per batch, not per dataset — the job of data collators.
  • Length-grouped batching: sort or bucket by length so batchmates match (group_by_length in the Trainer).
  • Packing: concatenate documents into fixed-length streams separated by EOS, eliminating padding entirely; universal in pretraining, with cross-contamination controlled by resetting the causal mask at boundaries (or accepted as noise, as in GPT-2/3 training).
  • Padding-free kernels: FlashAttention's varlen interface consumes concatenated sequences plus cumulative-length indices, removing the rectangle at the kernel level; PyTorch nested tensors aim at the same goal.

Truncation as a data decision

Truncation changes the training distribution, not only inference: BERT's own pretraining used 128-token sequences for 90% of steps for exactly this cost reason (Devlin et al., 2019). For pair tasks, HuggingFace's truncation="only_second" and the stride/overflow machinery (return_overflowing_tokens=True) implement the sliding-window standard from SQuAD-style QA, where each window's offsets let predictions be mapped back — tying this lesson to offsets.

What to learn next