Asking to pad but the tokenizer does not have a padding token
GPT-style models were trained without a pad token. Set tokenizer.pad_token = tokenizer.eos_token, and switch to left padding if you are generating text.
Updated
The error
ValueError: Asking to pad but the tokenizer does not have a padding token. Please select a token to use as `pad_token` `(tokenizer.pad_token = tokenizer.eos_token e.g.)` or add a new pad token via `tokenizer.add_special_tokens({'pad_token': '[PAD]'})`.It appears the first time you tokenize more than one string at once, or pass padding=True:
tokenizer(["short text", "a considerably longer piece of text"], padding=True, return_tensors="pt")What it means
Padding means filling the shorter sequences in a batch with a filler token so every row has the same length, because a tensor has to be rectangular. To do that, the tokenizer needs to know which token id to use as filler. This tokenizer does not have one defined, so it refuses rather than picking one at random.
It is a configuration gap, not a broken install.
Why it happens
The difference comes from how the two families of models were trained.
BERT-style encoder models process fixed-length batches with padding from day one, so [PAD] is part of their vocabulary and their config names it. Decoder-only models — GPT-2, Llama, Mistral, Qwen, Falcon — were trained on continuous streams of text chopped into equal-length chunks. Nothing was ever padded, so no pad token was needed and none was defined.
You hit the gap the moment you batch. One sequence at a time works fine; two sequences of different lengths need padding, and there is the error.
Padding is safe for the model's output because the attention mask tells it which positions to ignore. The mask is the part that makes padding invisible, and it is why the next section insists on passing it along.
How to fix it
1. Reuse the end-of-sequence token as the pad token. This is the standard fix and is what the error message itself suggests.
from transformers import AutoTokenizer
tokenizer = AutoTokenizer.from_pretrained("gpt2")
if tokenizer.pad_token is None:
tokenizer.pad_token = tokenizer.eos_token
batch = tokenizer(["short text", "a considerably longer piece of text"],
padding=True, return_tensors="pt")
print(batch["input_ids"].shape)
print(batch["attention_mask"])torch.Size([2, 7])
tensor([[1, 1, 0, 0, 0, 0, 0],
[1, 1, 1, 1, 1, 1, 1]])The zeros in the mask are the padded positions. Nothing in the model attends to them.
2. If you are generating text in batches, also switch to left padding. This one is easy to miss and it produces bad output rather than an error.
tokenizer.pad_token = tokenizer.eos_token
tokenizer.padding_side = "left" # required for batched generationA decoder-only model continues from the last position. With padding on the right, the last position of a short sequence is a pad token, so the model is asked to continue from filler and the output degrades. With padding on the left, every real sequence ends at the right-hand edge and generation starts from actual content. For training or classification, right padding — the default — is correct.
3. Tell generate which id is padding, so it can stop cleanly and stop warning you:
model.generation_config.pad_token_id = tokenizer.pad_token_id4. If you would rather add a genuinely new pad token, resize the embeddings. Skipping the resize causes a crash later, because the model has no row in its embedding table for the new id.
tokenizer.add_special_tokens({"pad_token": "[PAD]"})
model.resize_token_embeddings(len(tokenizer)) # both lines, always together
model.config.pad_token_id = tokenizer.pad_token_idForgetting resize_token_embeddings gives you CUDA error: device-side assert triggered at the first forward pass, with a traceback that points nowhere useful.
5. When fine-tuning, make sure padded positions are excluded from the loss. Setting a label to -100 tells PyTorch's cross-entropy to skip that position, and Hugging Face's collators do this using the pad token id:
from transformers import DataCollatorForLanguageModeling
collator = DataCollatorForLanguageModeling(tokenizer=tokenizer, mlm=False)There is a real trade-off to know about here. If pad_token equals eos_token, this masking cannot distinguish filler from genuine end-of-sequence markers, so real EOS tokens get masked too and the model may never learn to stop. When you are fine-tuning a chat model and stopping behaviour matters, add a distinct [PAD] token as in step 4, or build your label mask from the attention mask rather than from token equality:
labels = batch["input_ids"].clone()
labels[batch["attention_mask"] == 0] = -100 # mask padding, keep real EOS intactFor inference only, pad_token = eos_token is entirely fine and this concern does not apply.
6. Add truncation while you are at it. Padding handles sequences that are too short; truncation handles sequences that are too long.
batch = tokenizer(texts, padding=True, truncation=True, max_length=512, return_tensors="pt")How to prevent it
Put the tokenizer setup in one function and use it everywhere, so training, evaluation and serving cannot drift apart:
def load_tokenizer(name: str, for_generation: bool = False):
tok = AutoTokenizer.from_pretrained(name)
if tok.pad_token is None:
tok.pad_token = tok.eos_token
if for_generation:
tok.padding_side = "left"
return tokTwo habits go with it. After tokenizing a batch for the first time, print input_ids.shape and the attention_mask once and read them — five seconds of looking prevents a class of silent bugs where padding is present and ignored incorrectly. And pass the whole tokenizer output into the model with **batch rather than picking out input_ids alone, so the mask always travels with the ids.
Related errors
- The attention mask and the pad token id were not set — the warning that appears next when the mask is dropped
- CUDA error: device-side assert triggered — what happens when you add tokens without resizing the embeddings
- Context length exceeded