Error database

Asking to pad but the tokenizer does not have a padding token

GPT-style models were trained without a pad token. Set tokenizer.pad_token = tokenizer.eos_token, and switch to left padding if you are generating text.

The message you saw
Asking to pad but the tokenizer does not have a padding token

By Updated

The error

Output
ValueError: Asking to pad but the tokenizer does not have a padding token. Please select a token to use as `pad_token` `(tokenizer.pad_token = tokenizer.eos_token e.g.)` or add a new pad token via `tokenizer.add_special_tokens({'pad_token': '[PAD]'})`.

It appears the first time you tokenize more than one string at once, or pass padding=True:

python
tokenizer(["short text", "a considerably longer piece of text"], padding=True, return_tensors="pt")

What it means

Padding means filling the shorter sequences in a batch with a filler token so every row has the same length, because a tensor has to be rectangular. To do that, the tokenizer needs to know which token id to use as filler. This tokenizer does not have one defined, so it refuses rather than picking one at random.

It is a configuration gap, not a broken install.

Why it happens

The difference comes from how the two families of models were trained.

BERT-style encoder models process fixed-length batches with padding from day one, so [PAD] is part of their vocabulary and their config names it. Decoder-only models — GPT-2, Llama, Mistral, Qwen, Falcon — were trained on continuous streams of text chopped into equal-length chunks. Nothing was ever padded, so no pad token was needed and none was defined.

You hit the gap the moment you batch. One sequence at a time works fine; two sequences of different lengths need padding, and there is the error.

Padding is safe for the model's output because the attention mask tells it which positions to ignore. The mask is the part that makes padding invisible, and it is why the next section insists on passing it along.

How to fix it

1. Reuse the end-of-sequence token as the pad token. This is the standard fix and is what the error message itself suggests.

python
from transformers import AutoTokenizer

tokenizer = AutoTokenizer.from_pretrained("gpt2")
if tokenizer.pad_token is None:
    tokenizer.pad_token = tokenizer.eos_token

batch = tokenizer(["short text", "a considerably longer piece of text"],
                  padding=True, return_tensors="pt")
print(batch["input_ids"].shape)
print(batch["attention_mask"])
Output
torch.Size([2, 7])
tensor([[1, 1, 0, 0, 0, 0, 0],
        [1, 1, 1, 1, 1, 1, 1]])

The zeros in the mask are the padded positions. Nothing in the model attends to them.

2. If you are generating text in batches, also switch to left padding. This one is easy to miss and it produces bad output rather than an error.

python
tokenizer.pad_token = tokenizer.eos_token
tokenizer.padding_side = "left"        # required for batched generation

A decoder-only model continues from the last position. With padding on the right, the last position of a short sequence is a pad token, so the model is asked to continue from filler and the output degrades. With padding on the left, every real sequence ends at the right-hand edge and generation starts from actual content. For training or classification, right padding — the default — is correct.

3. Tell generate which id is padding, so it can stop cleanly and stop warning you:

python
model.generation_config.pad_token_id = tokenizer.pad_token_id

4. If you would rather add a genuinely new pad token, resize the embeddings. Skipping the resize causes a crash later, because the model has no row in its embedding table for the new id.

python
tokenizer.add_special_tokens({"pad_token": "[PAD]"})
model.resize_token_embeddings(len(tokenizer))     # both lines, always together
model.config.pad_token_id = tokenizer.pad_token_id

Forgetting resize_token_embeddings gives you CUDA error: device-side assert triggered at the first forward pass, with a traceback that points nowhere useful.

5. When fine-tuning, make sure padded positions are excluded from the loss. Setting a label to -100 tells PyTorch's cross-entropy to skip that position, and Hugging Face's collators do this using the pad token id:

python
from transformers import DataCollatorForLanguageModeling

collator = DataCollatorForLanguageModeling(tokenizer=tokenizer, mlm=False)

There is a real trade-off to know about here. If pad_token equals eos_token, this masking cannot distinguish filler from genuine end-of-sequence markers, so real EOS tokens get masked too and the model may never learn to stop. When you are fine-tuning a chat model and stopping behaviour matters, add a distinct [PAD] token as in step 4, or build your label mask from the attention mask rather than from token equality:

python
labels = batch["input_ids"].clone()
labels[batch["attention_mask"] == 0] = -100     # mask padding, keep real EOS intact

For inference only, pad_token = eos_token is entirely fine and this concern does not apply.

6. Add truncation while you are at it. Padding handles sequences that are too short; truncation handles sequences that are too long.

python
batch = tokenizer(texts, padding=True, truncation=True, max_length=512, return_tensors="pt")

How to prevent it

Put the tokenizer setup in one function and use it everywhere, so training, evaluation and serving cannot drift apart:

python
def load_tokenizer(name: str, for_generation: bool = False):
    tok = AutoTokenizer.from_pretrained(name)
    if tok.pad_token is None:
        tok.pad_token = tok.eos_token
    if for_generation:
        tok.padding_side = "left"
    return tok

Two habits go with it. After tokenizing a batch for the first time, print input_ids.shape and the attention_mask once and read them — five seconds of looking prevents a class of silent bugs where padding is present and ignored incorrectly. And pass the whole tokenizer output into the model with **batch rather than picking out input_ids alone, so the mask always travels with the ids.