Error database

The attention mask and the pad token id were not set (HuggingFace)

model.generate() received no attention mask, so it cannot tell padding from real text. Pass the whole tokenizer output with **inputs.

The message you saw
The attention mask and the pad token id were not set (HuggingFace)

By Updated

The error

Output
The attention mask and the pad token id were not set. As a consequence, you may observe unexpected behavior. Please pass your input's `attention_mask` to obtain reliable results.
Setting `pad_token_id` to `eos_token_id`:50256 for open-end generation.

Newer versions of Transformers word it more precisely:

Output
The attention mask is not set and cannot be inferred from input because pad token is same as eos token. As a consequence, you may observe unexpected behavior. Please pass your input's `attention_mask` to obtain reliable results.
Output
A decoder-only architecture is being used, but right-padding was detected! For correct generation results, please set `padding_side='left'` when initializing the tokenizer.

What it means

This is a warning, not a crash. Your code will produce output. The warning is telling you that the output may be wrong.

An attention mask is a row of 1s and 0s, one per token, marking which positions hold real text and which hold padding. Without it, the model treats padding as content — it attends to filler tokens and counts them when working out positions.

The Setting pad_token_id to eos_token_id line is the library filling a gap for you so generation can stop somewhere. It is informational, and setting the value yourself makes it go away.

Why it happens

Calling generate with only the ids is the usual route in:

python
input_ids = tokenizer(prompt, return_tensors="pt").input_ids   # the mask is discarded here
output = model.generate(input_ids, max_new_tokens=50)

The tokenizer returned both input_ids and attention_mask; taking .input_ids threw half of it away.

Transformers tries to reconstruct the mask by treating every occurrence of the pad token id as padding. That guess works when the pad token is distinct. It fails when pad_token has been set equal to eos_token, which is the standard fix for the missing pad token error. After that change the library cannot tell a padding token from a real end-of-sequence token in your prompt. That is precisely the case the second message describes.

The consequences are real for batched generation and mild for a single unpadded prompt. In a batch, padded positions influence attention and shift the position indices of your real tokens, so short prompts in a batch produce noticeably worse output than the same prompts run one at a time. That symptom — "it works alone but not in a batch" — is this warning being ignored.

How to fix it

1. Pass the entire tokenizer output. One character of difference, and it fixes the warning properly.

python
from transformers import AutoTokenizer, AutoModelForCausalLM

tokenizer = AutoTokenizer.from_pretrained("gpt2")
model = AutoModelForCausalLM.from_pretrained("gpt2")

inputs = tokenizer("The capital of France is", return_tensors="pt")
outputs = model.generate(**inputs, max_new_tokens=20)          # **inputs, not inputs.input_ids
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
Output
The capital of France is Paris, and the city is home to the largest number of French

**inputs unpacks input_ids and attention_mask into the call. That is the whole fix for the common case.

2. Set the pad token id explicitly to silence the second line.

python
model.generation_config.pad_token_id = tokenizer.eos_token_id

Setting it on the generation config applies to every call, which is tidier than passing pad_token_id= each time.

3. For batched generation, configure padding properly. All three settings belong together.

python
tokenizer = AutoTokenizer.from_pretrained("gpt2")
tokenizer.pad_token = tokenizer.eos_token
tokenizer.padding_side = "left"                 # required for decoder-only generation

prompts = ["The capital of France is", "In one sentence, photosynthesis is"]
inputs = tokenizer(prompts, return_tensors="pt", padding=True).to(model.device)

outputs = model.generate(**inputs, max_new_tokens=20,
                         pad_token_id=tokenizer.eos_token_id)

for row in outputs:
    print(tokenizer.decode(row, skip_special_tokens=True))

Left padding matters because a decoder-only model continues from the final position. With right padding, the shorter prompt ends in filler and the model continues from filler. This is the difference between a coherent completion and drifting text, and it is silent — no error, only worse results.

4. If you are moving tensors to a device by hand, move the mask too.

python
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)   # moves every field
outputs = model.generate(**inputs, max_new_tokens=50)

Moving input_ids alone leaves the mask on the CPU and produces Expected all tensors to be on the same device.

5. Confirm the mask is doing what you expect. Reading it once removes all doubt:

python
batch = tokenizer(["hi", "a much longer prompt here"], padding=True, return_tensors="pt")
print(batch["attention_mask"])
Output
tensor([[0, 0, 0, 1, 1],
        [1, 1, 1, 1, 1]])

With padding_side="left", the zeros appear at the start. That is correct.

6. Do not silence it with the logging level. Turning the warning off with transformers.logging.set_verbosity_error() hides the message and keeps the wrong behaviour. Fix the call instead.

How to prevent it

Make **inputs the only way you call generate in your codebase. There is no situation where extracting input_ids alone is better, and the habit removes this category of bug entirely.

Set up the tokenizer once, in one place, with pad token and padding side decided by whether you are generating or training. Then confirm your batching is honest with a one-time check: run two prompts individually, run the same two as a batch, and compare the outputs. They should be very close. If the short one is markedly worse in the batch, padding is being handled incorrectly and this warning was the clue.