Error database

UserWarning: do_sample is set to False. However, temperature is set to... (generation flags)

You passed sampling parameters while greedy decoding is active, so they are ignored. Set do_sample=True to make temperature and top_p count — and set max_new_tokens while you are at it.

The message you saw
UserWarning: do_sample is set to False. However, temperature is set to... (generation flags)

By Updated

The error

Output
UserWarning: `do_sample` is set to `False`. However, `temperature` is set to `0.7` -- this flag is only used in sample-based generation modes. You should set `do_sample=True` or unset `temperature`.

Recent transformers versions compress it to:

Output
The following generation flags are not valid and may be ignored: ['temperature', 'top_p']. Set `TRANSFORMERS_VERBOSITY=info` for more details.

What it means

generate() has two families of decoding. Greedy decoding (do_sample=False, the default for most models) always picks the highest-scoring token — deterministic, no randomness anywhere. Sampling (do_sample=True) draws tokens from the probability distribution, and only there do temperature and top_p mean anything. You passed sampling knobs to the greedy path. Nothing crashes; your settings are silently doing nothing, which is arguably worse.

Why it happens

API habits transfer badly. Hosted LLM APIs apply temperature without a separate switch, so people write model.generate(**inputs, temperature=0.7) expecting the same. Copied snippets mix flags from both modes. And some models ship generation configs whose defaults interact with your arguments in non-obvious ways.

How to fix it

1. Want varied, creative output? Turn sampling on.

python
out = model.generate(
    **inputs,
    do_sample=True,
    temperature=0.7,
    top_p=0.9,
    max_new_tokens=200,
)

Now temperature does what you think: lower is safer and more repetitive, higher is riskier and more diverse.

2. Want deterministic output? Drop the sampling knobs.

python
out = model.generate(**inputs, max_new_tokens=200)

Same behaviour, no warning, and the code no longer implies randomness it does not have.

3. Set max_new_tokens explicitly — the silent companion problem. The historical default generation length is tiny (20 new tokens), which is why "my model stops mid-sentence" so often accompanies this warning. Always state how much output you want.

4. Persist your intent in the generation config when the same settings apply everywhere:

python
model.generation_config.do_sample = True
model.generation_config.temperature = 0.7
model.generation_config.top_p = 0.9

Calls to generate then need only the inputs, and the flags cannot half-apply.

How to prevent it

Decide per use-case: sampling for open-ended text, greedy (or beams) for extraction and structured output. Keep the decoding flags together in one visible place, and treat this warning as a real bug report — it means your generation is not configured the way your code claims.