UserWarning: do_sample is set to False. However, temperature is set to... (generation flags)
You passed sampling parameters while greedy decoding is active, so they are ignored. Set do_sample=True to make temperature and top_p count — and set max_new_tokens while you are at it.
Updated
The error
UserWarning: `do_sample` is set to `False`. However, `temperature` is set to `0.7` -- this flag is only used in sample-based generation modes. You should set `do_sample=True` or unset `temperature`.
Recent transformers versions compress it to:
The following generation flags are not valid and may be ignored: ['temperature', 'top_p']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
What it means
generate() has two families of decoding. Greedy decoding (do_sample=False, the default for most models) always picks the highest-scoring token — deterministic, no randomness anywhere. Sampling (do_sample=True) draws tokens from the probability distribution, and only there do temperature and top_p mean anything. You passed sampling knobs to the greedy path. Nothing crashes; your settings are silently doing nothing, which is arguably worse.
Why it happens
API habits transfer badly. Hosted LLM APIs apply temperature without a separate switch, so people write model.generate(**inputs, temperature=0.7) expecting the same. Copied snippets mix flags from both modes. And some models ship generation configs whose defaults interact with your arguments in non-obvious ways.
How to fix it
1. Want varied, creative output? Turn sampling on.
out = model.generate(
**inputs,
do_sample=True,
temperature=0.7,
top_p=0.9,
max_new_tokens=200,
)Now temperature does what you think: lower is safer and more repetitive, higher is riskier and more diverse.
2. Want deterministic output? Drop the sampling knobs.
out = model.generate(**inputs, max_new_tokens=200)Same behaviour, no warning, and the code no longer implies randomness it does not have.
3. Set max_new_tokens explicitly — the silent companion problem. The historical default generation length is tiny (20 new tokens), which is why "my model stops mid-sentence" so often accompanies this warning. Always state how much output you want.
4. Persist your intent in the generation config when the same settings apply everywhere:
model.generation_config.do_sample = True
model.generation_config.temperature = 0.7
model.generation_config.top_p = 0.9Calls to generate then need only the inputs, and the flags cannot half-apply.
How to prevent it
Decide per use-case: sampling for open-ended text, greedy (or beams) for extraction and structured output. Keep the decoding flags together in one visible place, and treat this warning as a real bug report — it means your generation is not configured the way your code claims.
Related errors
- probability tensor contains inf, nan or element < 0 — when sampling itself breaks
- The attention mask and pad token id were not set — the other generate() warning worth obeying
- Token indices sequence length is longer than maximum