Temperature
In one sentence Temperature is the dial that controls how adventurous a language model's word choices are — low is predictable, high is varied and risky.
Updated
Temperature controls how much randomness goes into picking each next token: low values stick to the safest choice, high values spread the bets.
Ordering at a restaurant is the analogy. Temperature 0 is ordering your usual thali every single visit — reliable, identical, never a surprise. Temperature 1 is choosing in proportion to genuine appeal: usually the favourites, sometimes something different. Temperature 2 is letting a bored friend order "anything" — occasionally delightful, often regrettable.
Mechanically, the model's raw scores (logits) are divided by the temperature before softmax converts them to probabilities. Dividing by a small number stretches the gaps between scores, concentrating probability on the top choice; dividing by a large number shrinks the gaps toward a flat distribution.
next-token probabilities for "The capital of France is"
T = 0.2: " Paris" 0.99 " a" 0.005 ...
T = 1.0: " Paris" 0.87 " a" 0.06 " the" 0.03
T = 1.8: " Paris" 0.44 " a" 0.17 " the" 0.12 " home" 0.08Practical settings: 0 to 0.3 for extraction, code and anything needing reproducibility; around 0.7 to 1.0 for conversation and writing; higher only for brainstorming. Two honest caveats. Temperature 0 removes sampling randomness but is not a truth switch — a wrong belief is delivered with the same certainty. And most APIs pair it with top-p, which trims the long tail of bad options before temperature spreads probability onto it.
Where to go next
- Full lesson: Prompt engineering
- Related terms: top-p, top-k, softmax, logits