Generative AI

Temperature and sampling

Temperature is the dial that decides how adventurous a model is when choosing each next word, and turning it down reduces variety without making the model any more truthful.

On this page 11
  1. The short answer
  2. The analogy you have already lived
  3. Why there is a dial at all
  4. What the dial actually does
  5. What this looks like in real text
  6. The thing people get most wrong
  7. A gentler way to control it
  8. Where you have already seen this
  9. What is honestly messy
  10. Remember this
  11. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

Temperature is a dial that decides how adventurous the model is when it picks each next word.

The analogy you have already lived

Think about the restaurant you go to most often. You know the menu by heart.

On a boring day you order the same dish you always order. Every visit, the same thing. That is temperature zero.

On a normal day you usually order your favourite. Now and then you try the second or third thing you like. That is temperature around 0.7.

On a very strange day you close your eyes and run your finger down the page. You order whatever it lands on, including things you have never once wanted. That is temperature 2.0.

Same menu. Same you. A different rule for choosing.

Why there is a dial at all

The what is an LLM lesson said something easy to skim past: the model does not produce an answer. It produces a score for every word it knows, thousands of them, for every single position.

Something else has to take those scores and pick one word. That picking step is called sampling, and temperature is its main setting.

This surprises people. The model is not being creative or careful. The scores are the same every time. The dial changes only how those scores are turned into a choice.

What the dial actually does

Suppose the model has read "The best thing about Bengaluru is the" and produced these scores.

Here is roughly how often each word gets picked, out of a hundred tries, at three settings.

wordtemperature 0temperature 0.7temperature 2.0
weather1006532
food02121
traffic0815
coffee0412
people029
rain006
startups005
biryani001

Read the first column. At temperature zero the top word wins every time, and nothing else is ever chosen. The answer is fixed.

Read the last column. At temperature 2.0 the top word has been dragged down. It falls from near-certain to about one time in three. Words the model rated poorly become genuinely possible. "biryani" was almost impossible before. Now it happens.

The dial does not add new options. It flattens or sharpens the ones already there.

What this looks like in real text

Here is a very small model, generating three times, at three settings. This is real output, produced by the program in the developer section.

   temperature 0.01   the chai was hot and sweet . the chai was hot and
   temperature 0.7    the chai was sweet . the coffee was bitter . the chai
   temperature 2.0    the coffee chai . . the the coffee was hot was sweet

Look at the first line. It is grammatical and it has begun repeating itself. That is the classic failure of a very low temperature. The model finds the single most likely path and walks it in a circle.

The second line varies and stays sensible. That is why most defaults sit near this value.

The third line has fallen apart. Words that should never follow each other are being chosen, because their scores were flattened until they became plausible.

The thing people get most wrong

Turning the temperature down does not make the model more truthful.

This is worth stating twice, because the belief is everywhere. Low temperature makes the model pick its highest-scoring word. If its highest-scoring word is wrong, low temperature makes it wrong consistently, with total confidence, every time you ask.

A hallucination at temperature zero is still a hallucination. You have removed the variety, not the error. If anything, the confident repeatability makes it harder to notice.

What low temperature genuinely buys you is repeatability, which matters enormously when you are testing, comparing prompts or extracting structured data.

A gentler way to control it

There is a second approach, and it is usually better than turning the temperature up.

Instead of flattening every score, throw away the poor options first, then choose normally from what remains.

  • Top-k keeps only the k highest-scoring words. Set k to 3 and the model can never pick the fourth-best word, no matter what.
  • Top-p, also called nucleus sampling, keeps the smallest group of top words whose chances add up to p. Set p to 0.9 and it keeps enough words to cover ninety percent of the model's confidence, then stops.

The difference between the two matters. Top-k always keeps the same number of options. Top-p keeps more when the model is unsure and fewer when it is confident, which is usually what you want.

Most modern tools use top-p, often alongside a temperature.

Where you have already seen this

  • Asking a chatbot the same question twice and getting different wording. That is sampling above zero.
  • A coding assistant that gives you the same completion again and again. Code tools run at very low temperature on purpose, because there is usually one right answer.
  • A "regenerate" button. It re-runs the sampler, which is only useful because the sampler has randomness in it.
  • An AI writing tool with a "creative" slider. That slider is a temperature dial with a friendlier name.

What is honestly messy

Temperature zero is almost deterministic, not fully. Two words can score exactly the same, and something has to break the tie. On top of that, the arithmetic runs on hardware that is not perfectly consistent. Results shift a little with how many other requests are running alongside yours.

So the same prompt at temperature zero, on the same model, can occasionally differ. This catches out people building tests. It is a real effect and nobody is hiding it from you.

Remember this

  • Temperature reshapes the model's own scores. It cannot invent an option that was not already there.
  • Low temperature gives repeatability, not correctness.
  • Top-p trims the poor options instead of flattening everything, and is usually the better lever.

What to learn next

Developer — Code and libraries.

Three programs, all pure NumPy, all running in milliseconds. Together they show exactly what every temperature, top-k and top-p setting you have ever typed is doing.

Setup

bash
pip install numpy

What temperature does to the numbers

temperature.py
import numpy as np

# One real decision point. The model has read "The best thing about Bengaluru is the"
# and produced one raw score (a "logit") for every word it knows.
tokens = ["weather", "food", "traffic", "coffee", "people", "rain", "startups", "biryani"]
logits = np.array([4.0, 3.2, 2.5, 2.0, 1.5, 0.5, 0.2, -3.0])

def softmax(x, temperature):
    z = x / temperature                 # temperature divides BEFORE the softmax
    z = z - z.max()                     # subtract the max so exp() cannot overflow
    e = np.exp(z)
    return e / e.sum()

print(f"{'token':>9}  {'T=0.01':>8} {'T=0.7':>8} {'T=1.0':>8} {'T=2.0':>8}")
for i, tok in enumerate(tokens):
    row = [softmax(logits, T)[i] for T in (0.01, 0.7, 1.0, 2.0)]
    print(f"{tok:>9}  " + " ".join(f"{p:8.4f}" for p in row))
Output
    token    T=0.01    T=0.7    T=1.0    T=2.0
  weather    1.0000   0.6523   0.5146   0.3174
     food    0.0000   0.2080   0.2312   0.2128
  traffic    0.0000   0.0765   0.1148   0.1499
   coffee    0.0000   0.0375   0.0696   0.1168
   people    0.0000   0.0183   0.0422   0.0909
     rain    0.0000   0.0044   0.0155   0.0552
 startups    0.0000   0.0029   0.0115   0.0475
  biryani    0.0000   0.0000   0.0005   0.0096

Four things in that table are worth pointing at.

The logits never change. Only the division changes. The model did identical work in all four columns.

T=0.01 is a one-hot vector. Every probability is 0 except the top one. This is why temperature 0 is implemented as argmax rather than by actually dividing by zero, which would raise a ZeroDivisionError or produce nan.

Look at biryani. It goes from 0.0005 at T=1.0 to 0.0096 at T=2.0 — roughly twenty times more likely. High temperature does its most dramatic work on the tail, which is precisely where the bad words live.

Look at food. It barely moves: 0.2312 to 0.2128. Words near the top of the distribution are relatively unaffected. Temperature is not a uniform "make everything more random" knob. It redistributes mass from the head to the tail.

The same thing, in generated text

generate.py
import numpy as np
from collections import defaultdict

corpus = """the chai was hot and sweet . the chai was strong and sweet .
the coffee was hot and bitter . the coffee was strong and sweet .
the chai was hot and strong . the coffee was sweet and hot .
the chai was sweet . the coffee was bitter . the chai was hot .""".split()

vocab = sorted(set(corpus))
nxt = defaultdict(lambda: defaultdict(int))
for a, b in zip(corpus, corpus[1:]):
    nxt[a][b] += 1                       # count which word follows which

def logits_after(word):
    counts = np.array([nxt[word][w] for w in vocab], dtype=float)
    return np.log(counts + 0.2)          # 0.2 keeps unseen words rare but reachable

def generate(T, seed, n=11):
    rng = np.random.default_rng(seed)
    word, out = "the", ["the"]
    for _ in range(n):
        z = logits_after(word) / T
        z -= z.max()
        p = np.exp(z); p /= p.sum()
        word = vocab[rng.choice(len(vocab), p=p)]
        out.append(word)
    return " ".join(out)

for T in (0.01, 0.7, 2.0):
    for seed in (0, 1, 2):
        print(f"T={T:<5} seed={seed}  {generate(T, seed)}")
Output
T=0.01  seed=0  the chai was hot and sweet . the chai was hot and
T=0.01  seed=1  the chai was hot and sweet . the chai was hot and
T=0.01  seed=2  the chai was hot and sweet . the chai was hot and
T=0.7   seed=0  the coffee was bitter . the coffee was strong and sweet and
T=0.7   seed=1  the chai was hot hot . the coffee was hot . the
T=0.7   seed=2  the chai was sweet . the coffee was bitter . the chai
T=2.0   seed=0  the coffee chai . . the the coffee was hot was sweet
T=2.0   seed=1  the coffee was bitter was coffee hot sweet and strong . the
T=2.0   seed=2  the chai coffee was and strong strong and . coffee the coffee

This is a bigram model, not a transformer, and that is deliberate — everything here is checkable by hand from the corpus above. The sampler is identical to the one in a real system.

The three T=0.01 rows are byte-for-byte identical. Different random seeds, same output. That is what determinism looks like.

They also loop. "the chai was hot and sweet . the chai was hot and" — it has entered a cycle and would stay there forever. Greedy decoding does this in real models too, and it is the reason nobody generates long-form text at temperature 0.

The T=0.7 rows vary and stay readable. Row 1 has a stumble — "hot hot" — which is honest. Small models produce these.

The T=2.0 rows are broken in a specific way. They are not random noise; they still use words from the corpus. But "the the coffee was hot was sweet" shows the grammar going. Flattening the distribution let words with almost no support win.

Top-k and top-p, from scratch

truncation.py
import numpy as np

tokens = ["weather", "food", "traffic", "coffee", "people", "rain", "startups", "biryani"]
logits = np.array([4.0, 3.2, 2.5, 2.0, 1.5, 0.5, 0.2, -3.0])

def softmax(x, T=1.0):
    z = (x / T) - (x / T).max()
    e = np.exp(z)
    return e / e.sum()

def top_k(p, k):
    keep = np.argsort(p)[-k:]                       # indices of the k largest
    q = np.zeros_like(p); q[keep] = p[keep]
    return q / q.sum()                              # renormalise what is left

def top_p(p, threshold):
    order = np.argsort(p)[::-1]                     # largest first
    running = np.cumsum(p[order])
    cut = np.searchsorted(running, threshold) + 1   # smallest set reaching the threshold
    keep = order[:cut]
    q = np.zeros_like(p); q[keep] = p[keep]
    return q / q.sum()

base = softmax(logits)
print(f"{'token':>9} {'plain':>8} {'top-k=3':>8} {'top-p=0.9':>10}")
for i, t in enumerate(tokens):
    print(f"{t:>9} {base[i]:8.4f} {top_k(base,3)[i]:8.4f} {top_p(base,0.9)[i]:10.4f}")
print()
print("words still possible: plain", int((base > 0).sum()),
      "| top-k=3", int((top_k(base,3) > 0).sum()),
      "| top-p=0.9", int((top_p(base,0.9) > 0).sum()))
Output
    token    plain  top-k=3  top-p=0.9
  weather   0.5146   0.5979     0.5532
     food   0.2312   0.2687     0.2486
  traffic   0.1148   0.1334     0.1234
   coffee   0.0696   0.0000     0.0749
   people   0.0422   0.0000     0.0000
     rain   0.0155   0.0000     0.0000
 startups   0.0115   0.0000     0.0000
  biryani   0.0005   0.0000     0.0000

words still possible: plain 8 | top-k=3 3 | top-p=0.9 4

The surviving words get larger probabilities than they had before, because the discarded mass is redistributed among them. weather rises from 0.5146 to 0.5979 under top-k.

Top-k kept 3 words. Top-p kept 4, because it needed coffee to reach 0.9 of the total. That one-word difference is the entire argument for top-p: it adapts to how confident the model is at this particular position, and top-k does not.

To see the failure mode of top-k, imagine a position where the model is genuinely certain — completing "New York" after "New". Top-k with k=50 keeps 49 words that should have been impossible. Top-p keeps one.

The order of operations, which is not obvious

Real libraries apply these in a fixed order, and changing it changes the output:

   raw logits
     -> repetition and frequency penalties  (subtract from logits of seen tokens)
     -> divide by temperature
     -> top-k filter
     -> top-p filter
     -> softmax and sample

Temperature comes before truncation. That means raising the temperature can push a word into the top-p set that would otherwise have been cut. The two settings are not independent, which is why tuning both at once is confusing and why most guidance says to change one and leave the other at its default.

Common mistakes

Using temperature to reduce hallucination. It does not. It makes the model's favourite answer come out every time. If that answer is invented, you now get the same invented answer reliably. Ground the model with RAG instead.

Setting temperature 0 and top-p 0.9 together. At temperature 0 the sampler is argmax, so top-p has nothing to do. Harmless, and a sign that the settings were copied rather than chosen.

Expecting temperature 0 to be reproducible across runs. It usually is, and it is not guaranteed. Exact ties must be broken somehow. More importantly, floating-point reduction order on a GPU depends on batch composition, so the same prompt can produce marginally different logits when the server is busy. If you need bit-exact reproducibility, you need a fixed batch size and a fixed kernel configuration, not temperature 0 on its own.

Turning temperature above about 1.2 for anything that must be parsed. JSON extraction, classification labels, function arguments. Structure is a low-entropy thing, and flattening the distribution is exactly the wrong move. Use temperature 0 and structured output constraints.

Assuming higher temperature means more creative. Past a point it means less coherent, which is different. Genuine variety comes from a better prompt, more examples, or several samples at a moderate temperature — not from turning one dial to its maximum.

Sensible starting points

TaskTemperaturetop-pWhy
Extraction, classification, JSON0not usedone right answer, needs to be repeatable
Code generation0 to 0.20.95mostly one right answer
Factual question answering0 to 0.30.9reduce wandering, does not add truth
Summarising0.3 to 0.50.9slight variety, stays faithful
Chat and general writing0.70.9 to 0.95the common default
Brainstorming, fiction0.9 to 1.10.95variety is the point

Treat these as places to start measuring from, not as settled answers. The right value depends on the model and on your prompt.

Try it yourself

In temperature.py, change biryani's logit from -3.0 to 1.0 and re-run. Watch how much of the change lands on the low-temperature columns versus the high-temperature ones. Predict the direction first.

Then in generate.py, add top-p filtering to the sampler and run it at T=2.0. You should see something interesting: high temperature plus tight top-p produces text that is varied but stays inside the corpus vocabulary, because the tail was cut before the flattening could promote it. That combination — warm temperature with a nucleus filter — is what most production settings actually are.

What to learn next

Researcher — Mathematics and papers.

The tempered distribution

Given logits z over a vocabulary of size V, temperature T > 0 defines:

text
p_i(T) = exp(z_i / T) / sum over j = 1..V of exp(z_j / T)
  • z_i is the raw score for token i, the pre-softmax output of the final linear layer.
  • T is the temperature.

Equivalently, p(T) is the base distribution p(1) raised elementwise to the power 1/T and renormalised. This is the Gibbs form, and two limits follow:

text
T -> 0+    p(T) -> the one-hot vector at argmax(z)      (greedy decoding)
T -> inf   p(T) -> the uniform distribution over V

The Shannon entropy H(p(T)) = - sum_i p_i log p_i is monotonically non-decreasing in T. Temperature is therefore an entropy dial in a precise sense, which is why it is a reasonable single control despite its crudeness.

Note that T = 0 is not evaluable in the formula — division by zero. Implementations special-case it to argmax. A library that instead uses a small epsilon can produce inf or nan in float16 when logits are large.

Truncation samplers

Top-k (Fan, Lewis & Dauphin, 2018, arXiv:1805.04833). Retain the k highest-probability tokens, renormalise, sample. Fixed cardinality regardless of the shape of the distribution, which is its defect: k is simultaneously too large at confident positions and too small at genuinely ambiguous ones.

Top-p / nucleus (Holtzman et al., 2020, arXiv:1904.09751). Retain the smallest set S such that:

text
sum over i in S of p_i  >=  p_threshold

with S built by descending probability. Cardinality adapts to the entropy at each position.

The paper's contribution is more than the algorithm. It diagnosed neural text degeneration: maximisation-based decoding, including beam search, produces text whose per-token likelihood is far higher than that of human text, and which is repetitive and dull. Human text sits in a band of moderate surprise. Optimising likelihood directly leaves that band. This is the central empirical fact of the decoding literature.

Typical sampling (Meister et al., 2022, arXiv:2202.00666) follows that observation to its conclusion, selecting tokens whose information content -log p_i is closest to the distribution's entropy — targeting the expected surprise rather than the maximum probability.

Truncation as desmoothing (Hewitt, Manning & Liang, 2022, arXiv:2210.15191) provides the theoretical justification. A trained model is treated as the true distribution smoothed with a uniform component, an artefact of the training objective's aversion to assigning zero probability. Truncation removes that smoothing. Epsilon-sampling (absolute probability floor) and eta-sampling (entropy-dependent floor) follow from this framing.

Min-p (Nguyen et al., 2024, arXiv:2407.01082) sets the threshold relative to the top token's probability: keep tokens with p_i >= min_p * max_j p_j. It is more stable than top-p at high temperature, and is now widely available in open-weight serving stacks.

Repetition control, and its cost

Three mechanisms, all operating on logits before temperature:

text
repetition penalty:  z_i <- z_i / r   if z_i > 0 else z_i * r,  for tokens already generated
frequency penalty:   z_i <- z_i - alpha * count(i)
presence penalty:    z_i <- z_i - beta * [ count(i) > 0 ]
  • r > 1 is the repetition penalty (Keskar et al., 2019, CTRL).
  • count(i) is the number of times token i has appeared so far.
  • alpha, beta are the frequency and presence coefficients.

Be clear about what these are: they distort the model's distribution to patch a symptom. The multiplicative repetition penalty is particularly awkward, because its effect depends on the sign of the logit, which is an artefact of the parameterisation rather than anything meaningful.

They also damage legitimate repetition. Code with repeated identifiers, lists with a repeated structure, and any language with high function-word frequency all suffer. Apply sparingly and measure.

Contrastive search (Su et al., 2022, arXiv:2202.06417) is the more principled alternative: penalise a candidate by its maximum cosine similarity to the representations of previously generated tokens, trading off against model confidence. It reduces degeneration without distorting the probability model by fiat.

Beam search, and where it belongs

Beam search maintains b partial hypotheses and expands by cumulative log probability. It is correct for tasks with a sharply peaked target distribution — translation, speech recognition, constrained summarisation — where the goal genuinely is the most likely sequence.

For open-ended generation it fails, and Holtzman et al. explain why: the most likely sequence is not the most human-like one. Beam search also exhibits the beam search curse, where increasing b degrades quality on translation past a point, because longer beams find degenerate short outputs that score well under length-normalised likelihood.

Sampling temperature is not calibration temperature

These are conflated constantly and they are different objects.

Calibration temperature scaling (Guo et al., 2017, arXiv:1706.04599) fits a single scalar T on a held-out validation set to minimise negative log likelihood, so that predicted confidences match observed accuracies. It is a post-hoc correction with a fitted value, and it does not change the argmax, so accuracy is unchanged.

Sampling temperature is a free control knob at inference, chosen by taste, that changes which token is produced.

Same arithmetic, entirely different purpose. A model calibrated at T = 1.4 is not a model that should be sampled at 1.4.

Speculative decoding preserves the distribution

Worth knowing because it is the one acceleration technique that is not an approximation.

Speculative decoding (Leviathan, Kalman & Matias, 2023, arXiv:2211.17192; Chen et al., 2023) drafts k tokens with a small model, verifies them in one forward pass of the large model, and accepts a prefix using a modified rejection-sampling rule. The rule is constructed so the output distribution is provably identical to sampling from the large model directly.

Two implications. Speed-ups of two to three times cost nothing in quality, which is unusual. And the technique interacts with temperature: acceptance rates fall at low temperature, because the draft model's distribution diverges more sharply from the target's when both are near-deterministic.

Determinism, precisely

At T = 0, decoding is deterministic given fixed logits. Logits are not fixed in practice, for reasons unrelated to sampling:

  • Reduction order. Floating-point addition is not associative. GPU reductions partition work by block, and the partitioning depends on tensor shapes.
  • Batch-dependent kernels. Serving stacks select different matrix-multiply kernels by batch size. Your request's logits therefore depend on who else is being served alongside you.
  • Mixed precision. bf16 and fp16 accumulation amplify these differences relative to fp32.
  • Ties. Exact ties are broken by index order in most implementations, which is stable, but the tie itself may appear or disappear under the above.

Consequently a seed parameter gives best-effort reproducibility, not a guarantee. For regression tests, assert on semantic properties rather than exact strings, or pin batch size and precision.

Key references

  • Fan, A., Lewis, M. & Dauphin, Y. (2018). Hierarchical Neural Story Generation. arXiv:1805.04833
  • Holtzman, A. et al. (2020). The Curious Case of Neural Text Degeneration. ICLR. arXiv:1904.09751
  • Keskar, N. et al. (2019). CTRL: A Conditional Transformer Language Model for Controllable Generation. arXiv:1909.05858
  • Meister, C. et al. (2022). Typical Decoding for Natural Language Generation. arXiv:2202.00666
  • Su, Y. et al. (2022). A Contrastive Framework for Neural Text Generation. arXiv:2202.06417
  • Hewitt, J., Manning, C. & Liang, P. (2022). Truncation Sampling as Language Model Desmoothing. arXiv:2210.15191
  • Leviathan, Y., Kalman, M. & Matias, Y. (2023). Fast Inference from Transformers via Speculative Decoding. arXiv:2211.17192
  • Nguyen, M. et al. (2024). Min-p Sampling for Creative and Coherent LLM Outputs. arXiv:2407.01082
  • Guo, C. et al. (2017). On Calibration of Modern Neural Networks. arXiv:1706.04599

Open problems

No principled way to choose a decoding configuration. Temperature, top-p, min-p and penalties are selected by sweeping against a preference model or human judgement. There is no theory connecting task properties to optimal decoding parameters, and reported defaults transfer poorly between model families.

The degeneration diagnosis is better than the cure. Holtzman et al. established that likelihood maximisation and human-likeness diverge. Every subsequent method — nucleus, typical, contrastive, min-p — is a heuristic for staying in the right entropy band. None derives from a model of what human text distributions actually are.

Alignment tuning has changed the landscape and the guidance has not caught up. Instruction-tuned and RLHF-trained models have substantially lower output entropy than base models. Sampling advice developed on base models — much of the literature above — is being applied to models with very different distributions, and the empirical picture for aligned models is thin.

What to learn next

  • Hallucination — the error mode decoding cannot fix.
  • vLLM — where sampling parameters meet a real serving stack.
  • Structured output — constrained decoding as a different kind of control.

What to learn next

These follow on from what you just read.

  • Generative AI

    What is RAG?

    RAG means the model searches your documents first and then answers using what it found, instead of answering from memory.

  • Generative AI

    Vector databases

    A vector database stores text as numbers that capture meaning, then finds the closest matches to your question in milliseconds.

  • Generative AI

    Fine-tuning

    Fine-tuning continues training an already-trained model on your own examples, which changes how it behaves — and is the wrong tool for most problems beginners reach for it with.