How Text Is Generated

Top-k and nucleus sampling

Two ways to throw away the model's worst options before picking a word — keep a fixed number of them, or keep enough to cover a fixed share of the model's confidence.

On this page 7
  1. What the model hands you
  2. The two fixes
  3. How it works
  4. When each one fails
  5. Where you have already seen this
  6. Remember this
  7. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Top-k keeps a fixed number of the model's best guesses. Nucleus sampling keeps as many as it takes to cover a fixed share of the model's confidence.

Picture ordering at a busy roadside restaurant with a menu of two hundred items. You do not read all two hundred. Your eye lands on the five or six things that could plausibly be dinner, and you pick from those.

Now picture a stall that only sells three things. Your shortlist is three, not six, because there is nothing else worth listing.

That is the whole difference between the two methods. One always shortlists the same number of options. The other lets the shortlist grow and shrink depending on how many options are genuinely reasonable.

What the model hands you

At every step the model produces a score for every word it knows. Turn those scores into percentages and you get a ranked list.

The top handful are usually sensible. Below them stretches an enormous tail of words that are almost, but not quite, impossible.

Here is the trap. That tail is long. A model might know a hundred thousand words, and each junk word might carry a one-in-a-million chance. Add up a hundred thousand of those and the junk as a group is not rare at all.

Pick words at random from the full list often enough and you will eventually pick from the tail. One absurd word derails everything after it, because the model then treats its own mistake as context.

The two fixes

Top-k says: rank the words, keep the best k of them, delete the rest, then choose among the survivors.

Nucleus sampling, usually written as top-p, says: rank the words, and keep adding them to the shortlist until their chances together reach p. Then stop and choose among those.

The names are less important than the shapes. Top-k has a fixed shortlist. Top-p has a stretchy one.

How it works

   the model's ranked guesses after "I would like a cup of"

     chai     49 out of 100   ################
     coffee   33 out of 100   ###########
     water     8 out of 100   ###
     juice     6 out of 100   ##
     milk      4 out of 100   #
     petrol    almost none
     sand      almost none
     regret    almost none

   top-k keeping three  -> chai, coffee, water.        Always three.
   top-p reaching ninety -> chai, coffee, water, juice. As many as it takes.

Both cut off "petrol". That is the job.

When each one fails

Top-k fails when the model is honestly unsure. If fifteen words are all roughly equally good, keeping three throws away twelve fine choices. The writing then gets predictable.

Top-p fails in the other direction. If the model is very confident, top-p correctly keeps one word. But raise the creativity dial and top-p starts sweeping in nonsense. A flattened list needs many more words to reach the same total.

Neither is wrong. They fail on different steps, which is why people often use both together.

Where you have already seen this

  • A chat assistant giving a different answer to the same question each time.
  • A creative-writing tool with a slider labelled something like "randomness".
  • API settings named top_p and top_k that you have set without knowing what they cut.
  • A model that produces one bizarre word and then follows it off a cliff.

Remember this

  • The model's list of words has a long tail of near-impossible junk that adds up.
  • Top-k keeps a fixed number of the best; top-p keeps enough to reach a fixed share.
  • Top-k is too rigid when the model is unsure; top-p is too generous when the list is flattened.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install numpy

Both filters are about six lines each. Writing them yourself is the fastest way to stop guessing what your API parameters do.

Both filters, and where each one breaks

truncation.py
import numpy as np

VOCAB = ["chai", "coffee", "water", "juice", "milk", "petrol", "sand", "regret"]
LOGITS = np.array([4.0, 3.6, 2.2, 1.9, 1.5, -1.0, -2.0, -3.5])   # after "I would like a cup of"


def softmax(z):
    e = np.exp(z - z.max())
    return e / e.sum()


def show(title, probs):
    print(title)
    for w, p in sorted(zip(VOCAB, probs), key=lambda t: -t[1]):
        if p > 0:
            print(f"   {w:<8}{p:6.3f}  {'#' * round(p * 40)}")
    kept = int((probs > 0).sum())
    print(f"   -> {kept} of {len(VOCAB)} tokens can be chosen\n")


def top_k(logits, k):
    out = logits.copy()
    cutoff = np.sort(logits)[-k]          # the k-th largest value
    out[logits < cutoff] = -np.inf
    return out


def top_p(logits, p):
    order = np.argsort(-logits)           # most likely first
    probs = softmax(logits)[order]
    cum = np.cumsum(probs)
    # keep everything up to and INCLUDING the token that crosses p
    remove = cum - probs > p
    out = logits.copy()
    out[order[remove]] = -np.inf
    return out


show("raw distribution (temperature 1, nothing filtered)", softmax(LOGITS))
show("top-k, k = 3", softmax(top_k(LOGITS, 3)))
show("top-p, p = 0.90", softmax(top_p(LOGITS, 0.90)))

# --- where top-k goes wrong: a genuinely uncertain step -------------------
FLAT = np.array([1.05, 1.0, 0.98, 0.95, 0.93, 0.90, 0.88, 0.85])
print("A step where the model is honestly unsure (all options close):")
show("  top-k, k = 3  -> throws away five reasonable words", softmax(top_k(FLAT, 3)))
show("  top-p, p = 0.90 -> keeps them", softmax(top_p(FLAT, 0.90)))

# --- and where top-p goes wrong: a step the model is certain about --------
SHARP = np.array([9.0, 1.0, 0.9, 0.8, 0.7, 0.6, 0.5, 0.4])
print("A step where the model is certain:")
show("  top-p, p = 0.90 -> one token, correctly", softmax(top_p(SHARP, 0.90)))
show("  top-k, k = 3    -> drags in two words with ~0.03% mass", softmax(top_k(SHARP, 3)))

# --- what sampling actually does ------------------------------------------
rng = np.random.default_rng(7)
draws = rng.choice(len(VOCAB), size=10000, p=softmax(top_p(LOGITS, 0.90)))
counts = np.bincount(draws, minlength=len(VOCAB))
print("10,000 draws from the top-p 0.90 distribution (seed 7):")
for i in np.argsort(-counts):
    if counts[i]:
        print(f"   {VOCAB[i]:<8}{counts[i]:>6}")
Output
raw distribution (temperature 1, nothing filtered)
   chai     0.488  ####################
   coffee   0.327  #############
   water    0.081  ###
   juice    0.060  ##
   milk     0.040  ##
   petrol   0.003  
   sand     0.001  
   regret   0.000  
   -> 8 of 8 tokens can be chosen

top-k, k = 3
   chai     0.545  ######################
   coffee   0.365  ###############
   water    0.090  ####
   -> 3 of 8 tokens can be chosen

top-p, p = 0.90
   chai     0.511  ####################
   coffee   0.342  ##############
   water    0.084  ###
   juice    0.063  ###
   -> 4 of 8 tokens can be chosen

A step where the model is honestly unsure (all options close):
  top-k, k = 3  -> throws away five reasonable words
   chai     0.347  ##############
   coffee   0.330  #############
   water    0.323  #############
   -> 3 of 8 tokens can be chosen

  top-p, p = 0.90 -> keeps them
   chai     0.139  ######
   coffee   0.132  #####
   water    0.130  #####
   juice    0.126  #####
   milk     0.123  #####
   petrol   0.120  #####
   sand     0.117  #####
   regret   0.114  #####
   -> 8 of 8 tokens can be chosen

A step where the model is certain:
  top-p, p = 0.90 -> one token, correctly
   chai     1.000  ########################################
   -> 1 of 8 tokens can be chosen

  top-k, k = 3    -> drags in two words with ~0.03% mass
   chai     0.999  ########################################
   coffee   0.000  
   water    0.000  
   -> 3 of 8 tokens can be chosen

10,000 draws from the top-p 0.90 distribution (seed 7):
   chai      5099
   coffee    3369
   water      892
   juice      640

Reading the output

Filtering renormalises. chai was 0.488 in the raw distribution and 0.545 after top-k. Deleting options does not shrink the survivors' chances; it enlarges them, because the probabilities are recomputed over what is left. Truncation makes the model more decisive, not less.

Top-k gave three tokens on both distributions. Three when three was right, three when eight would have been right. The number k knows nothing about the step it is applied to. That is the flaw the next lesson attacks.

Top-p gave one token on the sharp step and eight on the flat one. Same setting, completely different shortlist, decided by the model's own confidence. This is why top-p became the default almost everywhere.

Top-k on the sharp step kept two tokens at 0.000. They round to zero at three decimals but they are not zero. Over ten thousand generated tokens, a per-step chance of three in ten thousand fires roughly three times. That is a real derailment every few paragraphs.

The 10,000 draws match the filtered probabilities. chai at 0.511 produced 5,099 of 10,000. Sampling is unbiased with respect to the distribution it is given, so all the interesting behaviour lives in how you shape that distribution before sampling.

Checking your implementation against a real library

Two lines of your own code, or a library's twenty. It is worth knowing they agree.

verify.py
import numpy as np, torch
from transformers import TopKLogitsWarper, TopPLogitsWarper

VOCAB = ["chai", "coffee", "water", "juice", "milk", "petrol", "sand", "regret"]
LOGITS = np.array([4.0, 3.6, 2.2, 1.9, 1.5, -1.0, -2.0, -3.5])


def softmax(z):
    e = np.exp(z - z.max())
    return e / e.sum()


def top_k(logits, k):
    out = logits.copy()
    out[logits < np.sort(logits)[-k]] = -np.inf
    return out


def top_p(logits, p):
    order = np.argsort(-logits)
    probs = softmax(logits)[order]
    out = logits.copy()
    out[order[np.cumsum(probs) - probs > p]] = -np.inf
    return out


t = torch.tensor(LOGITS, dtype=torch.float32).unsqueeze(0)
ids = torch.zeros((1, 1), dtype=torch.long)
for k in (1, 2, 3, 5):
    hf = TopKLogitsWarper(k)(ids, t.clone())[0].numpy()
    print(f"top_k={k}  same survivors as HuggingFace:",
          np.array_equal(np.isfinite(hf), np.isfinite(top_k(LOGITS, k))))
for p in (0.5, 0.9, 0.95, 0.99):
    hf = TopPLogitsWarper(p, min_tokens_to_keep=1)(ids, t.clone())[0].numpy()
    print(f"top_p={p} same survivors as HuggingFace:",
          np.array_equal(np.isfinite(hf), np.isfinite(top_p(LOGITS, p))))
Output
top_k=1  same survivors as HuggingFace: True
top_k=2  same survivors as HuggingFace: True
top_k=3  same survivors as HuggingFace: True
top_k=5  same survivors as HuggingFace: True
top_p=0.5 same survivors as HuggingFace: True
top_p=0.9 same survivors as HuggingFace: True
top_p=0.95 same survivors as HuggingFace: True
top_p=0.99 same survivors as HuggingFace: True

Written against transformers 5.6.2 and torch 2.5.1. These two warper classes have been stable for years; the surrounding generation API has not been, so pin what you rely on.

Order of operations, which changes the result

HuggingFace applies these in a fixed order: temperature, then top-k, then top-p, then min-p. That order matters and is not arbitrary.

Temperature first means top-p sees the reheated distribution, so raising temperature widens the nucleus. If top-p ran first, temperature could only redistribute mass inside an already-chosen shortlist, and the two knobs would barely interact.

Different servers have historically used different orders. When your local model behaves unlike the hosted one at identical settings, this is the first thing to check.

Common mistakes

Setting top_p=1.0 and thinking it is off. It is off. top_p=0.0 is the confusing case: implementations keep at least one token, so it degenerates to greedy rather than erroring.

Leaving top_k=50 on by default. It was the old HuggingFace default and it is still baked into many model config files. It quietly caps every step at fifty options. On a model with a hundred-thousand-token vocabulary that is a much harder cut than it sounds.

Sampling from unnormalised scores. After masking you must renormalise. np.random.choice raises an error if probabilities do not sum to one; torch.multinomial does not, and will happily sample from something meaningless.

Tuning top-p to fix repetition. Repetition is a different failure with a different fix, covered in repetition penalties. Raising top-p to escape a loop mostly adds noise.

Comparing settings across models. A good top_p depends on vocabulary size and how sharp a model's distributions are. Numbers copied from a blog post about a different model are a starting point, not a setting.

Try it yourself

Set k = 1 and confirm the output equals greedy decoding. Then take FLAT, apply top_p at 0.5, 0.9 and 0.99, and count survivors at each. Plot that count against p. The curve you get explains why 0.9 and 0.95 are the values everybody ends up using.

What to learn next

Researcher — Mathematics and papers.

Definitions

Let $P(x \mid x_{<t})$ be the model's next-token distribution over vocabulary $V$, and let $x^{(1)}, x^{(2)}, \dots$ be tokens sorted by descending probability.

Top-k (Fan et al., 2018, arxiv.org/abs/1805.04833) keeps $V_k = {x^{(1)}, \dots, x^{(k)}}$ and renormalises:

$$ P'(x) = \frac{P(x)}{\sum_{y \in V_k} P(y)} \ \text{for} \ x \in V_k, \qquad 0 \ \text{otherwise} $$

Nucleus sampling (Holtzman et al., ICLR 2020, arxiv.org/abs/1904.09751) keeps the smallest set $V_p$ satisfying

$$ \sum_{x \in V_p} P(x) \ge p $$

so $|V_p|$ varies per step. Temperature $T$ is applied to the logits before either filter:

$$ P_T(x) = \frac{\exp(z_x / T)}{\sum_{y} \exp(z_y / T)} $$

Note that $T$ and $p$ are not independent. Raising $T$ flattens $P$, which enlarges $V_p$ at fixed $p$, so tuning them separately is misleading. This coupling is a large part of the case for the confidence-relative methods in the next lesson.

Why sampling from the full distribution fails

Holtzman et al. named this neural text degeneration and gave it two halves.

Maximisation-based decoding — greedy and beam search — produces text that repeats and is measurably less surprising than human text. Human writing does not sit at the mode of a language model; it wanders around it.

Pure ancestral sampling produces the opposite failure. The tail of a large-vocabulary distribution carries non-trivial aggregate mass. Sampling from it eventually selects a token the model considers implausible, and autoregressive conditioning turns one error into a permanent change of direction.

The diagnostic they introduced is worth reusing: compare the distribution of per-token probabilities in generated text against human text. Beam search output sits far too high; unfiltered sampling has a tail human text does not have. Nucleus sampling matches human statistics far better on both.

The truncation family

MethodKept setAdapts to confidence
Top-kfixed size $k$no
Top-psmallest set with mass $\ge p$partly
Typicaltokens with surprise near the entropyyes
Epsilon$P(x) > \epsilon$no
Eta$P(x) > \min!\left(\epsilon, \sqrt{\epsilon\, e^{-H}}\right)$yes
Min-p$P(x) \ge m \cdot \max_y P(y)$yes
Top-hprefix whose cumulative entropy stays under $\tau = \alpha H$yes

Epsilon and eta sampling come from Hewitt et al., Truncation Sampling as Language Model Desmoothing, 2022 (arxiv.org/abs/2210.15191). It frames truncation as undoing the smoothing that maximum-likelihood training bakes in. That framing is the most useful theoretical account available: the model's tail is not a belief the model holds, it is an artefact of never being allowed to assign zero.

All seven are implemented in transformers as logits processors, and vLLM implements a subset directly in its sampler.

Practical settings, with the caveat they deserve

Common defaults cluster around $p \in [0.9, 0.95]$ with $T \in [0.7, 1.0]$ for open-ended text, and $T \to 0$ for extraction, classification and code. These are conventions, not results. There is no published evidence that a single $(T, p)$ pair transfers across model families, and vocabulary size alone should make you doubt it.

Two effects worth knowing when interpreting benchmarks:

  • Truncation reduces measured diversity across samples while raising per-sample quality. Any benchmark reporting only one of the two can be gamed by moving $p$.
  • Reasoning models trained with reinforcement learning are often more sensitive to sampling settings than instruction-tuned models. Model cards for them now frequently specify a required $(T, p, k)$ triple. Overriding it is a real quality regression, not a stylistic choice.

Papers

What to learn next