Top-k and nucleus sampling
Two ways to throw away the model's worst options before picking a word — keep a fixed number of them, or keep enough to cover a fixed share of the model's confidence.
- 15 min read
- 3 reading levels
- Updated
Read these first
On this page 7
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Top-k keeps a fixed number of the model's best guesses. Nucleus sampling keeps as many as it takes to cover a fixed share of the model's confidence.
Picture ordering at a busy roadside restaurant with a menu of two hundred items. You do not read all two hundred. Your eye lands on the five or six things that could plausibly be dinner, and you pick from those.
Now picture a stall that only sells three things. Your shortlist is three, not six, because there is nothing else worth listing.
That is the whole difference between the two methods. One always shortlists the same number of options. The other lets the shortlist grow and shrink depending on how many options are genuinely reasonable.
What the model hands you
At every step the model produces a score for every word it knows. Turn those scores into percentages and you get a ranked list.
The top handful are usually sensible. Below them stretches an enormous tail of words that are almost, but not quite, impossible.
Here is the trap. That tail is long. A model might know a hundred thousand words, and each junk word might carry a one-in-a-million chance. Add up a hundred thousand of those and the junk as a group is not rare at all.
Pick words at random from the full list often enough and you will eventually pick from the tail. One absurd word derails everything after it, because the model then treats its own mistake as context.
The two fixes
Top-k says: rank the words, keep the best k of them, delete the rest, then choose among the survivors.
Nucleus sampling, usually written as top-p, says: rank the words, and keep adding them to the shortlist until their chances together reach p. Then stop and choose among those.
The names are less important than the shapes. Top-k has a fixed shortlist. Top-p has a stretchy one.
How it works
the model's ranked guesses after "I would like a cup of"
chai 49 out of 100 ################
coffee 33 out of 100 ###########
water 8 out of 100 ###
juice 6 out of 100 ##
milk 4 out of 100 #
petrol almost none
sand almost none
regret almost none
top-k keeping three -> chai, coffee, water. Always three.
top-p reaching ninety -> chai, coffee, water, juice. As many as it takes.Both cut off "petrol". That is the job.
When each one fails
Top-k fails when the model is honestly unsure. If fifteen words are all roughly equally good, keeping three throws away twelve fine choices. The writing then gets predictable.
Top-p fails in the other direction. If the model is very confident, top-p correctly keeps one word. But raise the creativity dial and top-p starts sweeping in nonsense. A flattened list needs many more words to reach the same total.
Neither is wrong. They fail on different steps, which is why people often use both together.
Where you have already seen this
- A chat assistant giving a different answer to the same question each time.
- A creative-writing tool with a slider labelled something like "randomness".
- API settings named
top_pandtop_kthat you have set without knowing what they cut. - A model that produces one bizarre word and then follows it off a cliff.
Remember this
- The model's list of words has a long tail of near-impossible junk that adds up.
- Top-k keeps a fixed number of the best; top-p keeps enough to reach a fixed share.
- Top-k is too rigid when the model is unsure; top-p is too generous when the list is flattened.
What to learn next
- Min-p and typical sampling — truncation that adapts to how sure the model is.
- Temperature and sampling — the knob these filters are applied after.
- Reading logprobs — inspecting the distribution these filters cut.
Developer — Code and libraries.
Setup
pip install numpyBoth filters are about six lines each. Writing them yourself is the fastest way to stop guessing what your API parameters do.
Both filters, and where each one breaks
import numpy as np
VOCAB = ["chai", "coffee", "water", "juice", "milk", "petrol", "sand", "regret"]
LOGITS = np.array([4.0, 3.6, 2.2, 1.9, 1.5, -1.0, -2.0, -3.5]) # after "I would like a cup of"
def softmax(z):
e = np.exp(z - z.max())
return e / e.sum()
def show(title, probs):
print(title)
for w, p in sorted(zip(VOCAB, probs), key=lambda t: -t[1]):
if p > 0:
print(f" {w:<8}{p:6.3f} {'#' * round(p * 40)}")
kept = int((probs > 0).sum())
print(f" -> {kept} of {len(VOCAB)} tokens can be chosen\n")
def top_k(logits, k):
out = logits.copy()
cutoff = np.sort(logits)[-k] # the k-th largest value
out[logits < cutoff] = -np.inf
return out
def top_p(logits, p):
order = np.argsort(-logits) # most likely first
probs = softmax(logits)[order]
cum = np.cumsum(probs)
# keep everything up to and INCLUDING the token that crosses p
remove = cum - probs > p
out = logits.copy()
out[order[remove]] = -np.inf
return out
show("raw distribution (temperature 1, nothing filtered)", softmax(LOGITS))
show("top-k, k = 3", softmax(top_k(LOGITS, 3)))
show("top-p, p = 0.90", softmax(top_p(LOGITS, 0.90)))
# --- where top-k goes wrong: a genuinely uncertain step -------------------
FLAT = np.array([1.05, 1.0, 0.98, 0.95, 0.93, 0.90, 0.88, 0.85])
print("A step where the model is honestly unsure (all options close):")
show(" top-k, k = 3 -> throws away five reasonable words", softmax(top_k(FLAT, 3)))
show(" top-p, p = 0.90 -> keeps them", softmax(top_p(FLAT, 0.90)))
# --- and where top-p goes wrong: a step the model is certain about --------
SHARP = np.array([9.0, 1.0, 0.9, 0.8, 0.7, 0.6, 0.5, 0.4])
print("A step where the model is certain:")
show(" top-p, p = 0.90 -> one token, correctly", softmax(top_p(SHARP, 0.90)))
show(" top-k, k = 3 -> drags in two words with ~0.03% mass", softmax(top_k(SHARP, 3)))
# --- what sampling actually does ------------------------------------------
rng = np.random.default_rng(7)
draws = rng.choice(len(VOCAB), size=10000, p=softmax(top_p(LOGITS, 0.90)))
counts = np.bincount(draws, minlength=len(VOCAB))
print("10,000 draws from the top-p 0.90 distribution (seed 7):")
for i in np.argsort(-counts):
if counts[i]:
print(f" {VOCAB[i]:<8}{counts[i]:>6}")raw distribution (temperature 1, nothing filtered) chai 0.488 #################### coffee 0.327 ############# water 0.081 ### juice 0.060 ## milk 0.040 ## petrol 0.003 sand 0.001 regret 0.000 -> 8 of 8 tokens can be chosen top-k, k = 3 chai 0.545 ###################### coffee 0.365 ############### water 0.090 #### -> 3 of 8 tokens can be chosen top-p, p = 0.90 chai 0.511 #################### coffee 0.342 ############## water 0.084 ### juice 0.063 ### -> 4 of 8 tokens can be chosen A step where the model is honestly unsure (all options close): top-k, k = 3 -> throws away five reasonable words chai 0.347 ############## coffee 0.330 ############# water 0.323 ############# -> 3 of 8 tokens can be chosen top-p, p = 0.90 -> keeps them chai 0.139 ###### coffee 0.132 ##### water 0.130 ##### juice 0.126 ##### milk 0.123 ##### petrol 0.120 ##### sand 0.117 ##### regret 0.114 ##### -> 8 of 8 tokens can be chosen A step where the model is certain: top-p, p = 0.90 -> one token, correctly chai 1.000 ######################################## -> 1 of 8 tokens can be chosen top-k, k = 3 -> drags in two words with ~0.03% mass chai 0.999 ######################################## coffee 0.000 water 0.000 -> 3 of 8 tokens can be chosen 10,000 draws from the top-p 0.90 distribution (seed 7): chai 5099 coffee 3369 water 892 juice 640
Reading the output
Filtering renormalises. chai was 0.488 in the raw distribution and 0.545 after top-k. Deleting options does not shrink the survivors' chances; it enlarges them, because the probabilities are recomputed over what is left. Truncation makes the model more decisive, not less.
Top-k gave three tokens on both distributions. Three when three was right, three when eight would have been right. The number k knows nothing about the step it is applied to. That is the flaw the next lesson attacks.
Top-p gave one token on the sharp step and eight on the flat one. Same setting, completely different shortlist, decided by the model's own confidence. This is why top-p became the default almost everywhere.
Top-k on the sharp step kept two tokens at 0.000. They round to zero at three decimals but they are not zero. Over ten thousand generated tokens, a per-step chance of three in ten thousand fires roughly three times. That is a real derailment every few paragraphs.
The 10,000 draws match the filtered probabilities. chai at 0.511 produced 5,099 of 10,000. Sampling is unbiased with respect to the distribution it is given, so all the interesting behaviour lives in how you shape that distribution before sampling.
Checking your implementation against a real library
Two lines of your own code, or a library's twenty. It is worth knowing they agree.
import numpy as np, torch
from transformers import TopKLogitsWarper, TopPLogitsWarper
VOCAB = ["chai", "coffee", "water", "juice", "milk", "petrol", "sand", "regret"]
LOGITS = np.array([4.0, 3.6, 2.2, 1.9, 1.5, -1.0, -2.0, -3.5])
def softmax(z):
e = np.exp(z - z.max())
return e / e.sum()
def top_k(logits, k):
out = logits.copy()
out[logits < np.sort(logits)[-k]] = -np.inf
return out
def top_p(logits, p):
order = np.argsort(-logits)
probs = softmax(logits)[order]
out = logits.copy()
out[order[np.cumsum(probs) - probs > p]] = -np.inf
return out
t = torch.tensor(LOGITS, dtype=torch.float32).unsqueeze(0)
ids = torch.zeros((1, 1), dtype=torch.long)
for k in (1, 2, 3, 5):
hf = TopKLogitsWarper(k)(ids, t.clone())[0].numpy()
print(f"top_k={k} same survivors as HuggingFace:",
np.array_equal(np.isfinite(hf), np.isfinite(top_k(LOGITS, k))))
for p in (0.5, 0.9, 0.95, 0.99):
hf = TopPLogitsWarper(p, min_tokens_to_keep=1)(ids, t.clone())[0].numpy()
print(f"top_p={p} same survivors as HuggingFace:",
np.array_equal(np.isfinite(hf), np.isfinite(top_p(LOGITS, p))))top_k=1 same survivors as HuggingFace: True top_k=2 same survivors as HuggingFace: True top_k=3 same survivors as HuggingFace: True top_k=5 same survivors as HuggingFace: True top_p=0.5 same survivors as HuggingFace: True top_p=0.9 same survivors as HuggingFace: True top_p=0.95 same survivors as HuggingFace: True top_p=0.99 same survivors as HuggingFace: True
Written against transformers 5.6.2 and torch 2.5.1. These two warper classes have been stable for years; the surrounding generation API has not been, so pin what you rely on.
Order of operations, which changes the result
HuggingFace applies these in a fixed order: temperature, then top-k, then top-p, then min-p. That order matters and is not arbitrary.
Temperature first means top-p sees the reheated distribution, so raising temperature widens the nucleus. If top-p ran first, temperature could only redistribute mass inside an already-chosen shortlist, and the two knobs would barely interact.
Different servers have historically used different orders. When your local model behaves unlike the hosted one at identical settings, this is the first thing to check.
Common mistakes
Setting top_p=1.0 and thinking it is off. It is off. top_p=0.0 is the confusing case: implementations keep at least one token, so it degenerates to greedy rather than erroring.
Leaving top_k=50 on by default. It was the old HuggingFace default and it is still baked into many model config files. It quietly caps every step at fifty options. On a model with a hundred-thousand-token vocabulary that is a much harder cut than it sounds.
Sampling from unnormalised scores. After masking you must renormalise. np.random.choice raises an error if probabilities do not sum to one; torch.multinomial does not, and will happily sample from something meaningless.
Tuning top-p to fix repetition. Repetition is a different failure with a different fix, covered in repetition penalties. Raising top-p to escape a loop mostly adds noise.
Comparing settings across models. A good top_p depends on vocabulary size and how sharp a model's distributions are. Numbers copied from a blog post about a different model are a starting point, not a setting.
Try it yourself
Set k = 1 and confirm the output equals greedy decoding. Then take FLAT, apply top_p at 0.5, 0.9 and 0.99, and count survivors at each. Plot that count against p. The curve you get explains why 0.9 and 0.95 are the values everybody ends up using.
What to learn next
- Min-p and typical sampling — truncation that adapts to how sure the model is.
- Temperature and sampling — the knob these filters are applied after.
- Reading logprobs — inspecting the distribution these filters cut.
Researcher — Mathematics and papers.
Definitions
Let $P(x \mid x_{<t})$ be the model's next-token distribution over vocabulary $V$, and let $x^{(1)}, x^{(2)}, \dots$ be tokens sorted by descending probability.
Top-k (Fan et al., 2018, arxiv.org/abs/1805.04833) keeps $V_k = {x^{(1)}, \dots, x^{(k)}}$ and renormalises:
$$ P'(x) = \frac{P(x)}{\sum_{y \in V_k} P(y)} \ \text{for} \ x \in V_k, \qquad 0 \ \text{otherwise} $$
Nucleus sampling (Holtzman et al., ICLR 2020, arxiv.org/abs/1904.09751) keeps the smallest set $V_p$ satisfying
$$ \sum_{x \in V_p} P(x) \ge p $$
so $|V_p|$ varies per step. Temperature $T$ is applied to the logits before either filter:
$$ P_T(x) = \frac{\exp(z_x / T)}{\sum_{y} \exp(z_y / T)} $$
Note that $T$ and $p$ are not independent. Raising $T$ flattens $P$, which enlarges $V_p$ at fixed $p$, so tuning them separately is misleading. This coupling is a large part of the case for the confidence-relative methods in the next lesson.
Why sampling from the full distribution fails
Holtzman et al. named this neural text degeneration and gave it two halves.
Maximisation-based decoding — greedy and beam search — produces text that repeats and is measurably less surprising than human text. Human writing does not sit at the mode of a language model; it wanders around it.
Pure ancestral sampling produces the opposite failure. The tail of a large-vocabulary distribution carries non-trivial aggregate mass. Sampling from it eventually selects a token the model considers implausible, and autoregressive conditioning turns one error into a permanent change of direction.
The diagnostic they introduced is worth reusing: compare the distribution of per-token probabilities in generated text against human text. Beam search output sits far too high; unfiltered sampling has a tail human text does not have. Nucleus sampling matches human statistics far better on both.
The truncation family
| Method | Kept set | Adapts to confidence |
|---|---|---|
| Top-k | fixed size $k$ | no |
| Top-p | smallest set with mass $\ge p$ | partly |
| Typical | tokens with surprise near the entropy | yes |
| Epsilon | $P(x) > \epsilon$ | no |
| Eta | $P(x) > \min!\left(\epsilon, \sqrt{\epsilon\, e^{-H}}\right)$ | yes |
| Min-p | $P(x) \ge m \cdot \max_y P(y)$ | yes |
| Top-h | prefix whose cumulative entropy stays under $\tau = \alpha H$ | yes |
Epsilon and eta sampling come from Hewitt et al., Truncation Sampling as Language Model Desmoothing, 2022 (arxiv.org/abs/2210.15191). It frames truncation as undoing the smoothing that maximum-likelihood training bakes in. That framing is the most useful theoretical account available: the model's tail is not a belief the model holds, it is an artefact of never being allowed to assign zero.
All seven are implemented in transformers as logits processors, and vLLM implements a subset directly in its sampler.
Practical settings, with the caveat they deserve
Common defaults cluster around $p \in [0.9, 0.95]$ with $T \in [0.7, 1.0]$ for open-ended text, and $T \to 0$ for extraction, classification and code. These are conventions, not results. There is no published evidence that a single $(T, p)$ pair transfers across model families, and vocabulary size alone should make you doubt it.
Two effects worth knowing when interpreting benchmarks:
- Truncation reduces measured diversity across samples while raising per-sample quality. Any benchmark reporting only one of the two can be gamed by moving $p$.
- Reasoning models trained with reinforcement learning are often more sensitive to sampling settings than instruction-tuned models. Model cards for them now frequently specify a required $(T, p, k)$ triple. Overriding it is a real quality regression, not a stylistic choice.
Papers
- Fan et al., Hierarchical Neural Story Generation, 2018 — arxiv.org/abs/1805.04833
- Holtzman et al., The Curious Case of Neural Text Degeneration, ICLR 2020 — arxiv.org/abs/1904.09751
- Hewitt et al., Truncation Sampling as Language Model Desmoothing, 2022 — arxiv.org/abs/2210.15191
- Meister et al., Locally Typical Sampling, TACL 2023 — arxiv.org/abs/2202.00666
What to learn next
- Min-p and typical sampling — truncation that adapts to how sure the model is.
- Temperature and sampling — the knob these filters are applied after.
- Reading logprobs — inspecting the distribution these filters cut.