Top-p (nucleus sampling)
In one sentence Top-p sampling keeps only the smallest set of next-token candidates whose probabilities add up to p, then samples within that set.
Updated
Top-p sampling, also called nucleus sampling, cuts the candidate list to the smallest group of tokens whose combined probability reaches p, and samples only from that group.
Think of a buffet where dishes are arranged by popularity. Top-p 0.9 says: walk down the line, adding dishes until you have covered 90% of what people actually eat, then choose only among those. On a day when one biryani dominates, your plate comes from two or three dishes. At a varied wedding spread, the 90% might span fifteen. The cutoff adapts to the situation — that adaptiveness is the entire point.
Contrast with top-k, which always keeps a fixed number of candidates regardless of how confident the model is. Top-p keeps few candidates when the model is sure and many when it is genuinely uncertain:
after "The capital of France is" → " Paris" alone covers 0.9 → 1 candidate
after "My favourite hobby is" → dozens needed to reach 0.9 → wide, varied choiceWhat it buys you: the long tail of thousands of terrible-but-possible tokens is deleted before sampling, which prevents the rare bizarre word that pure temperature sampling occasionally lets through. Typical values run 0.9 to 0.95; 1.0 disables the filter.
Practical convention: tune temperature or top-p, not both aggressively at once — most API defaults (temperature ~0.7-1.0 with top-p ~0.9-1.0) are already sensible for chat.
Where to go next
- Full lesson: How LLMs work
- Related terms: temperature, top-k, logits, hallucination