Top-k sampling
In one sentence Top-k sampling restricts each next-token choice to the k highest-probability candidates, a fixed-size shortlist.
Updated
Top-k sampling keeps only the k most probable next tokens — a fixed-length shortlist — and samples from those alone.
It is hiring from a shortlist: whatever the applicant pool looks like, exactly the top 40 CVs reach the interview stage, and the job goes to someone among them. The model does the same at every generation step — rank all vocabulary tokens by probability, keep the top k (40 or 50 are common), renormalise, sample.
The rigidity of that fixed k is both the appeal and the flaw, and the flaw shows at both extremes of model confidence:
model very sure : " Paris" 0.97, rest are junk → k=40 still admits 39 junk options
model genuinely open: 200 plausible continuations → k=40 cuts off 160 fine onesTop-p was invented precisely to fix this, by making the shortlist size adapt to the model's confidence. In practice top-p has largely replaced top-k as the default filter, and several modern APIs no longer expose top-k at all. You will still meet it in local-model tooling (llama.cpp, Ollama expose it), in older papers, and combined with other filters — the filters stack: top-k first, then top-p, then temperature-scaled sampling among the survivors.
Boundary cases make the family clear: k = 1 is greedy-decoding — always take the single best token — and k = vocabulary size is no filter at all.
Where to go next
- Full lesson: How LLMs work
- Related terms: top-p, temperature, greedy-decoding, logits