Greedy decoding
In one sentence Greedy decoding always picks the single most probable next token — deterministic, fast, and prone to repetitive, short-sighted text.
Updated
Greedy decoding generates text by always choosing the single highest-probability next token, with no randomness at all.
It is driving across a city by taking, at every junction, whichever road looks widest right now. Each individual choice is locally reasonable. The route as a whole can still be poor — the widest first road may feed into the most jammed flyover. Greedy decoding has exactly this short-sightedness: the best next token repeatedly is not the best sentence, because one early safe-looking word can commit the model to a weaker continuation.
Its virtues are real: deterministic (same prompt, same output — valuable for tests and caching), cheapest possible decoding, and entirely adequate when the answer is highly constrained — extraction, classification labels, code completion where one continuation dominates.
Its signature failure is degeneration on open-ended text: bland phrasing, and loops where a phrase repeats because each repetition makes the next repetition more probable.
greedy, open-ended: "The market was busy. The market was busy. The market..."The alternatives each relax greediness differently. Sampling with temperature and top-p injects controlled randomness, fixing repetition and adding variety. Beam-search keeps several candidate routes alive to fix the short-sightedness. Note that "temperature 0" in most APIs is greedy decoding by another name.
Where to go next
- Full lesson: How LLMs work
- Related terms: beam-search, temperature, top-p, logits