AI glossary

Speculative decoding

In one sentence Speculative decoding has a small fast model draft several tokens ahead and the big model verify them in one pass — same output, delivered faster.

By Updated

Speculative decoding speeds up generation by letting a small "draft" model guess several tokens ahead, then having the large model check all the guesses in a single pass.

It is a senior lawyer with a junior drafting for her. The junior writes five sentences of routine boilerplate; the senior reads them in one glance and signs off. When the junior guesses wrong, the senior crosses out from the first error and writes the correction herself. The document that leaves the office is identical to one the senior would have written alone — but her expensive attention was spent verifying, which is fast, instead of composing, which is slow.

The technical basis: for a transformer, checking k proposed tokens costs about the same as generating one, because verification processes them in parallel while generation is one-at-a-time. So a draft model (a small distilled sibling, or extra prediction heads on the model itself) proposes 4-8 tokens; the big model scores them all at once; matching tokens are accepted and the first mismatch is corrected.

draft: "The capital of France is Paris ."
big model verifies all 7 at once → accepts 7 tokens for ~1 token's latency

A rejection-sampling rule guarantees the output distribution is exactly the big model's — this is a pure speed trick with no quality trade. Typical speedups run 2-3×, best on predictable text (code, boilerplate) where the drafter guesses well, and it is now standard in major serving stacks.

Where to go next