Speculative decoding
A small fast model guesses several tokens ahead and the big model checks them all in one pass, producing exactly the text the big model would have produced, several times faster.
- 14 min read
- 3 reading levels
- Updated
Read these first
On this page 7
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
A small model guesses the next few words. The big model checks all of them at once, and keeps the ones that were right.
Think about a junior colleague drafting a letter for you. They write four sentences. You read all four in one glance, keep the three that are fine, fix the fourth, and hand it back.
You still read everything. You still approved every word. But you wrote far less, so the letter was finished sooner.
The remarkable part is that the final letter is exactly the letter you would have written alone. Not similar. Identical in distribution.
Why this helps at all
Writing one token needs the model to fetch all of its own memory. That fetching is the slow part, and it costs the same whether the model produces one token or checks eight.
So checking eight tokens is nearly as cheap as producing one.
The small model is maybe twenty times cheaper to run. Letting it guess, then checking its guesses in bulk, trades expensive slow steps for cheap fast ones.
How it works
1. the small model writes 4 guesses, quickly
"the office is closed"
2. the big model checks all 4 in ONE pass
the -> agrees
office -> agrees
is -> agrees
closed -> disagrees, wants "shut"
3. keep the 3 it agreed with, replace the 4th with its own choice
-> "the office is shut"
4 tokens produced. One expensive pass used.If the small model guesses well, you get several tokens per expensive pass. If it guesses badly, you get one, which is what you had before. The floor is the old speed.
The part that sounds impossible
The output is not an approximation. It is drawn from exactly the same distribution as the big model working alone.
That is a real mathematical guarantee, not marketing. It comes from a careful rule about when to accept a guess. And what to do when one is rejected. Rejecting badly would bias the output toward whatever the small model likes. The rule is designed to prevent exactly that.
You do not lose quality. You lose only the memory and the complexity of running two models.
When it works and when it does not
It works beautifully when text is predictable. Boilerplate, code, formatted output, quoting a document back — the small model gets long runs right.
It works poorly on genuinely creative or unusual text. The small model disagrees constantly, and every disagreement wastes its work.
It also stops helping when the server is already busy. With many users, the expensive model is no longer waiting around. There is no idle capacity left to fill with checking.
Where you have already seen this
- Phone keyboards suggesting the next three words, which you accept with one tap.
- A hosted model that got noticeably faster with no announced model change.
- Code assistants completing a whole line at once rather than a character at a time.
Remember this
- A small model guesses ahead; the big model verifies many guesses in one pass.
- The output matches what the big model would have produced on its own.
- The gain comes from predictable text and an under-loaded server.
What to learn next
- Medusa, EAGLE and self-speculation — removing the second model entirely.
- vLLM — where you actually switch this on.
- Knowledge distillation — how a good draft model gets trained in the first place.
Developer — Code and libraries.
Setup
pip install numpyThe acceptance rule is the whole algorithm, and it fits in six lines. Running it a few hundred thousand times against a known target distribution is how you convince yourself the guarantee is real.
The algorithm, and a proof by counting
import numpy as np
VOCAB = ["the", "train", "is", "late", "again", "today"]
TARGET = np.array([0.40, 0.25, 0.15, 0.10, 0.06, 0.04]) # what the big model wants
DRAFT = np.array([0.35, 0.30, 0.20, 0.10, 0.03, 0.02]) # what the small model guesses
rng = np.random.default_rng(42)
def speculative_step(p, q, rng):
"""One token. p = target distribution, q = draft distribution.
Returns (token, accepted?) and provably samples from p."""
x = rng.choice(len(q), p=q) # the draft's guess, free
if rng.random() < min(1.0, p[x] / q[x]): # the acceptance test
return x, True
residual = np.maximum(p - q, 0.0) # what the draft under-supplied
residual /= residual.sum()
return rng.choice(len(p), p=residual), False # resample, still one target call
N = 200_000
tokens = np.empty(N, dtype=int)
accepted = 0
for i in range(N):
tokens[i], ok = speculative_step(TARGET, DRAFT, rng)
accepted += ok
empirical = np.bincount(tokens, minlength=len(VOCAB)) / N
print(f"{'token':<8}{'target p':>10}{'draft q':>10}{'observed':>10}")
for w, p, q, e in zip(VOCAB, TARGET, DRAFT, empirical):
print(f"{w:<8}{p:>10.3f}{q:>10.3f}{e:>10.3f}")
print(f"\nlargest gap between target and observed: {np.abs(empirical - TARGET).max():.4f}")
print(f"acceptance rate: {accepted / N:.3f}")
print(f"1 - total variation distance between p and q: "
f"{1 - 0.5 * np.abs(TARGET - DRAFT).sum():.3f}")
# ---- how many tokens does one verification pass actually buy? ----
GAMMA = 4
def one_round(rng):
"""Draft GAMMA tokens, verify them all in one target pass."""
for i in range(GAMMA):
_, ok = speculative_step(TARGET, DRAFT, rng)
if not ok:
return i + 1 # i accepted, plus the corrected one
return GAMMA + 1 # all accepted, plus one free bonus token
rounds = 100_000
produced = sum(one_round(rng) for _ in range(rounds))
alpha = accepted / N
expected = (1 - alpha ** (GAMMA + 1)) / (1 - alpha)
print(f"\ndrafting {GAMMA} tokens per round, {rounds:,} rounds")
print(f" measured tokens per target pass : {produced / rounds:.3f}")
print(f" formula (1-a^(g+1))/(1-a) : {expected:.3f}")token target p draft q observed the 0.400 0.350 0.400 train 0.250 0.300 0.251 is 0.150 0.200 0.150 late 0.100 0.100 0.100 again 0.060 0.030 0.059 today 0.040 0.020 0.040 largest gap between target and observed: 0.0012 acceptance rate: 0.901 1 - total variation distance between p and q: 0.900 drafting 4 tokens per round, 100,000 rounds measured tokens per target pass : 4.096 formula (1-a^(g+1))/(1-a) : 4.100
Seeded, so these exact numbers reproduce. Checked on NumPy 1.26 and 2.4, on Python 3.10 and 3.13.
Reading the output
The observed column matches the target column, not the draft. The draft wanted train at 0.300; the target wanted 0.250; the output landed on 0.251. The draft wanted today at 0.020, the target at 0.040, and the output gave 0.040. The draft's bias has been removed exactly. Largest deviation across the whole vocabulary is 0.0012, which is sampling noise at 200,000 draws.
This is the entire claim of the method, verified by counting rather than believed from a paper.
Acceptance rate 0.901 against a predicted 0.900. That is not a coincidence. The expected acceptance rate equals $1$ minus the total variation distance between the two distributions. The paper proves it; this run confirms it to three decimals. It also tells you what to optimise: pick a draft model whose distribution is close to the target's, and the entire speed-up follows.
4.096 tokens per target pass, against a formula prediction of 4.100. Four expensive passes' worth of work for one. That is the speed-up, in the only unit that matters.
The +1 in return GAMMA + 1 is free and easy to miss. When the target verifies four draft tokens, its forward pass also produced logits for the position after the fourth. That token needs no verification, so a fully-accepted round yields five tokens. Implementations that drop it lose a measurable fraction of the benefit.
The residual distribution, which is where people go wrong
On rejection, the naive move is to sample from the target distribution p. That is wrong, and it biases the result.
The reason is that acceptance already consumed part of p. Tokens the draft over-proposes were sometimes accepted, so the remaining probability that still needs delivering is p - q where that is positive, and zero elsewhere:
residual = np.maximum(p - q, 0.0)
residual /= residual.sum()Delete the maximum and use plain p and the whole guarantee vanishes, silently. Nothing crashes. Your output distribution is now somewhere between the draft and the target, and no test will tell you.
Greedy is the easy special case
At temperature zero the rule collapses to string comparison:
# accept each draft token while it equals the target's argmax at that position
take = 0
for j, d in enumerate(draft):
if target_argmax[j] == d:
take += 1
else:
breakAccept the matching prefix, take the target's own token at the first mismatch, discard the rest. This is what most production implementations use, and it is exactly lossless because greedy decoding has no distribution to preserve.
Using it for real
# vLLM: a smaller model from the same family as the drafter
from vllm import LLM
llm = LLM(
model="Qwen/Qwen3-32B",
speculative_config={"method": "draft_model",
"model": "Qwen/Qwen3-0.6B",
"num_speculative_tokens": 5},
)# HuggingFace Transformers: assisted generation
out = model.generate(**inputs, assistant_model=small_model, max_new_tokens=64)No output blocks for these two. Both need model downloads, and the vLLM path needs a GPU, so any output printed here would be invented rather than observed. The vLLM speculative-decoding API has changed several times. speculative_config is the current shape and it takes a method key. Check your installed version's docs rather than a blog post.
Common mistakes
Mismatched tokenizers. Draft and target must share a vocabulary, or every id means something different. Use a smaller model from the same family. vLLM has added a cross-vocabulary mode, and at the time of writing it supports greedy draft sampling only.
Drafting too many tokens. The optimal num_speculative_tokens depends on the acceptance rate. Every rejected draft token is wasted work, and beyond roughly 3 to 8 the wasted drafts outweigh the saved passes on most workloads. Measure your acceptance rate first.
Expecting a gain under heavy load. Speculation spends spare arithmetic capacity. At batch size 1 there is plenty. At batch size 64 the GPU is already busy and speculation can make total throughput worse while still improving one user's latency. It is a latency optimisation, not a throughput one.
Forgetting the draft model's memory. It needs weights and a KV cache of its own, taken out of the same GPU. On a tight memory budget this can cost you more concurrency than the speed is worth.
Benchmarking on the wrong text. Acceptance is high on boilerplate and low on creative prose. A speed-up measured on repetitive prompts does not transfer.
Try it yourself
Set DRAFT = TARGET and re-run. Acceptance goes to 1.0 and tokens per pass reaches GAMMA + 1, the theoretical ceiling. Then set DRAFT to a uniform distribution and watch acceptance and speed-up collapse together. Plot tokens-per-pass against acceptance rate and you have drawn the curve that decides whether this technique is worth deploying for you.
What to learn next
- Medusa, EAGLE and self-speculation — removing the second model entirely.
- vLLM — where you actually switch this on.
- Knowledge distillation — how a good draft model gets trained in the first place.
Researcher — Mathematics and papers.
The algorithm
From Leviathan, Kalman and Matias, Fast Inference from Transformers via Speculative Decoding, ICML 2023 (arxiv.org/abs/2211.17192). And concurrently from Chen et al., Accelerating Large Language Model Decoding with Speculative Sampling, 2023 (arxiv.org/abs/2302.01318).
Let $p$ be the target distribution and $q$ the draft distribution at a position. Draw $x \sim q$ and accept with probability
$$ \min!\left(1, \frac{p(x)}{q(x)}\right) $$
On rejection, draw from the normalised residual
$$ p'(x) = \frac{\max\big(0,\ p(x) - q(x)\big)}{\sum_{x'} \max\big(0,\ p(x') - q(x')\big)} $$
Theorem. The token returned is distributed exactly as $p$.
Proof sketch. $P(\text{accept and return } x) = q(x)\min(1, p(x)/q(x)) = \min(q(x), p(x))$. The total acceptance probability is $\beta = \sum_x \min(p(x), q(x)) = 1 - D_{TV}(p, q)$. The rejection branch contributes $(1-\beta) \cdot p'(x) = \max(0, p(x) - q(x))$. Summing the two branches gives $\min(p(x), q(x)) + \max(0, p(x) - q(x)) = p(x)$ for every $x$. $\square$
The counting experiment above is a direct empirical check of this identity, including the measured acceptance rate landing on $1 - D_{TV}$.
Expected speed-up
With $\gamma$ draft tokens per round and per-token acceptance rate $\alpha$ assumed independent across positions, the expected number of tokens per target forward pass is
$$ \mathbb{E}[#\text{tokens}] = \frac{1 - \alpha^{\gamma + 1}}{1 - \alpha} $$
Let $c$ be the cost of one draft step relative to one target step. Wall-clock speed-up is
$$ S = \frac{1 - \alpha^{\gamma+1}}{(1 - \alpha)(1 + c\gamma)} $$
Differentiating in $\gamma$ gives the optimum. At $\alpha = 0.8$, $c = 0.05$ the optimum sits near $\gamma = 5$ to $7$. At $\alpha = 0.5$ it drops to $\gamma = 2$ to $3$. The independence assumption is optimistic — real acceptance is autocorrelated, since a draft that has gone off-track keeps being wrong — so treat the formula as an upper bound.
The $(1 + c\gamma)$ term is why the technique fails at high batch sizes in a way the token-count formula hides. Under load the target step is compute-bound, so verifying $\gamma+1$ positions is no longer nearly free, and $c$ effectively rises.
Draft selection
| Draft source | Acceptance | Cost | Notes |
|---|---|---|---|
| Small model, same family | 0.6 – 0.8 | needs its own weights and cache | The original proposal |
| N-gram / prompt lookup | 0.3 – 0.9, workload-dependent | free | Excellent when output copies input |
| Medusa-style heads | 0.6 – 0.7 | small | Trained on the target's own hidden states |
| EAGLE-family | 0.8+ | small | Current state of the art |
| Early-exit from the target's own layers | varies | none | Self-speculation, no second model |
All of these are covered in the next lesson. The design axis is the same throughout: raise $\alpha$ while keeping $c$ small.
Tree verification
Verifying one linear draft wastes the pass whenever the first token is rejected. Verify a tree of candidates in one pass instead, using an attention mask where each node attends only to its ancestors. That raises the expected accepted length for a modest increase in verified positions. Introduced for this setting by SpecInfer (Miao et al., 2023, arxiv.org/abs/2305.09781) and now standard in Medusa and EAGLE.
Practical notes that papers understate
Acceptance is not stationary. It varies enormously within a single generation — high inside a quoted document, low at the start of a novel sentence. Dynamic schemes that adjust $\gamma$ from a running acceptance estimate are worth more than tuning a constant.
Sampling parameters change $\alpha$. Temperature applied to the target sharpens or flattens $p$ relative to $q$ and moves acceptance in either direction. Reported speed-ups at temperature 0 do not transfer to temperature 1.
Speculation and structured output interact badly. Draft tokens must satisfy the grammar as well as the target, so a tight grammar constraint lowers acceptance.
It is a latency technique. For a throughput-bound service, larger batches beat speculation. For an interactive single-user or low-concurrency service, speculation is often the single largest available win.
Papers
- Leviathan et al., Fast Inference from Transformers via Speculative Decoding, ICML 2023 — arxiv.org/abs/2211.17192
- Chen et al., Accelerating Large Language Model Decoding with Speculative Sampling, 2023 — arxiv.org/abs/2302.01318
- Stern et al., Blockwise Parallel Decoding for Deep Autoregressive Models, 2018 — arxiv.org/abs/1811.03115
- Miao et al., SpecInfer: Accelerating Generative LLM Serving with Tree-based Speculative Inference, 2023 — arxiv.org/abs/2305.09781
What to learn next
- Medusa, EAGLE and self-speculation — removing the second model entirely.
- vLLM — where you actually switch this on.
- Knowledge distillation — how a good draft model gets trained in the first place.