How Text Is Generated

Medusa, EAGLE and self-speculation

Speculative decoding without a second model — the big model guesses ahead using extra heads, its own hidden states, or plain text it has already seen.

On this page 8
  1. Why anyone wanted this
  2. The three families
  3. How copying works
  4. Verifying several guesses at once
  5. The honest limits
  6. Where you have already seen this
  7. Remember this
  8. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The big model guesses ahead for itself, instead of hiring a second small model. It does that by copying text it has already seen, or with small extra parts bolted on.

Think about filling in a long government form where your address appears four times. The first time you write it out carefully. The next three times your hand almost moves on its own, because you are copying something already on the page.

You did not need help from anyone. The information was already in front of you.

That is the simplest form of self-speculation, and it is astonishingly effective on the kind of text people actually generate: quoting a document, editing code, filling a template.

Why anyone wanted this

The previous lesson used a second, smaller model as the guesser. That works, and it has three annoyances.

You need a small model that shares the big one's vocabulary, which does not always exist. You have to load it, so it eats graphics memory that would otherwise hold conversations. And you have to run it, which costs time on every step.

Self-speculation removes all three. The guesses come from the big model itself. Or from something attached to it, or from text already on the page.

The three families

Copy from what is already there. Look at the last few words written. Find where that exact phrase appeared earlier in the prompt or the answer. Guess that whatever followed it then will follow it now. No training, no extra model, no memory. Free.

Extra heads. Bolt a few small predictors onto the big model. Each is trained to guess one, two or three words ahead. They read the same internal state the model already computed. Tiny compared to a whole second model. This is what the Medusa method does.

Reuse the model's own thinking. Take the internal state the big model produced for the current word. Let one small layer roll it forward to guess the next few. This is the EAGLE family, and it is the current best-performing approach.

How copying works

   prompt:  "... leaves Nagpur at 6:15 in the morning and reaches
             Delhi at 9:40 at night. ... leaves Nagpur at"

   last three words written:  "leaves Nagpur at"
   found earlier in the text: yes
   what followed it there:    "6:15 in the morning"

   -> guess those four words, and check all four in one pass

Nothing was trained. The answer was sitting in the prompt.

Verifying several guesses at once

A guesser can offer more than one option: maybe the next word is "closed", maybe it is "shut".

Checking each possibility separately would be wasteful. The big model can instead check a whole branching tree of options in one pass. It does that by controlling which words are allowed to see which others. Shared beginnings are computed once.

This is why the newer methods can afford to propose several alternatives instead of one line of text.

The honest limits

Copying only helps when the answer repeats the input. Ask for an original story and there is nothing to copy. The technique then does nothing at all, and costs nothing.

The trained approaches help more broadly, and they need training. Someone has to build and publish those extra parts for your specific model. When they exist, they are excellent. When they do not, you wait.

Where you have already seen this

  • A coding assistant refactoring a file, moving unusually fast where the code repeats.
  • An agent tool replaying long, similar sequences quickly.
  • A model card mentioning an accompanying draft module.

Remember this

  • Self-speculation removes the second model, and with it the extra memory and vocabulary problem.
  • Copying from earlier text is free and works whenever output repeats input.
  • Trained approaches like Medusa and EAGLE guess better, and someone must train them for your model.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install torch transformers

Prompt lookup, the training-free member of the family, in about twenty lines against a real model. distilgpt2 is roughly 340 MB and runs on CPU. The download prints a progress bar not shown below.

Copying, verified against plain greedy

prompt_lookup.py
import torch
from transformers import AutoTokenizer, AutoModelForCausalLM

tok = AutoTokenizer.from_pretrained("distilgpt2")
model = AutoModelForCausalLM.from_pretrained("distilgpt2").eval()

PROMPT = ("The Rajdhani Express leaves Nagpur at 6:15 in the morning and reaches "
          "Delhi at 9:40 at night. The Rajdhani Express leaves Nagpur at")
NEW = 16


def argmax_next(ids):
    with torch.no_grad():
        return model(ids).logits[0, -1].argmax().item()


def plain_greedy(ids, n):
    passes = 0
    for _ in range(n):
        ids = torch.cat([ids, torch.tensor([[argmax_next(ids)]])], dim=1)
        passes += 1
    return ids, passes


def lookup_draft(seq, ngram=3, k=4):
    """Find the last `ngram` tokens earlier in the text and copy what followed."""
    if len(seq) < ngram + 1:
        return []
    tail = seq[-ngram:]
    for i in range(len(seq) - ngram - 1, -1, -1):
        if seq[i:i + ngram] == tail:
            return seq[i + ngram: i + ngram + k]
    return []


def speculative_greedy(ids, n, ngram=3, k=4):
    passes, accepted, proposed = 0, 0, 0
    seq = ids[0].tolist()
    while len(seq) - ids.shape[1] < n:
        draft = lookup_draft(seq, ngram, k)
        proposed += len(draft)
        cand = torch.tensor([seq + draft])
        with torch.no_grad():
            logits = model(cand).logits[0]          # ONE pass scores every draft position
        passes += 1
        start = len(seq) - 1                        # this position predicts draft[0]
        take = 0
        for j, d in enumerate(draft):
            if logits[start + j].argmax().item() == d:
                take += 1
            else:
                break
        accepted += take
        seq = seq + draft[:take]
        seq.append(logits[start + take].argmax().item())   # the free corrected token
    return torch.tensor([seq[: ids.shape[1] + n]]), passes, accepted, proposed


ids = tok(PROMPT, return_tensors="pt").input_ids
a, p1 = plain_greedy(ids, NEW)
b, p2, acc, prop = speculative_greedy(ids, NEW)

print("plain greedy text     :", repr(tok.decode(a[0, ids.shape[1]:])))
print("speculative text      :", repr(tok.decode(b[0, ids.shape[1]:])))
print("identical             :", torch.equal(a, b))
print()
print(f"tokens generated      : {NEW}")
print(f"forward passes, plain : {p1}")
print(f"forward passes, spec  : {p2}")
print(f"draft tokens proposed : {prop}, accepted {acc}"
      f"  ({acc / prop:.0%} acceptance)" if prop else "")
print(f"speed-up in passes    : {p1 / p2:.2f}x")
Output
plain greedy text     : ' 6:15 in the morning and reaches Delhi at 9:40 at night.'
speculative text      : ' 6:15 in the morning and reaches Delhi at 9:40 at night.'
identical             : True

tokens generated      : 16
forward passes, plain : 16
forward passes, spec  : 4
draft tokens proposed : 16, accepted 16  (100% acceptance)
speed-up in passes    : 4.00x

Run against transformers 5.6.2 and torch 2.5.1 on CPU. Greedy decoding of a fixed prompt is deterministic, so this reproduces.

Reading the output

identical: True, and that is non-negotiable. Sixteen tokens, byte for byte the same as plain greedy. Every draft token was checked by the full model and kept only when it matched what the full model would have chosen. Speculation that changes the output is a bug, not a trade-off.

Sixteen passes became four. Four times fewer calls to the expensive model, with no training, no second model and no extra memory. The entire drafter is lookup_draft, which is a substring search.

One hundred percent acceptance, because the prompt contained the answer. The prompt states the timetable, then repeats the opening phrase. Every guess was a correct copy. This is not a rigged example. It is the shape of summarisation, question answering over a document, code refactoring, and agent traces that quote tool output.

Change the prompt to something original and this drops to nothing. lookup_draft returns an empty list when the trigram has not appeared before. The loop falls back to one token per pass, and you have plain greedy with a negligible search overhead. The floor is the old behaviour, which is what makes this safe to leave on.

start = len(seq) - 1 is the line to check twice. Position i predicts token i+1. The last real token's row predicts the first draft token, the first draft token's row predicts the second, and so on. Getting this index wrong produces a working program that silently accepts nothing.

Turning it on for real

python
# HuggingFace: the same idea, built in
out = model.generate(**inputs, prompt_lookup_num_tokens=10, max_new_tokens=100)
python
# vLLM
llm = LLM(model="...", speculative_config={"method": "ngram",
                                           "num_speculative_tokens": 4,
                                           "prompt_lookup_min": 2,
                                           "prompt_lookup_max": 5})

No output blocks: both need a model download, and vLLM needs a GPU. Printing invented output here would teach you to expect something that will not happen.

Verifying a tree instead of a line

The trained methods propose several alternatives at once. Checking each separately would waste the saving, so they pack a tree into one pass using a hand-built attention mask.

tree_mask.py
import numpy as np

# 3 context tokens, then a draft TREE: two guesses for the next token,
# and two guesses for the one after each of those.
NAMES  = ["c0", "c1", "c2", "A", "B", "A>X", "A>Y", "B>Z", "B>W"]
PARENT = [None, 0, 1, 2, 2, 3, 3, 4, 4]        # index of each node's parent

n = len(NAMES)
mask = np.zeros((n, n), dtype=int)
for i in range(n):
    j = i
    while j is not None:                              # a node sees itself and its ancestors
        mask[i, j] = 1
        j = PARENT[j]

print("tree attention mask (1 = may attend)")
print("        " + " ".join(f"{x:>4}" for x in NAMES))
for i, row in enumerate(mask):
    print(f"{NAMES[i]:>6}  " + " ".join(f"{v:>4}" for v in row))

paths = [("A", "A>X"), ("A", "A>Y"), ("B", "B>Z"), ("B", "B>W")]
print(f"\ncandidate continuations packed into this one pass: {len(paths)}")
print(f"positions in the tree pass          : {n}")
print(f"positions if run as separate passes : {len(paths)} x {3 + 2} = {len(paths) * 5}")
print(f"draft tokens stored once instead of {len(paths)}: 'A' and 'B'")
Output
tree attention mask (1 = may attend)
          c0   c1   c2    A    B  A>X  A>Y  B>Z  B>W
    c0     1    0    0    0    0    0    0    0    0
    c1     1    1    0    0    0    0    0    0    0
    c2     1    1    1    0    0    0    0    0    0
     A     1    1    1    1    0    0    0    0    0
     B     1    1    1    0    1    0    0    0    0
   A>X     1    1    1    1    0    1    0    0    0
   A>Y     1    1    1    1    0    0    1    0    0
   B>Z     1    1    1    0    1    0    0    1    0
   B>W     1    1    1    0    1    0    0    0    1

candidate continuations packed into this one pass: 4
positions in the tree pass          : 9
positions if run as separate passes : 4 x 5 = 20
draft tokens stored once instead of 4: 'A' and 'B'

Look at the A>X row. It attends to the three context tokens, to A, and to itself — and it cannot see B at all. Each candidate path sees exactly its own history and nothing from a sibling branch. Nine positions cover four candidate continuations, against twenty positions if each were run alone.

This mask is the only structural difference between linear and tree speculation. Everything else is bookkeeping about which path was accepted and how to trim the KV cache afterwards.

Choosing a method

MethodTraining neededExtra memoryBest on
Prompt lookup / n-gramnonenoneOutput that quotes input: summaries, edits, RAG, agents
Suffix decodingnonea suffix tree over historyRepetitive agent traces
Medusa headsfine-tune headssmallGeneral text, when heads exist for your model
EAGLE familytrain a small draftersmallGeneral text; the current best acceptance rates
Layer-skip self-draftingnone, or light tuningnoneWhen no drafter exists for your model

The decision procedure is short. Try prompt lookup first, because it costs nothing and is one flag. If your workload has no repetition, look for a published EAGLE-3 drafter for your exact model. If none exists, you are choosing between training one and doing without.

Common mistakes

Assuming a drafter transfers between models. Medusa heads and EAGLE drafters are trained against one specific checkpoint's hidden states. A drafter for the base model will not work on your fine-tune, and it will fail by accepting almost nothing rather than by crashing.

Setting the n-gram length to 1 or 2. Short triggers match everywhere and propose nonsense, so the acceptance rate collapses and you pay for wasted verification. Three to five is the usual range.

Measuring speed-up on the wrong prompts. Prompt lookup on original prose gives exactly zero benefit. Measure on your real traffic.

Leaving speculation on under heavy load. It spends spare arithmetic capacity, and at large batch sizes there is none. Several serving stacks now disable speculation dynamically above a concurrency threshold, and if yours does not, you should.

Try it yourself

Change ngram from 3 to 1 and re-run. Acceptance falls and passes rise, because a single-token trigger matches text that means something else. Then remove the second half of PROMPT so nothing repeats, and confirm the pass count returns to 16 with no error. Degrading to the old speed rather than breaking is the property that makes this deployable.

What to learn next

Researcher — Mathematics and papers.

The lineage

Blockwise parallel decoding (Stern et al., 2018, arxiv.org/abs/1811.03115) is the ancestor. Train $k$ auxiliary output heads to predict positions $t+1 \dots t+k$ from the same final hidden state, then verify the block in one pass. Everything below is a refinement.

Medusa (Cai et al., 2024, arxiv.org/abs/2401.10774) revives this for LLMs. Head $i$ is a residual block over the last hidden state $h_t$:

$$ p_i(\cdot) = \mathrm{softmax}\Big( W^{(2)}_i \big( \mathrm{SiLU}(W^{(1)}_i h_t) + h_t \big) \Big) $$

Top-$s_i$ candidates from each head are combined into a Cartesian tree and verified in one pass under a tree mask. Two training regimes: Medusa-1 freezes the backbone, giving lossless acceleration; Medusa-2 trains heads and backbone together for higher acceptance at the cost of changing the base model.

Medusa also introduced typical acceptance for temperature sampling. Accept a draft token when its target probability exceeds a threshold that scales with the step's entropy, rather than running exact rejection sampling. This is not distribution-preserving. Reported quality is unaffected on benchmarks, and the strict guarantee is gone. Know which one your stack uses before claiming losslessness.

The structural weakness is that all heads condition on the same $h_t$, so head $i$ cannot see what head $i-1$ chose. Hydra (2024, arxiv.org/abs/2402.05109) makes the heads sequentially dependent and improves acceptance for that reason.

EAGLE

EAGLE (Li et al., 2024, arxiv.org/abs/2401.15077) makes two observations. Autoregression at the feature level is easier than at the token level. And feature-level drafting is ambiguous unless you also condition on the token actually sampled one step ahead.

The drafter is a single transformer layer over the concatenation of second-top-layer features and shifted token embeddings. It predicts the next feature, which the target's own LM head turns into a distribution. Reported around $3\times$ over vanilla decoding on MT-bench, and $1.6\times$ over Medusa.

EAGLE-2 replaces the static draft tree with a dynamic one, expanding branches by the drafter's own confidence, on the finding that acceptance rates vary strongly by position and context.

EAGLE-3 (2025, arxiv.org/abs/2503.01840) drops feature prediction entirely in favour of direct token prediction. It also fuses low, middle and high layer features from the target, rather than using only the second-top layer. Its central claim is a scaling law: EAGLE's feature-prediction constraint capped the benefit of more training data, and removing it makes acceptance improve with data the way everything else in this field does. The training-time test procedure simulates multi-step drafting during training so the drafter learns to handle its own errors.

EAGLE 3.1 (2026) addresses "attention drift": the drafter's behaviour degrading over successive draft steps within one round. It normalises each target hidden state and feeds post-normalisation states into the next drafting step. The drafter then behaves more like a proper recursive application of itself. The vLLM team reports up to twice the acceptance length on long-context workloads against EAGLE-3, about $2.03\times$ per-user output throughput at concurrency 1, and $1.66\times$ at concurrency 16. Existing EAGLE-3 checkpoints remain compatible.

The concurrency-16 figure is the one to notice. Most speculative-decoding results are reported at batch size 1, where spare capacity is abundant. Holding a meaningful gain at realistic concurrency is the harder claim.

Model-free methods

Prompt lookup decoding (Saxena, 2023) has no paper and is a few dozen lines: match the last $n$ generated tokens against the context, copy the continuation. Available as prompt_lookup_num_tokens in transformers and method: "ngram" in vLLM.

SuffixDecoding (Oliaro et al., 2024, arxiv.org/abs/2411.04975, NeurIPS 2025 spotlight) generalises this properly. It maintains suffix trees over prompts and previous outputs, and adapts speculation length to the estimated acceptance likelihood. On agentic benchmarks including SWE-Bench and text-to-SQL it reports up to $5.3\times$, beating EAGLE-2 and EAGLE-3 on those workloads.

That result deserves a moment. A method with no neural drafter at all beats trained state-of-the-art drafters on repetitive agent traffic. Draft quality is workload-specific, and "state of the art" without naming the workload means very little in this area.

Self-drafting from the target's own layers

Draft & Verify (Zhang et al., 2023, arxiv.org/abs/2309.08168) drafts by running the target model with some intermediate layers skipped. The full model then verifies. No extra parameters, no extra memory, no training — reported up to $1.99\times$ on Llama-2. Which layers to skip is chosen by a Bayesian optimisation pass.

LayerSkip (Elhoushi et al., ACL 2024, arxiv.org/abs/2404.16710) trains for this explicitly. Layer dropout and an early-exit loss make early exits usable, and the KV cache is shared between the drafting and verification passes.

Multi-token prediction (Gloeckle et al., 2024, arxiv.org/abs/2404.19737) trains the model from the start with several output heads, improving both quality and self-speculation. DeepSeek-V3 ships MTP modules for exactly this, and vLLM exposes them as method: "mtp".

What to actually measure

Speed-up is not one number. Report:

  • Acceptance rate $\alpha$, per position, not averaged over positions.
  • Mean accepted length per verification pass, the quantity that maps to wall clock.
  • Drafting overhead $c$ as a fraction of a target step.
  • Concurrency, always. A speed-up at batch 1 says nothing about batch 32.
  • Output equality against the non-speculative baseline, for greedy. If it is not bit-identical, the implementation has a bug or is using a relaxed acceptance rule.

Papers

What to learn next