Looking Inside a Trained Model
Induction heads
Induction heads let a model complete a repeated pattern it has only seen once before in the same conversation, the mechanism behind in-context learning.
- 8 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Induction heads let a model finish a pattern it saw once already, earlier in the same text.
Someone reads you the start of a song's chorus. You have heard it once already, earlier in the song. Your brain completes the rest before they finish the line.
You did not memorise that chorus in school. You spotted the repeat, seconds ago, and copied it forward. A language model has a specific part built for exactly this trick, called an induction head.
Why it exists
A model meets brand-new information constantly. A made-up name, a password, an unfamiliar code, none of it in its training data. It still needs to use that information correctly, later in the same conversation.
Plain memorised facts cannot help here. The made-up name was never in the training data to memorise. The model needs a different skill: noticing a repeat, right now, and copying what came after it last time.
Induction heads are the mechanism that does this. They were discovered by researchers studying how models handle exactly this kind of never-seen-before pattern.
How it works
A completely random, made-up sequence of symbols:
[see it once] K7 Q2 X9 M4 P1 ...
[see it again] K7 Q2 X9 M4 P1 ...
^
after seeing "M4" a second time,
the model predicts "P1" next,
because that is what followed it last timeAn induction head watches for the current word to match something seen earlier. On a match, it checks whatever word came right after that earlier occurrence. It copies that word forward as the prediction.
Where you have already seen it
- A chatbot repeating your own words back correctly. Give it an unusual made-up term, use it again later. It reuses your exact term, not a similar-sounding real word.
- Code autocomplete. Suggesting the same made-up variable name you typed a few lines up, spelled exactly as you spelled it.
- Few-shot prompting. Show a model two or three examples of a pattern. It continues that exact pattern on a new example.
Remember this
- Induction heads let a model copy a pattern it saw once already, earlier in the same text.
- This is how a model handles brand-new information, like a made-up name, that it never memorised.
- It is one of the mechanisms behind "few-shot" prompting, learning a pattern from only a couple of examples.
What to learn next
- Activation patching — testing exactly which part of a model is doing this copying.
- Multi-head attention — the mechanism induction heads are built out of.
- The logit lens — watching a prediction sharpen, layer by layer, as you did here.
Developer — Code and libraries.
The clearest test for induction behaviour needs no real language at all: random tokens, repeated once. If the model gets much better at predicting the second copy, that is induction at work.
Setup
pip install transformers torchThe first run downloads gpt2, roughly 500 MB.
Measuring the "seen it before" advantage
import torch
from transformers import GPT2LMHeadModel, GPT2Tokenizer
torch.manual_seed(0)
tok = GPT2Tokenizer.from_pretrained("gpt2")
model = GPT2LMHeadModel.from_pretrained("gpt2")
model.eval()
# 30 random token IDs, repeated once. The model has never seen this exact
# sequence before, so it cannot rely on memorised facts.
random_ids = torch.randint(1000, 20000, (30,))
full = torch.cat([random_ids, random_ids]).unsqueeze(0)
with torch.no_grad():
logits = model(full, labels=full).logits[0]
losses = torch.nn.functional.cross_entropy(logits[:-1], full[0, 1:], reduction="none")
first_copy = losses[:29] # predicting inside the never-seen-before copy
second_copy = losses[29:] # predicting inside the already-seen copy
print(f"mean loss, first copy (never seen before): {first_copy.mean():.2f}")
print(f"mean loss, second copy (seen once already): {second_copy.mean():.2f}")mean loss, first copy (never seen before): 12.38 mean loss, second copy (seen once already): 0.74
Loss measures surprise: lower means the model was more confident, and correct. On the first copy, the model is guessing blind, since the tokens are random and unfamiliar. On the second copy, loss falls by more than 16 times. The model recognised the repeat and started copying.
Line by line
torch.randint(1000, 20000, (30,)) picks 30 random real GPT-2 token IDs. Using real vocabulary IDs, rather than made-up numbers, keeps the test fair to how the model actually operates on tokens.
labels=full makes the model compute cross-entropy loss against the input itself shifted by one position, the standard "predict the next token" training objective, reused here purely for measurement.
Slicing losses[:29] and losses[29:] splits the 59 prediction positions into "inside the first copy" and "inside the second copy", so the two halves can be compared directly.
Common mistakes
Using real words instead of random tokens for this test. Real words let the model succeed through memorised language statistics alone, muddying whether you are actually measuring induction or general fluency.
Testing with only one random sequence. A single run can land on an easy or hard sequence by chance. Run several different seeds and average, before drawing a conclusion.
Confusing this behavioural test with locating the responsible attention head. This experiment proves induction-like behaviour exists somewhere in the model. It does not, by itself, say which layer or head is doing it.
Try it yourself
Change random_ids to only 5 tokens instead of 30, and rerun. Compare the loss drop to the original 30-token version.
A shorter repeated pattern is an easier copying task, so expect the improvement to look different. Try a few lengths and see how the effect changes as the pattern gets longer.
What to learn next
- Activation patching — pinpointing which layer's computation is responsible for this drop in loss.
- Multi-head attention — the query-key-value mechanism an induction head is built from.
- Visualising attention maps — looking directly at which positions a head attends to.
Researcher — Mathematics and papers.
The behavioural signature
The standard diagnostic, following Olsson et al. (2022), measures in-context learning score: the drop in negative log-likelihood between early-sequence and late-sequence token predictions, on random or shuffled data repeated within the context:
induction_score = loss(early tokens) - loss(late tokens)A large positive score indicates the model is using earlier context to predict later tokens far better than it could from unigram statistics alone, exactly the pattern reproduced in the developer block.
The mechanistic circuit
Olsson et al. (2022), and the earlier Elhage et al. (2021) framework it builds on, describe induction as a two-head circuit, composed via the residual stream rather than existing as one head's isolated behaviour:
- A previous-token head, in an earlier layer, writes information about "what token preceded me" into each position's residual stream.
- An induction head, in a later layer, reads that information via K-composition. Its key vector is built partly from the previous-token head's output. It effectively searches for "a position whose previous token matches my current token", then copies forward whatever came after that earlier match.
This two-step structure is why induction heads count as a clean example of a genuine multi-component circuit, rather than a single attention pattern. It is also why they became an early success case for mechanistic interpretability as a field.
Why this specific circuit matters
Olsson et al. (2022) report a striking correlation. The training step where induction heads emerge, in small attention-only models, coincides closely with a visible bend in the loss curve. It also coincides with a sharp, independently measured jump in in-context learning ability.
The authors observed this pattern across the model sizes and architectures they tested. They present it as evidence that induction heads are a major contributor to in-context learning broadly, not only to literal token-copying, extending informally to analogical and pattern-based few-shot behaviour.
Complexity
Locating induction heads exhaustively requires computing the induction score per head: O(n_layers * n_heads) forward passes, or one pass with per-head ablation, over a suite of synthetic repeated-sequence probes. This is cheap relative to training, and is typically a one-time offline analysis per model checkpoint.
Key references
- Elhage, N. et al. (2021). A Mathematical Framework for Transformer Circuits. Anthropic.
- Olsson, C. et al. (2022). In-context Learning and Induction Heads. arXiv:2209.11895
Current state and open problems
Induction heads are among the best-understood circuits in transformer interpretability, largely because the behavioural test is simple, synthetic, and produces a clean, reproducible signal, exactly as in the developer block above.
What remains less settled is how far the story generalises. Real induction heads in large, real-world-trained models are often messier than in the small toy models the circuit was first described in. They are sometimes split across several heads, or entangled with other behaviours the same head performs. Whether "induction-like" behaviour in large models is genuinely one mechanism, or a family of related ones, is an open question current activation-patching studies are still working through.
What to learn next
- Activation patching — the causal-intervention technique used to confirm a circuit like this one.
- Probing hidden states — a complementary technique for testing what information a layer carries.
- Visualising attention maps — inspecting the previous-token and induction attention patterns directly.