Cross-attention
Queries come from one sequence, keys and values from another, which is how a model reads a picture while writing a caption or reads Hindi while writing English.
- 12 min read
- 3 reading levels
- Updated
Read these first
On this page 7
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Cross-attention lets one sequence read a completely different sequence while it works.
Think about copying notes from the blackboard. You write one word in your notebook. You glance up at the board, find your place, write the next word. Then glance up again.
Two separate things are in play. The board holds information you did not write and cannot change. Your notebook is what you are building, one word at a time. Your eyes keep travelling from one to the other.
Cross-attention is that glance. The questions come from what you are writing. The answers come from the board.
How it differs from ordinary attention
Everything so far in this section has been self-attention — a sentence looking at itself. One sequence, playing all three roles.
Cross-attention splits the roles across two sequences.
- The queries come from the sequence being built.
- The keys and values come from the sequence being read.
Nothing else about the mechanism changes. Same scoring, same sharing out, same mixing. Only the source of the inputs moved.
Why it exists
Translation was the original problem. You have a Hindi sentence and you want an English one.
The two are not the same length. The word order is different. One English word can correspond to two Hindi words, and vice versa.
Older systems squashed the whole Hindi sentence into one fixed summary, then unrolled English out of it. Short sentences worked. Long ones fell apart, because everything had to survive that single squashed summary.
Cross-attention removed the squash. The English side keeps the whole Hindi sentence available and looks up whichever part it needs, word by word.
THE BOARD (read only) YOUR NOTEBOOK (being written)
main chai peeta hoon I drink tea
^ ^ ^ ^
| | | |
+-----+-----+------------------------+
"tea" asks: which word here means a drink?
answer: mostly "chai"The shape gives it away
Self-attention produces a square grid. Four words looking at four words is a four-by-four grid.
Cross-attention produces a rectangle. Three English words looking at four Hindi words is three rows and four columns. If you ever see a rectangular attention picture, you are looking at cross-attention.
Why nothing needs masking here
Causal masking exists so a word cannot see its own answer. On the board, there is no answer to peek at — the Hindi sentence is the question, not the answer.
So the English side still masks itself, because it is being written left to right. But when it looks at the board, it may look anywhere. Beginning, middle, end. All of it is fair.
Where you have already seen this
- Google Translate producing a sentence while keeping the original available.
- Automatic subtitles, where the words being written read from the sound that was heard.
- Image captioning, where the caption reads parts of the picture instead of parts of a sentence.
- A chatbot answering about a PDF you uploaded. Some designs use cross-attention instead of pasting the text into the prompt.
Remember this
- Queries come from what is being written; keys and values come from what is being read.
- The attention picture is a rectangle, not a square.
- The read-only side needs no causal mask, because there is nothing there to cheat with.
What to learn next
- Why attention costs grow with the square of length — where the bill comes from.
- Vision-language models — cross-attention joining pictures to words.
- Whisper — cross-attention joining sound to words.
Developer — Code and libraries.
Setup
pip install numpyA rectangular alignment, hand-built
import numpy as np
np.set_printoptions(precision=3, suppress=True)
def softmax(x, axis=-1):
x = x - x.max(axis=axis, keepdims=True)
e = np.exp(x)
return e / e.sum(axis=axis, keepdims=True)
# Source side: a Hindi sentence the encoder has already read.
src = ["main", "chai", "peeta", "hoon", "<eos>"]
# Slot meanings: [speaker, beverage, action]
K_src = np.array([[1., 0., 0.],
[0., 1., 0.],
[0., 0., 1.],
[0., 0., 0.3],
[0., 0., 0.]])
V_src = K_src.copy()
# Target side: what the decoder has produced so far.
tgt = ["I", "drink", "tea"]
Q_tgt = np.array([[4., 0., 0.], # "I" is hunting for the speaker
[0., 0., 4.], # "drink" is hunting for the action
[0., 4., 0.]]) # "tea" is hunting for the beverage
logits = Q_tgt @ K_src.T / np.sqrt(K_src.shape[-1])
A = softmax(logits)
ctx = A @ V_src
print("query comes from the target, key and value come from the source")
print("Q_tgt ", Q_tgt.shape, " K_src", K_src.shape, " A", A.shape, " ctx", ctx.shape)
print("\nalignment (row = English word, column = Hindi word it read):")
print(" " + "".join(f"{w:>9s}" for w in src))
for w, row in zip(tgt, A):
print(f"{w:>8s} " + "".join(f"{v:9.3f}" for v in row))
print("\neach row still sums to 1:", A.sum(axis=1))
print("\nnote the shape: A is 3x5, not square.")
print("self-attention would give", f"{len(tgt)}x{len(tgt)}", "on the target side.")
# The encoder is read once; the decoder queries it on every step.
print("\ncost of decoding one more target token:")
for made in (1, 2, 3, 10):
print(f" {made:>2} tokens produced -> {made} new query rows x {len(src)} cached keys "
f"= {made * len(src)} scores; K_src and V_src are never recomputed")query comes from the target, key and value come from the source
Q_tgt (3, 3) K_src (5, 3) A (3, 5) ctx (3, 3)
alignment (row = English word, column = Hindi word it read):
main chai peeta hoon <eos>
I 0.716 0.071 0.071 0.071 0.071
drink 0.066 0.066 0.668 0.133 0.066
tea 0.071 0.716 0.071 0.071 0.071
each row still sums to 1: [1. 1. 1.]
note the shape: A is 3x5, not square.
self-attention would give 3x3 on the target side.
cost of decoding one more target token:
1 tokens produced -> 1 new query rows x 5 cached keys = 5 scores; K_src and V_src are never recomputed
2 tokens produced -> 2 new query rows x 5 cached keys = 10 scores; K_src and V_src are never recomputed
3 tokens produced -> 3 new query rows x 5 cached keys = 15 scores; K_src and V_src are never recomputed
10 tokens produced -> 10 new query rows x 5 cached keys = 50 scores; K_src and V_src are never recomputedWhat to take from the alignment
The three-by-five shape is the signature. Three target rows, five source columns. A @ V_src gives back three rows, so the target keeps its own length. Cross-attention never changes how long the output is.
"drink" spread across "peeta" and "hoon". It put 0.668 on "peeta" and 0.133 on "hoon". In Hindi the verb is split across those two words, and a soft alignment can represent that. A hard one-to-one alignment could not. This is why attention beat the alignment tables that came before it.
Rows sum to one; columns do not. Nothing forces a source word to be used, and nothing stops one being used by every target word. Source words that no target word reads are a real failure mode in translation. It shows up as dropped content in the output.
The cost line is the reason encoders are worth having. K_src and V_src are computed once for the whole source and cached. Every decoding step reuses them. If you re-encode the source on every step, you have thrown away the main advantage of the design.
In PyTorch
Cross-attention is the same function with different inputs:
import torch, torch.nn as nn, torch.nn.functional as F
d, H = 32, 4
q_proj, k_proj, v_proj = (nn.Linear(d, d, bias=False) for _ in range(3))
tgt = torch.randn(1, 3, d) # being written
mem = torch.randn(1, 7, d) # already read, held fixed
def heads(t):
B, T, _ = t.shape
return t.view(B, T, H, d // H).transpose(1, 2)
q, k, v = heads(q_proj(tgt)), heads(k_proj(mem)), heads(v_proj(mem))
out = F.scaled_dot_product_attention(q, k, v) # no is_causal: the memory is fully visible
print(q.shape, k.shape, out.shape)torch.Size([1, 4, 3, 8]) torch.Size([1, 4, 7, 8]) torch.Size([1, 4, 3, 8])
Query has 3 tokens, key has 7, and the output has 3. That is the target's length, carrying information from the memory. nn.MultiheadAttention(query, key, value) accepts the same split. That is why its call signature takes three tensors rather than one.
Where a cross-attention block sits
A classic encoder-decoder decoder layer has three sublayers in this order:
- Causal self-attention over the target produced so far.
- Cross-attention into the encoder output.
- A feedforward layer.
Each is wrapped in a residual connection and a normalisation step. That is the pattern in walking through one transformer block.
Cross-attention outside translation
The mechanism does not care that both sides are text.
- Whisper (speech to text) cross-attends from text tokens into encoded audio frames. See Whisper.
- Stable Diffusion cross-attends from image patches into text embeddings, which is how a prompt steers the picture. See diffusion models.
- Flamingo-style vision-language models insert cross-attention layers into a frozen language model. It can then read image features. See vision-language models.
In all three, the source and target are different kinds of data entirely. The only requirement is that keys and values end up in the same width as the queries.
Common mistakes
Applying a causal mask to the memory. The memory is fully available. Masking it blinds the decoder to most of the source and produces fluent output that ignores the input.
Forgetting the memory's padding mask. Batched sources have padding, and padded frames must be blocked. Otherwise the model attends to filler and quality drops without any error.
Recomputing the encoder every decoding step. Run the encoder once. Cache K_src and V_src per layer. Some frameworks do this for you; verify rather than assume.
Mismatched widths. Queries come from the target's width, keys and values from the memory's. If the two towers differ in width, the key and value projections must map into the query width. This is the single most common bug when bolting a vision encoder onto a language model.
Try it yourself
Make the source longer than the target and then shorter, and print A.shape each time. Confirm the output row count always follows the target.
Then set Q_tgt[2] = [0., 0., 0.] so "tea" asks for nothing. The row becomes uniform, one over five everywhere. A query with no preference gets the average of the whole source. That is exactly the squashed summary cross-attention was invented to avoid.
What to learn next
- Why attention costs grow with the square of length — where the bill comes from.
- Vision-language models — cross-attention joining pictures to words.
- Whisper — cross-attention joining sound to words.
Researcher — Mathematics and papers.
Definition
With target states $Y \in \mathbb{R}^{T_y \times d}$ and encoder memory $Z \in \mathbb{R}^{T_z \times d_z}$:
$$ \operatorname{CrossAttn}(Y, Z) = \operatorname{softmax}!\left( \frac{(Y W_Q)(Z W_K)^\top}{\sqrt{d_k}} \right) Z W_V $$
The attention matrix is $T_y \times T_z$. Setting $Z = Y$ recovers self-attention exactly, so self-attention is the special case, not the general one.
Lineage
Cross-attention is the original attention. Bahdanau et al. (2015), arXiv:1409.0473, introduced it to fix the fixed-length bottleneck in translation. Scoring used a small feedforward network. Luong et al. (2015), arXiv:1508.04025, replaced that with a dot product and catalogued global against local variants. Vaswani et al. (2017) kept cross-attention and added self-attention alongside it. The paper's title refers to removing recurrence, not to inventing attention.
Self-attention is the later idea. Textbook orderings that present it first invert the history.
Cost and caching
Per decoding step $t$, per layer:
- Cross-attention: $O(T_z d)$. Constant in $t$, since the memory does not grow.
- Causal self-attention: $O(t d)$. Grows with what has been generated.
The encoder KV cache is written once, size $2 L T_z d_k$ bytes, and read every step. In a decoder-only model that content lives in the same growing cache as the generated tokens. The two designs converge in cost when $T_z \gg T_y$. The real difference is architectural separation, not arithmetic — a point made concrete in decoder-only vs encoder-decoder.
Alignment quality is not free
Attention weights are frequently displayed as word alignments. Two papers together establish the careful position: Jain and Wallace (2019) and Wiegreffe and Pinter (2019). Attention weights are one of many distributions that produce the same output. They are evidence about information flow, not a faithful explanation of the decision. For translation, cross-attention correlates well with human alignments in some heads and poorly in others. The alignment-extraction literature selects or supervises specific heads rather than averaging them.
Cross-attention as a modality bridge
Three designs dominate for connecting a non-text encoder to a language model:
- Interleaved cross-attention layers. Flamingo (Alayrac et al., 2022, arXiv:2204.14198) freezes both towers and inserts gated cross-attention layers. The gate is initialised to zero, so the language model is unchanged at step zero.
- Prefix projection. BLIP-2 (Li et al., 2023, arXiv:2301.12597) and the LLaVA family take a different route. Encoder features are projected into the token embedding space and prepended. The language model then uses ordinary self-attention. Simpler, and it consumes context length.
- Perceiver resampling. A fixed set of learned latent queries cross-attends into a variable-length input. The result is a constant-length summary (Jaegle et al., 2021, arXiv:2103.03206). This decouples cost from input length and is how many systems handle video.
The perceiver pattern is worth noting for a structural reason. The queries carry no information about the input at all. They are learned constants. Everything content-dependent arrives through the keys and values.
Papers
- Bahdanau et al., Neural Machine Translation by Jointly Learning to Align and Translate, 2015 — arxiv.org/abs/1409.0473
- Luong et al., Effective Approaches to Attention-based NMT, 2015 — arxiv.org/abs/1508.04025
- Jaegle et al., Perceiver, 2021 — arxiv.org/abs/2103.03206
- Alayrac et al., Flamingo, 2022 — arxiv.org/abs/2204.14198
- Jain and Wallace, Attention is not Explanation, 2019 — arxiv.org/abs/1902.10186
- Wiegreffe and Pinter, Attention is not not Explanation, 2019 — arxiv.org/abs/1908.04626
What to learn next
- Why attention costs grow with the square of length — where the bill comes from.
- Vision-language models — cross-attention joining pictures to words.
- Whisper — cross-attention joining sound to words.