Attention Mechanics

Queries, keys and values

Attention splits every word into three roles - what it is looking for, what it advertises about itself, and what it hands over when chosen.

On this page 7
  1. Why this had to be split into three
  2. Why three roles and not one
  3. How it works
  4. The part worth sitting with
  5. Where you have already seen this
  6. Remember this
  7. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Every word in a sentence plays three roles at once. It asks a question, it advertises what it is, and it carries something to give away.

Picture the spice rack in your kitchen. Each jar has a label stuck on the front, and something quite different inside it. You walk up thinking "I need something sour today". You read the labels. You do not open every jar. Then you take the powder from the one that matched.

Three different things happened there. Your need was one thing. The labels were another. What came out of the jar was a third.

Attention gives every word all three. The need is called the query, meaning the thing this word is looking for. The label is called the key, meaning what this word advertises to others. The contents are called the value, meaning what this word actually contributes when it is chosen.

Why this had to be split into three

Read this sentence: "the cat drank the milk because it was thirsty".

What does "it" mean? You knew instantly. It is the cat, because cats get thirsty and milk does not.

Now change one word: "the cat drank the milk because it was sweet". Suddenly "it" is the milk. The sentence barely moved, and the answer flipped completely.

Older language models read words in order and kept one running summary of everything so far. By the time they reached "it", the details of "cat" and "milk" had been squashed into that summary. They had no way to go back and look.

Attention removed that limit. Every word can look at every other word directly, no matter how far apart they sit.

Why three roles and not one

You might ask why a word cannot use the same description for all three jobs. It is a fair question, and the spice rack answers it.

The label on the jar is short and made for scanning. The contents are the actual useful thing. Your need is a third thing again, and it is not a label or a spice.

Words work the same way. To a search, "cat" advertises something short: "I am a living creature". What it hands over when picked is different: "a small furry animal, the subject here". And what "cat" is itself looking for is different again.

Keeping them separate lets one word match many different questions. It still contributes something specific each time.

How it works

   "it"  ---> makes a QUERY:  "I need something alive"
                    |
                    v
   every word offers a KEY (its label):
      the      -> "filler word"
      cat      -> "alive"
      drank    -> "an action"
      milk     -> "drinkable"
                    |
                    v
   compare the query against every key
   the closest match wins most of the weight
                    |
                    v
   collect the VALUES, weighted by how well each key matched
                    |
                    v
   "it" now carries mostly what "cat" contributed

Nothing here is hand-written. The model works out what makes a good query, a good key and a good value entirely from examples.

The part worth sitting with

The word "it" does not get replaced by "cat". Nothing is copied or substituted.

Instead, "it" ends up carrying a blend, and most of that blend came from "cat". A little came from every other word too. This blending is what the whole mechanism does, over and over, at every layer.

Where you have already seen this

  • Google Translate getting "bank" right depending on whether the sentence mentions a river or money.
  • ChatGPT answering a question about paragraph two of a long document you pasted.
  • Autocorrect on your phone choosing between "there" and "their".
  • Search inside your email finding a message from the words you half-remember.

Remember this

  • Each word makes a query, a key and a value.
  • The query asks, the key advertises, and the value is what gets handed over.
  • Queries are matched against keys. Values are what actually gets collected.
  • Splitting the three lets one word answer many different questions in different ways.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install numpy

Nothing else is needed. The whole idea fits in NumPy, and building it there first makes the PyTorch version read like a shorthand.

Query, key and value, hand-written so you can check every number

Real embeddings are hundreds of numbers wide and impossible to read. So this example uses three slots per word, with meanings we choose ourselves.

qkv.py
import numpy as np

np.set_printoptions(precision=3, suppress=True)

# One sentence. Each word gets three hand-written numbers, so you can check
# the arithmetic yourself. Slot meanings: [alive, drinkable, filler-word]
words = ["the", "cat", "drank", "the", "milk", "because", "it", "was"]
E = np.array([
    [0., 0., 1.],   # the       - a filler word
    [1., 0., 0.],   # cat       - alive
    [0., 0., 0.],   # drank     - an action
    [0., 0., 1.],   # the
    [0., 1., 0.],   # milk      - drinkable
    [0., 0., 1.],   # because
    [0., 0., 0.],   # it
    [0., 0., 0.],   # was
])

def softmax(x):
    e = np.exp(x - x.max())          # subtract the max so exp never overflows
    return e / e.sum()

def attend(query, label):
    keys   = E                        # what each word advertises about itself
    values = E                        # what each word hands over if chosen
    scores = keys @ query             # one number per word: how well it matches
    weights = softmax(scores)
    out = weights @ values
    print(f"\n{label}")
    print("  query =", query)
    for w, word, s, p in zip(range(len(words)), words, scores, weights):
        bar = "#" * int(round(p * 40))
        print(f"   {word:8s} score={s:5.1f}  weight={p:.3f} {bar}")
    print("  blended output =", out)

# The word "it" asks a question. The question IS the query vector.
attend(np.array([4., 0., 0.]), "'it was thirsty'  -> ask for something ALIVE:")
attend(np.array([0., 4., 0.]), "'it was sweet'    -> ask for something DRINKABLE:")
Output

'it was thirsty'  -> ask for something ALIVE:
  query = [4. 0. 0.]
   the      score=  0.0  weight=0.016 #
   cat      score=  4.0  weight=0.886 ###################################
   drank    score=  0.0  weight=0.016 #
   the      score=  0.0  weight=0.016 #
   milk     score=  0.0  weight=0.016 #
   because  score=  0.0  weight=0.016 #
   it       score=  0.0  weight=0.016 #
   was      score=  0.0  weight=0.016 #
  blended output = [0.886 0.016 0.049]

'it was sweet'    -> ask for something DRINKABLE:
  query = [0. 4. 0.]
   the      score=  0.0  weight=0.016 #
   cat      score=  0.0  weight=0.016 #
   drank    score=  0.0  weight=0.016 #
   the      score=  0.0  weight=0.016 #
   milk     score=  4.0  weight=0.886 ###################################
   because  score=  0.0  weight=0.016 #
   it       score=  0.0  weight=0.016 #
   was      score=  0.0  weight=0.016 #
  blended output = [0.016 0.886 0.049]

Reading that output line by line

The keys never changed between the two runs. Same sentence, same words, same labels on the jars. Only the query moved, and the answer moved with it. That is the entire point of separating the three roles.

Every word got some weight, even "because". No word ever receives exactly zero. The softmax step turns scores into fractions that add up to one. A fraction is only zero when its score is minus infinity. Attention always spreads a little everywhere.

The 0.886 is not a probability of correctness. It is a mixing proportion. The output is 88.6 percent of what "cat" offered, plus small slices of everything else. Read attention weights as a recipe, never as a confidence score.

Two identical words got identical weights. "the" appears at position 0 and position 3, and both scored 0.016. Nothing in this calculation knows about order. Word order is added separately, before attention ever runs — see transformers.

Where the three actually come from

In a real model you do not hand-write queries. Each one is produced from the token's own vector by a learned matrix:

python
Q = X @ W_q        # X is (tokens, d_model); W_q is (d_model, d_head)
K = X @ W_k
V = X @ W_v

X holds one row per token, straight from the embedding table plus position information. The three matrices W_q, W_k and W_v are ordinary trainable weights, updated by backpropagation like everything else.

So a single token vector fans out into three different vectors. The model decides, during training, what each of them should emphasise.

Common mistakes

Thinking Q, K and V are three different inputs. In self-attention they all come from the same X. Only the matrices differ. In cross-attention they come from two places, and that is what makes cross-attention a different thing.

Reading attention weights as an explanation. A high weight tells you where information flowed, not why the model answered as it did. Attention maps are a useful debugging view and a poor interpretability claim.

Forgetting that W_v can output a different width from W_q and W_k. Query and key must match each other, because they get multiplied together. The value width is free. Most implementations set them equal out of habit, not necessity.

Assuming keys and values are the same thing because they came from the same token. They are two separate projections with two separate matrices, trained for two separate jobs.

Try it yourself

Add a fourth slot to every word for "is a plural noun". Then write a query that hunts for plurals. Predict which word wins before you run it.

Then set values = np.eye(8) instead of E and re-run. The output becomes the attention weights themselves, which makes it obvious that the value projection is a free choice.

What to learn next

Researcher — Mathematics and papers.

The projection view

Given input $X \in \mathbb{R}^{T \times d_{\text{model}}}$ for $T$ tokens, a single attention head defines three learned projections:

$$ Q = X W_Q, \qquad K = X W_K, \qquad V = X W_V $$

with $W_Q, W_K \in \mathbb{R}^{d_{\text{model}} \times d_k}$ and $W_V \in \mathbb{R}^{d_{\text{model}} \times d_v}$. Here $d_k$ is the query and key width, and $d_v$ is the value width. Only $d_k$ is constrained to be shared, because $Q$ and $K^\top$ are multiplied together.

Why the scoring function is really bilinear

The score between token $i$ and token $j$ expands to:

$$ s_{ij} = q_i \cdot k_j = x_i^\top W_Q W_K^\top x_j = x_i^\top \left( W_Q W_K^\top \right) x_j $$

So the pair $(W_Q, W_K)$ only ever appears as the single matrix $W_{QK} = W_Q W_K^\top \in \mathbb{R}^{d_{\text{model}} \times d_{\text{model}}}$. Attention scores are a bilinear form on the input space. Such a function is linear in each of its two arguments separately.

Two consequences follow, and both matter.

The factorisation is a rank constraint, not extra capacity. Since $\operatorname{rank}(W_{QK}) \le d_k$, choosing $d_k < d_{\text{model}}$ deliberately restricts which bilinear forms the head can express. This is the low-rank bottleneck analysed by Bhojanapalli et al. (2020). Head width, not head count, is what caps expressivity.

Storing $W_Q$ and $W_K$ separately costs $2 d_{\text{model}} d_k$ parameters against $d_{\text{model}}^2$ for the full form. At the usual $d_k = d_{\text{model}} / h$, the factored version is cheaper whenever $h > 2$.

The same argument applies on the output side. Elhage et al. (2021) formalise this as two circuits per head. The QK circuit $W_Q W_K^\top$ decides where to read from. The OV circuit $W_V W_O$ decides what gets written. The head's read pattern and its write content are governed by two independent low-rank matrices.

Why not skip the projections entirely

Using raw embeddings as both query and key gives $s_{ij} = x_i \cdot x_j$, which is symmetric: $s_{ij} = s_{ji}$. Language is not symmetric. "Delhi" should attend to "capital" far more than "capital" attends to "Delhi" in most contexts.

The learned $W_{QK}$ is unconstrained and generally asymmetric, which is exactly what is needed. It also lets the head select a subspace. Two tokens far apart overall can still match sharply on the directions that head cares about.

Sharing keys and values across heads

The three projections are not equally expensive at inference time. Queries are recomputed for every new token and thrown away. Keys and values are cached for the whole sequence, and that cache is what dominates memory during generation.

  • Multi-query attention (Shazeer, 2019, arXiv:1911.02150) keeps $h$ query heads but one shared $K$ and $V$. The cache shrinks by a factor of $h$.
  • Grouped-query attention (Ainslie et al., 2023, arXiv:2305.13245) interpolates: $g$ key-value groups with $h$ query heads, $1 \le g \le h$. Llama 3 8B uses $h = 32$, $g = 8$.
  • Multi-head latent attention, introduced in the DeepSeek-V2 report, caches a compressed latent per token. It reconstructs $K$ and $V$ on the fly, trading arithmetic for memory bandwidth.

All three are asymmetric treatments of the three roles. That asymmetry is only available because the roles were separated in the first place.

Provenance

The query-key-value framing predates transformers. Graves et al. (2014), Neural Turing Machines, use content-based addressing. A key is compared explicitly against a memory bank. Bahdanau et al. (2015) introduced learned alignment for sequence models. It scored with a small feedforward network rather than a dot product. Luong et al. (2015) replaced that network with a dot product. Vaswani et al. (2017), Attention Is All You Need, added the separate value projection and the scaling term. It removed recurrence entirely.

Papers

What to learn next