AI glossary

Multi-head attention

In one sentence Multi-head attention runs several attention operations in parallel, letting each head watch for a different kind of relationship between tokens.

By Updated

Multi-head attention splits attention into several parallel "heads", each free to track a different kind of relationship between tokens.

One editor reading a manuscript in a single pass must check everything at once: grammar, pronoun references, factual consistency, tone. A publishing house instead sends copies to specialists — one reads only for grammar, one only for continuity, one for tone — then merges their notes. Multi-head attention is the publishing house. Each head runs the full attention mechanism independently, with its own learned way of asking "which tokens matter to me?"

Mechanically, the model's working vector for each token is split across heads. A model with dimension 4,096 and 32 heads gives each head a 128-dimensional slice. Each head computes its own queries, keys and values, attends, and the heads' outputs are concatenated and mixed back together by one final learned projection.

Why not one big head? Attention produces a weighted average, and one averaging operation can express only one pattern of "who matters to whom" per position. Eight heads can hold eight patterns at once — studies of trained models find heads that track subject-verb links, coreference (which "it" means what), adjacent-word syntax, and rare-token copying. Redundancy is real too: many heads can be pruned with modest damage, an observation that leads directly to grouped-query-attention and its cheaper cousins.

Where to go next