Attention
In one sentence Attention is the step where a model decides which other words in the input matter most for the word it is handling right now.
Updated
Attention is the mechanism that lets a model weigh which parts of the input matter most for the part it is working on right now.
Read this sentence: "The trophy did not fit in the suitcase because it was too big." To know what "it" means, your eyes flick back to "trophy". Now change one word: "...because it was too small." Suddenly "it" is the suitcase, and you looked back at a different word. You did not read the sentence left to right and hope. You looked back, selectively. That looking-back is attention.
Inside the model, every token sends out a query ("what am I looking for?") and offers a key ("what do I contain?"). Every query is compared against every key, which produces a score for each pair. Those scores are turned into weights that add up to one, and the model builds a new representation of each token as a weighted blend of all the others. High weight means "this word matters to me".
What the weights look like
processing the word "it"
The trophy did not fit in the suitcase because it
0.01 0.58 0.01 0.01 0.04 0.01 0.01 0.22 0.02 0.09
▲
most of the meaning of "it" is pulled from hereBecause every token can look at every other token in one step, a transformer handles a long sentence in parallel instead of word by word. That parallelism is what made training on internet-scale text possible. The cost is that comparing every token with every other token grows with the square of the sequence length, which is a large part of why long contexts are expensive.
Where to go next
- Full lesson: Attention
- Related terms: transformer, token, embedding, llm