Inside a Transformer Section 075

Attention Mechanics

Open up the attention layer and look at every number that moves through it, one step at a time.

9 of 9 lessons published Three reading levels on every lesson

Start with “Queries, keys and values”

Lessons in order

Work top to bottom. Each lesson assumes the one above it.

  1. Queries, keys and values
  2. Scaled dot-product attention
  3. Attention from scratch in NumPy
  4. Multi-head attention
  5. Causal masking
  6. Cross-attention
  7. Why attention costs grow with the square of length
  8. Softmax overflow inside attention
  9. QK normalisation