Inside a Transformer Section 080

Fast Attention and Long Context

Every trick used to make attention faster, cheaper, or able to read a whole book without running out of memory.

12 of 12 lessons published Three reading levels on every lesson

Start with “FlashAttention”

Lessons in order

Work top to bottom. Each lesson assumes the one above it.

  1. FlashAttention
  2. Memory-bound vs compute-bound
  3. Multi-query attention
  4. Grouped-query attention
  5. Multi-head latent attention
  6. Sliding-window attention
  7. Attention sinks and streaming
  8. Sparse and block attention patterns
  9. Linear attention
  10. Mamba and state space models
  11. Ring attention and sequence parallelism
  12. Advertised context vs usable context