Inside a Transformer Section 080
Fast Attention and Long Context
Every trick used to make attention faster, cheaper, or able to read a whole book without running out of memory.
12 of 12 lessons published Three reading levels on every lesson
Lessons in order
Work top to bottom. Each lesson assumes the one above it.
- FlashAttention
- Memory-bound vs compute-bound
- Multi-query attention
- Grouped-query attention
- Multi-head latent attention
- Sliding-window attention
- Attention sinks and streaming
- Sparse and block attention patterns
- Linear attention
- Mamba and state space models
- Ring attention and sequence parallelism
- Advertised context vs usable context