Inside a Transformer Section 079
How Text Is Generated
A model only ever predicts one next token. Everything you see as an answer comes from how that choice is made, over and over.
15 of 15 lessons published Three reading levels on every lesson
Lessons in order
Work top to bottom. Each lesson assumes the one above it.
- Prefill and decode
- The KV cache
- How much memory the KV cache eats
- PagedAttention
- Top-k and nucleus sampling
- Min-p and typical sampling
- Repetition and frequency penalties
- Beam search, and why chat models dropped it
- Constrained decoding and grammars
- Reading logprobs
- Speculative decoding
- Medusa, EAGLE and self-speculation
- Stop tokens and stopping criteria
- Why temperature zero is still not deterministic
- Streaming tokens correctly