Skip to main content
Learn AI Create the Future — with Pranay
  • Learn
  • Paths
  • Projects
  • Practice
  • Glossary
  • Errors
  • LeetCode
  • Main site

Main site →
Select language

Translation is unavailable right now. Your browser’s own “Translate page” usually works.

Machine translation by Google

Topics

  • 🧱 Foundations 2 sections
  • 🤖 Machine Learning 11 sections
  • 📊 Statistics and Experiments 4 sections
  • 🧭 Doing the Work 9 sections
  • 🧠 Deep Learning 11 sections
  • 🧰 Libraries and Frameworks 4 sections
  • 💬 Language and NLP 15 sections
  • 👁️ Computer Vision 13 sections
  • 🎙️ Speech, Forecasting and Recommenders 3 sections
  • ✨ Generative AI and LLMs 4 sections
  • 🔬 Inside a Transformer 11 sections
    • Attention Mechanics 9 lessons
    • Inside a Transformer Block 12 lessons
    • How Models Know Word Order 10 lessons
    • Tokeniser Internals 13 lessons
    • How Text Is Generated 15 lessons
      • Overview
      • Prefill and decode
      • The KV cache
      • How much memory the KV cache eats
      • PagedAttention
      • Top-k and nucleus sampling
      • Min-p and typical sampling
      • Repetition and frequency penalties
      • Beam search, and why chat models dropped it
      • Constrained decoding and grammars
      • Reading logprobs
      • Speculative decoding
      • Medusa, EAGLE and self-speculation
      • Stop tokens and stopping criteria
      • Why temperature zero is still not deterministic
      • Streaming tokens correctly
    • Fast Attention and Long Context 12 lessons
    • Mixture of Experts 10 lessons
    • How Models Are Actually Trained 8 lessons
    • Post-training and Alignment 14 lessons
    • Quantised LLM Inference 8 lessons
    • Looking Inside a Trained Model 9 lessons
  • ⚙️ Production and MLOps 16 sections
  • 🌍 AI in the Real World 12 sections
  • ⚖️ Safety, Ethics and Law 2 sections

See all topics

  1. Home
  2. Learn
  3. How Text Is Generated

🔬 Inside a Transformer · Section 079

🎲 How Text Is Generated

A model only ever predicts one next token. Everything you see as an answer comes from how that choice is made, over and over.

Every lesson in this section is written by Pranay Mahendrakar.

15 of 15 lessons published · Three reading levels on every lesson

Start with “Prefill and decode”

Lessons in order

Work top to bottom. Each lesson assumes the one above it.

  1. 01 Prefill and decode
  2. 02 The KV cache
  3. 03 How much memory the KV cache eats
  4. 04 PagedAttention
  5. 05 Top-k and nucleus sampling
  6. 06 Min-p and typical sampling
  7. 07 Repetition and frequency penalties
  8. 08 Beam search, and why chat models dropped it
  9. 09 Constrained decoding and grammars
  10. 10 Reading logprobs
  11. 11 Speculative decoding
  12. 12 Medusa, EAGLE and self-speculation
  13. 13 Stop tokens and stopping criteria
  14. 14 Why temperature zero is still not deterministic
  15. 15 Streaming tokens correctly
Previous Tokeniser Internals Next Fast Attention and Long Context

Learn AI

Learn Artificial Intelligence the easy way.

Learn. Build. Ask. Share.

Free forever. No sign-up, no ads, no tracking.

Learn

  • All topics
  • Learning paths
  • Projects
  • Practice playground
  • AI glossary
  • Error database
  • LeetCode solutions

Start here

  • Python for AI
  • Mathematics for AI
  • Machine Learning
  • Classic Algorithms in Depth
  • Linear Models and Regularisation
  • Ensembles and Gradient Boosting

This site

  • Search
  • New lessons feed
  • Sitemap
  • About Pranay Mahendrakar
  • pranaymahendrakar.com

© 2026 Pranay Mahendrakar. Written for people who are starting from zero.

10 new lessons every day. If something here is wrong or confusing, that is worth fixing — say so.