All sections

The whole curriculum, grouped into 14 areas. Start at the top if you are new. If you came here for one thing, filter to its area or search — everything is on this page or one click from it.

14 areas 117 sections 1,162 of 1,167 lessons published 10 new lessons every day

Showing Inside a Transformer — 11 sections, 120 lessons. Show all 14 areas

Inside a Transformer

Open the model up: attention, tokenisers, how text is generated, and how these things are trained and shrunk.

11 sections 115 of 120 lessons published

Attention Mechanics

Open up the attention layer and look at every number that moves through it, one step at a time.

9 published

  1. Queries, keys and values
  2. Scaled dot-product attention
  3. Attention from scratch in NumPy
  4. Multi-head attention
  5. Causal masking
  6. Cross-attention
  7. Why attention costs grow with the square of length
  8. Softmax overflow inside attention
  9. QK normalisation

Inside a Transformer Block

A transformer is one small block repeated many times. This section takes that block apart piece by piece.

12 published

  1. Walking through one transformer block
  2. The residual stream
  3. Layer normalisation
  4. Pre-norm vs post-norm
  5. RMSNorm
  6. The feedforward layer
  7. SwiGLU and gated feedforward layers
  8. Tied input and output embeddings
  9. Decoder-only vs encoder-decoder
  10. Counting a model's parameters by hand
  11. FLOPs per token
  12. Building a tiny GPT from scratch

How Models Know Word Order

Attention sees a bag of words with no order at all, so position has to be added by hand. Here is every way people do it.

10 published

  1. Why a transformer cannot tell word order
  2. Sinusoidal position encoding
  3. Learned position embeddings
  4. Rotary position embeddings (RoPE)
  5. Stretching RoPE for longer context
  6. ALiBi
  7. Relative position embeddings
  8. Models with no positional encoding at all
  9. What breaks when you exceed the trained length
  10. Positions in images and video

Tokeniser Internals

Models never see letters. This section shows exactly how your text becomes numbers, and every strange bug that causes.

13 published

  1. Byte pair encoding, implemented
  2. WordPiece
  3. The unigram tokeniser
  4. SentencePiece
  5. Byte-level BPE
  6. Training a tokeniser on your own text
  7. Choosing a vocabulary size
  8. Special tokens
  9. Chat templates
  10. Glitch tokens
  11. Why models miscount letters and digits
  12. Counting tokens before you send them
  13. Adding tokens and resizing embeddings

How Text Is Generated

A model only ever predicts one next token. Everything you see as an answer comes from how that choice is made, over and over.

15 published

  1. Prefill and decode
  2. The KV cache
  3. How much memory the KV cache eats
  4. PagedAttention
  5. Top-k and nucleus sampling
  6. Min-p and typical sampling
  7. Repetition and frequency penalties
  8. Beam search, and why chat models dropped it
  9. Constrained decoding and grammars
  10. Reading logprobs
  11. Speculative decoding
  12. Medusa, EAGLE and self-speculation
  13. Stop tokens and stopping criteria
  14. Why temperature zero is still not deterministic
  15. Streaming tokens correctly

Fast Attention and Long Context

Every trick used to make attention faster, cheaper, or able to read a whole book without running out of memory.

12 published

  1. FlashAttention
  2. Memory-bound vs compute-bound
  3. Multi-query attention
  4. Grouped-query attention
  5. Multi-head latent attention
  6. Sliding-window attention
  7. Attention sinks and streaming
  8. Sparse and block attention patterns
  9. Linear attention
  10. Mamba and state space models
  11. Ring attention and sequence parallelism
  12. Advertised context vs usable context

Mixture of Experts

How a model can have hundreds of billions of parameters but only use a few of them for each word you type.

10 published

  1. What a mixture of experts really is
  2. The router
  3. Load balancing and expert collapse
  4. Expert capacity and dropped tokens
  5. Shared and fine-grained experts
  6. Active parameters vs total parameters
  7. Serving a MoE across GPUs
  8. A working MoE layer in PyTorch
  9. Mixture of depths
  10. Early exit and layer skipping

How Models Are Actually Trained

What really happens in the months of compute that turn a pile of text and random numbers into a model that can write.

8 published

  1. The next-token objective
  2. Cross-entropy and perplexity
  3. Packing documents into fixed-length batches
  4. Warmup, cosine and WSD schedules
  5. Loss spikes and gradient clipping
  6. Scaling laws and compute-optimal training
  7. Building a pretraining corpus
  8. Reading a training run while it happens

Post-training and Alignment

A freshly pretrained model just continues text. This is every step that turns it into something that answers you helpfully.

14 published

  1. Instruction tuning
  2. Formatting an SFT dataset
  3. Reward models
  4. The KL penalty and the reference model
  5. Direct preference optimisation
  6. Building preference data
  7. GRPO
  8. RL with verifiable rewards
  9. Thinking tokens and reasoning models
  10. Distilling a large model into a small one
  11. Adapters beyond LoRA
  12. Catastrophic forgetting
  13. Merging model weights
  14. Reward hacking and sycophancy

Quantised LLM Inference

Shrinking a language model so it fits and runs fast, and knowing exactly what that costs you in quality.

3 published, 5 on the way

  1. Compressing the KV cache
  2. fp32, bf16, fp8 and int4
  3. GPTQ
  4. AWQ soon
  5. GGUF and llama.cpp quantisation soon
  6. Measuring what quantisation costs you soon
  7. Separating prefill from decode soon
  8. CUDA graphs and kernel fusion soon

Looking Inside a Trained Model

A trained model is not a black box you must accept. These are the tools people use to read what is happening inside it.

9 published

  1. The logit lens
  2. Induction heads
  3. Activation patching
  4. Probing hidden states
  5. Superposition and polysemantic neurons
  6. Sparse autoencoders for feature discovery
  7. Steering a model with activation vectors
  8. Where a model stores a fact
  9. Visualising attention maps

Other ways to use this site

Learning paths

An ordered route through the lessons for one job: ML developer, AI engineer, computer vision engineer.

Projects

Build real things with full code, a dataset and a deployment guide.

AI glossary

Every term you keep seeing, defined in one plain sentence first.

Error database

Paste the error you got. Find out what it means and how to fix it.