Inside a Transformer Section 082

How Models Are Actually Trained

What really happens in the months of compute that turn a pile of text and random numbers into a model that can write.

8 of 8 lessons published Three reading levels on every lesson

Start with “The next-token objective”

Lessons in order

Work top to bottom. Each lesson assumes the one above it.

  1. The next-token objective
  2. Cross-entropy and perplexity
  3. Packing documents into fixed-length batches
  4. Warmup, cosine and WSD schedules
  5. Loss spikes and gradient clipping
  6. Scaling laws and compute-optimal training
  7. Building a pretraining corpus
  8. Reading a training run while it happens