Inside a Transformer Section 083
Post-training and Alignment
A freshly pretrained model just continues text. This is every step that turns it into something that answers you helpfully.
14 of 14 lessons published Three reading levels on every lesson
Lessons in order
Work top to bottom. Each lesson assumes the one above it.
- Instruction tuning
- Formatting an SFT dataset
- Reward models
- The KL penalty and the reference model
- Direct preference optimisation
- Building preference data
- GRPO
- RL with verifiable rewards
- Thinking tokens and reasoning models
- Distilling a large model into a small one
- Adapters beyond LoRA
- Catastrophic forgetting
- Merging model weights
- Reward hacking and sycophancy