Inside a Transformer Section 085

Looking Inside a Trained Model

A trained model is not a black box you must accept. These are the tools people use to read what is happening inside it.

9 of 9 lessons published Three reading levels on every lesson

Start with “The logit lens”

Lessons in order

Work top to bottom. Each lesson assumes the one above it.

  1. The logit lens
  2. Induction heads
  3. Activation patching
  4. Probing hidden states
  5. Superposition and polysemantic neurons
  6. Sparse autoencoders for feature discovery
  7. Steering a model with activation vectors
  8. Where a model stores a fact
  9. Visualising attention maps