Inside a Transformer Section 077

How Models Know Word Order

Attention sees a bag of words with no order at all, so position has to be added by hand. Here is every way people do it.

10 of 10 lessons published Three reading levels on every lesson

Start with “Why a transformer cannot tell word order”

Lessons in order

Work top to bottom. Each lesson assumes the one above it.

  1. Why a transformer cannot tell word order
  2. Sinusoidal position encoding
  3. Learned position embeddings
  4. Rotary position embeddings (RoPE)
  5. Stretching RoPE for longer context
  6. ALiBi
  7. Relative position embeddings
  8. Models with no positional encoding at all
  9. What breaks when you exceed the trained length
  10. Positions in images and video