Inside a Transformer Section 077
How Models Know Word Order
Attention sees a bag of words with no order at all, so position has to be added by hand. Here is every way people do it.
10 of 10 lessons published Three reading levels on every lesson
Lessons in order
Work top to bottom. Each lesson assumes the one above it.
- Why a transformer cannot tell word order
- Sinusoidal position encoding
- Learned position embeddings
- Rotary position embeddings (RoPE)
- Stretching RoPE for longer context
- ALiBi
- Relative position embeddings
- Models with no positional encoding at all
- What breaks when you exceed the trained length
- Positions in images and video