RoPE (rotary position embedding)
In one sentence RoPE tells a transformer where each token sits in the sequence by rotating its query and key vectors by position-dependent angles.
Updated
RoPE, short for rotary position embedding, encodes each token's position by rotating its attention vectors by an angle proportional to that position.
A transformer is position-blind: attention sees a set of tokens, so "dog bites man" and "man bites dog" would look identical without extra help. Some positional-encoding must inject word order.
RoPE's trick is geometric. Picture each token's query and key vectors as clock hands. Token 1's hand is rotated by one small step, token 2's by two steps, token 50's by fifty. When attention compares two tokens, the score depends on the angle between their hands — and that angle depends only on how far apart they are, not on where they sit absolutely. Tokens 3-and-7 relate exactly like tokens 103-and-107.
That relative-distance property is why RoPE won. Language cares about relative structure — an adjective modifies the noun next to it, wherever the pair appears in the document. RoPE bakes this in with no extra parameters and no addition step: the rotation is applied directly to queries and keys inside each attention layer.
Introduced by Su et al. in 2021, it is the de facto standard in modern open LLMs — Llama, Qwen, Mistral, DeepSeek. A practical bonus made it famous: the rotation frequencies can be rescaled after training ("RoPE scaling", YaRN), stretching a model's usable context-window far past its training length with light fine-tuning.
Where to go next
- Full lesson: Transformers
- Related terms: positional-encoding, attention, context-window, transformer