AI glossary

Positional encoding

In one sentence Positional encoding injects word-order information into a transformer, which otherwise sees its input as an unordered set.

By Updated

Positional encoding is any scheme that tells a transformer where each token sits in the sequence, because attention alone cannot tell.

Drop a stack of numbered pages and pick them up unnumbered, and the novel becomes a shuffled mess — the pages' content never changed, only your knowledge of their order. A transformer reads all tokens in parallel, and attention compares tokens as an unordered set. Without position information, "the dog chased the cat" and "the cat chased the dog" would be indistinguishable. Positional encoding writes the page numbers back on.

The schemes differ in how the number is written:

sinusoidal (2017)   add a fixed wave-pattern vector to each token's embedding
learned positions   train one embedding per position slot (BERT, GPT-2)
relative encodings  encode distances between tokens, not absolute slots
RoPE (modern)       rotate attention vectors by position-dependent angles
ALiBi               penalise attention scores by distance, no vectors at all

The trend line matters more than the details: absolute schemes ("this is token #7") lost to relative ones ("these tokens are 4 apart"), because language structure is relative and because relative schemes extrapolate better past the training length — the key to stretching a context-window. RoPE is the current default in open models.

One reassurance for readers of the original paper: the sinusoidal formulas look intimidating, but they are one idea — give every position a unique, smoothly varying fingerprint the network can learn to read.

Where to go next