How Models Know Word Order

Sinusoidal position encoding

The original transformer stamped each slot with a set of waves at different speeds, so every position gets a unique pattern and nothing has to be learned.

On this page 7
  1. The problem it was invented for
  2. How it works
  3. The property that makes it clever
  4. Where you have already seen this idea
  5. What is honestly hard here
  6. Remember this
  7. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Sinusoidal position encoding gives every slot in a sentence its own pattern of waves, and adds that pattern to the word sitting there.

Think of an old analogue clock. It has a second hand sweeping fast, a minute hand moving slower, and an hour hand crawling. Read all three together and you know the exact time, with no two moments looking alike.

Now think of a wall of clocks, each running at a different speed. Word one gets a snapshot of every dial. Word two gets the next snapshot. Every slot ends up with a fingerprint made of hand positions.

That wall of clocks is the whole idea. The fastest dial changes between neighbouring words. The slowest barely moves over a thousand words.

The problem it was invented for

The first fix anyone reaches for is to write the slot number into the word: 0, 1, 2, 3.

That fails in two ways. Slot 4000 arrives as a huge number that drowns out the meaning of the word. And a model trained up to slot 512 has never seen the number 4000. It has no idea what to do with it.

Waves fix both. A wave never leaves the range from minus one to plus one, no matter how far you go. And a wave repeats, so its shape at a far-away position still looks familiar.

How it works

Take a set of dials. Each dial turns at its own rate. The first spins once every six words. The last takes tens of thousands of words for a single turn.

   slot:      0      1      2      3      4      5   ...

   fast dial  ^      v      ^      v      ^      v      changes every step
   medium     ^      ^      -      v      v      -      changes every few steps
   slow       ^      ^      ^      ^      ^      ^      barely moves
   slowest    ^      ^      ^      ^      ^      ^      almost flat

Read down a column and you get that slot's fingerprint. Read across a row and you see one dial turning.

The fast dials say "these two words are next to each other". The slow dials say "these two are in roughly the same part of the document".

The property that makes it clever

Here is the part that is genuinely elegant. Turning a dial forward by five steps is the same operation, no matter where the dial started.

So the model can learn one fixed rule that means "look five words back". That rule works at word ten and at word ten thousand, without being taught separately for each.

Nothing about this is learned. The whole thing is a fixed formula, computed once, added to the word vectors, and never trained. That is why it costs zero parameters.

Where you have already seen this idea

  • The hands of a clock telling you a time with three numbers.
  • An odometer where the last digit spins and the first digit crawls.
  • Musical notes, where octaves stack frequencies to make each pitch distinct.
  • Longitude and latitude giving a unique fingerprint to every spot on Earth.

What is honestly hard here

The claim "waves make each position unique" sounds like hand-waving until you check it. The reason it works is that dial speeds are chosen to be very spread out. Two positions would need to match on every dial at once to collide, and that takes tens of thousands of steps.

The other honest point: this scheme handles longer inputs without crashing, but that is not the same as handling them well. That distinction gets its own lesson later.

Remember this

  • Every slot gets a fingerprint made of waves running at different speeds.
  • It is a fixed formula, so it costs no learned parameters at all.
  • Shifting by a fixed number of steps is the same operation everywhere, which is why relative distance is easy to learn.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install numpy

The encoding, written out

sinusoidal.py
import numpy as np
np.set_printoptions(precision=3, suppress=True)

def sinusoidal(n_pos, d_model, base=10000.0):
    pos = np.arange(n_pos)[:, None]                 # (n_pos, 1)
    i = np.arange(d_model // 2)[None, :]            # one i per sin/cos pair
    freq = 1.0 / (base ** (2 * i / d_model))        # small i turns fast, large i turns slowly
    pe = np.zeros((n_pos, d_model))
    pe[:, 0::2] = np.sin(pos * freq)                # even columns
    pe[:, 1::2] = np.cos(pos * freq)                # odd columns
    return pe

small = sinusoidal(n_pos=4, d_model=8)
print("first four positions, 8 dimensions each:")
print(small)

print("\nwavelength of each sin/cos pair, d_model=8 (tokens per full cycle):")
for i in range(4):
    print(f"  pair {i}: {2 * np.pi * 10000.0 ** (2 * i / 8):10.1f}")

# A realistic width. Now look at how similarity falls off with the gap.
PE = sinusoidal(n_pos=70, d_model=64)
print("\ndot product between two positions, d_model=64:")
print("  gap   from pos 0   from pos 30")
for k in [0, 1, 2, 4, 8, 16, 32]:
    print(f"  {k:>3}   {PE[0] @ PE[k]:9.3f}   {PE[30] @ PE[30 + k]:10.3f}")
Output
first four positions, 8 dimensions each:
[[ 0.     1.     0.     1.     0.     1.     0.     1.   ]
 [ 0.841  0.54   0.1    0.995  0.01   1.     0.001  1.   ]
 [ 0.909 -0.416  0.199  0.98   0.02   1.     0.002  1.   ]
 [ 0.141 -0.99   0.296  0.955  0.03   1.     0.003  1.   ]]

wavelength of each sin/cos pair, d_model=8 (tokens per full cycle):
  pair 0:        6.3
  pair 1:       62.8
  pair 2:      628.3
  pair 3:     6283.2

dot product between two positions, d_model=64:
  gap   from pos 0   from pos 30
    0      32.000       32.000
    1      30.917       30.917
    2      28.304       28.304
    4      23.934       23.934
    8      22.414       22.414
   16      19.373       19.373
   32      19.554       19.554

Reading that output carefully

Position 0 is all zeros and ones. Every sine of zero is zero, every cosine of zero is one. That row is the reference the others are measured against.

The last two columns barely move. At d_model=8 the slowest pair has a wavelength of 6283 tokens, so across four positions it moves from 0.000 to 0.003. Those channels are almost useless for short text and carry the coarse "which chapter am I in" signal for long text.

The two dot-product columns are identical. The similarity between positions 0 and 4 equals the similarity between 30 and 34. The encoding depends only on the gap, not on where you start. That is the whole point.

The decline is not perfectly monotone. Gap 32 scores 19.554, slightly above gap 16 at 19.373. Similarity falls off with distance in the broad sense, then wobbles. Anyone who tells you the dot product decreases smoothly forever has not printed it. This is exactly why the encoding is added and then reshaped by learned projections, rather than used as a distance measure on its own.

The shift property, proved rather than asserted

The paper's argument is that PE[pos + k] is a fixed linear function of PE[pos]. Since each pair is a point on a circle, moving forward by k is a rotation, and a rotation is a matrix:

shift.py
import numpy as np

def sinusoidal(n_pos, d_model, base=10000.0):
    pos = np.arange(n_pos)[:, None]
    i = np.arange(d_model // 2)[None, :]
    freq = 1.0 / (base ** (2 * i / d_model))
    pe = np.zeros((n_pos, d_model))
    pe[:, 0::2] = np.sin(pos * freq)
    pe[:, 1::2] = np.cos(pos * freq)
    return pe

def shift_matrix(k, d_model, base=10000.0):
    """One matrix that turns PE[p] into PE[p+k], for every p."""
    M = np.zeros((d_model, d_model))
    for i in range(d_model // 2):
        a = k / base ** (2 * i / d_model)          # how far this pair turns in k steps
        s, c = np.sin(a), np.cos(a)
        M[2*i:2*i+2, 2*i:2*i+2] = [[c, -s], [s, c]]
    return M

PE = sinusoidal(1000, 64)
M = shift_matrix(5, 64)
print("positions 0-99   :", np.abs(PE[:100] @ M - PE[5:105]).max())
print("positions 900-989:", np.abs(PE[900:990] @ M - PE[905:995]).max())
Output
positions 0-99   : 1.3711254354120683e-14
positions 900-989: 9.646450305211829e-14

One matrix. Every position. Error at the level of float64 rounding, at position 900 as much as at position 0.

This matters because W_Q and W_K inside attention can, in principle, learn something equivalent to that matrix. "Attend five tokens back" becomes one reusable operation instead of a separate fact per position.

Common mistakes

Getting the even and odd interleaving wrong. The paper puts sine in even channels and cosine in odd, pairing channel 2i with 2i+1. Many implementations instead concatenate all sines then all cosines. Both work when used consistently, but mixing the two conventions between training and inference silently destroys the encoding. Print the first four rows and compare.

Forgetting to scale the token embeddings. The original transformer multiplies embeddings by sqrt(d_model) before adding the encoding. Skip it and the positional signal, which is bounded by 1, dominates freshly initialised embeddings. This is one line and it is easy to lose in a rewrite.

Assuming it extrapolates well. The formula is defined for any position, so nothing crashes. Quality at unseen lengths is a separate question, and the answer is disappointing. See the length-extrapolation lesson.

Learning it instead of fixing it. Making these values trainable turns them into a learned table with a clever initialisation. That is a valid choice, and it is a different method with different properties.

Try it yourself

Change base from 10000 to 100 and reprint the wavelengths. Every dial now turns much faster, so distant positions start colliding. Find the smallest sequence length where two positions become nearly identical, and you have rediscovered why the base is large.

What to learn next

Researcher — Mathematics and papers.

Definition

For position $pos$ and dimension index $i \in {0, \dots, d/2 - 1}$, with model width $d$ and base $b = 10000$:

$$ PE_{(pos,\,2i)} = \sin!\left(\frac{pos}{b^{2i/d}}\right), \qquad PE_{(pos,\,2i+1)} = \cos!\left(\frac{pos}{b^{2i/d}}\right) $$

Write $\theta_i = b^{-2i/d}$ for the angular frequency of pair $i$. The wavelength is $\lambda_i = 2\pi/\theta_i = 2\pi b^{2i/d}$, a geometric progression from $2\pi$ to $2\pi b$. With $b = 10000$ that spans roughly 6.28 to 62832 tokens.

The result is added to the token embedding, after the embedding is scaled by $\sqrt{d}$:

$$ h^{(0)}{pos} = \sqrt{d}\, E[x{pos}] + PE_{pos} $$

The shift theorem

For each pair $i$, define $v_i(pos) = (\sin \theta_i pos,\; \cos \theta_i pos)^\top$. Then

$$ v_i(pos + k) = \begin{pmatrix} \cos \theta_i k & \sin \theta_i k \ -\sin \theta_i k & \cos \theta_i k \end{pmatrix} v_i(pos) $$

The matrix depends on $k$ and $i$, never on $pos$. Stacking the $d/2$ blocks gives a single block-diagonal $M_k \in \mathbb{R}^{d \times d}$ with $PE_{pos+k} = M_k\, PE_{pos}$ for all $pos$. This is Vaswani et al.'s stated motivation: relative attention is expressible as a fixed linear map.

The inner product

$$ PE_{pos}^\top PE_{pos + k} = \sum_{i=0}^{d/2 - 1} \cos(\theta_i k) $$

Every dependence on $pos$ cancels; only the gap $k$ survives. That is the identity the developer output verifies numerically.

Treating $i/d$ as continuous, the sum approaches

$$ \frac{d}{2} \int_0^1 \cos!\left(k\, b^{-2u}\right) du $$

which decays with $k$ but is not monotone. Yan et al. (2019), TENER, point out the related and more damaging fact: the property is destroyed once the query and key projections are applied, because $PE_{pos}^\top W_Q W_K^\top PE_{pos+k}$ carries no such cancellation for general $W_Q, W_K$. The elegant relative structure exists at the input and is not preserved by the layer that consumes it.

Why the base is 10000

The base sets the longest wavelength, $2\pi b$. Two positions collide when every pair returns to nearly the same angle, which requires alignment across the whole geometric spread. Larger $b$ pushes the coarsest resolution further out at the cost of the slowest channels being nearly constant over any realistic window.

The value 10000 is not derived; the paper reports it as a choice. It reappears as RoPE's default $\theta$ base, and modern long-context models raise it substantially — Llama 3 uses 500000, and many long-context finetunes go higher still. Base scaling for context extension is covered in the RoPE scaling lesson.

Empirical standing

  • Vaswani et al. (2017) reported near-identical results between sinusoidal and learned absolute embeddings on WMT translation, and chose sinusoidal for the hope of length extrapolation.
  • Press et al. (2022), Train Short, Test Long (ALiBi), measured that hope directly and found sinusoidal extrapolation poor: perplexity rises sharply past the training length rather than holding.
  • Kazemnejad et al. (2023) rank absolute position embeddings, sinusoidal included, at the bottom for downstream length generalisation in decoder-only models.

The scheme is therefore historically central and largely superseded. It still appears in encoder-only vision and audio models, in diffusion timestep embeddings, and anywhere a fixed, parameter-free positional signal over a bounded range is wanted.

Cost

Zero parameters. $O(nd)$ to build the table, computed once and cached. Memory $O(Ld)$ for maximum length $L$, which is negligible next to the weights.

Papers

What to learn next