How Models Know Word Order

Why a transformer cannot tell word order

Attention treats a sentence as a bag of words, so "dog bites man" and "man bites dog" reach it identically. Position has to be added by hand.

Read these first

On this page 7
  1. Why this happens
  2. Why older designs did not have this problem
  3. How the problem gets fixed
  4. Where you have already felt this
  5. What is honestly hard here
  6. Remember this
  7. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

A transformer, on its own, cannot tell which word came first. It sees a pile of words, not a line of them.

Picture a bowl of Scrabble tiles. Someone spells D-O-G, B-I-T-E-S, M-A-N, then tips the tiles into the bowl and shakes it. The letters are all still there. The order is gone forever.

That bowl is what raw attention sees. You have surely noticed that "dog bites man" and "man bites dog" mean opposite things. To the attention step, they are the same bowl.

Why this happens

Attention works by comparing every word with every other word. For each pair, it asks one question: how related are these two?

It never asks where either word sat. There is no step in the calculation that looks at a slot number.

So if you shuffle the input words, every pair-comparison gives the same answer as before. The results come out shuffled the same way, and nothing else changes.

This property has a name: permutation equivariance — shuffle the input, and the output shuffles identically, with no other difference.

Why older designs did not have this problem

Before transformers, language models read one word at a time in order. An RNN is a network that keeps a running memory and updates it word by word. It could not lose the order even if it wanted to. Order was baked into the reading.

That was also its weakness. Reading strictly left to right means you cannot process the whole sentence at once, so training is slow.

Transformers threw out sequential reading to gain speed. Losing word order was the bill that came with it.

How the problem gets fixed

You stamp each word with a marker that says where it sat. Then "dog" in slot 0 and "dog" in slot 2 arrive as different things.

   without a stamp:
      dog   bites   man        ->   { dog, bites, man }   <- a bag
      man   bites   dog        ->   { dog, bites, man }   <- the same bag

   with a stamp:
      dog#0 bites#1 man#2      ->   { dog#0, bites#1, man#2 }
      man#0 bites#1 dog#2      ->   { man#0, bites#1, dog#2 }   <- different

Everything else in this section is an argument about what that stamp should be. Add it to the word? Rotate the word by it? Subtract a penalty for distance? Each answer is a lesson ahead.

Where you have already felt this

  • Google Translate handling "the cat chased the dog" versus "the dog chased the cat".
  • A chatbot following "first do A, then do B" in the right order.
  • Code completion knowing that a - b is not b - a.
  • A model reading a date as 03/11 rather than 11/03.

What is honestly hard here

Reading this the first time, most people assume the model must know the order, because the words arrive in an ordered list. It does not. The list order is thrown away in the very first attention step.

That feels wrong until you watch it happen. The code below makes it happen in ten lines. Read it twice if it does not land the first time. That is normal.

Remember this

  • Attention compares pairs of words and never looks at their slots.
  • Shuffling the input shuffles the output and changes nothing else.
  • Position has to be added deliberately, and there are many competing ways to do it.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install numpy

Nothing else. The claim is about the attention formula itself, so NumPy is enough to test it.

Proving it in one script

permutation.py
import numpy as np

np.set_printoptions(precision=3, suppress=True)

# Three word vectors. Deliberately hand-made so you can see what moves.
vocab = {"dog": [1.0, 0.0, 0.0, 0.0],
         "bites": [0.0, 1.0, 0.0, 0.0],
         "man": [0.0, 0.0, 1.0, 0.0]}

# One fixed attention head. Same weights for both sentences.
rng = np.random.default_rng(0)
Wq, Wk, Wv = (rng.normal(size=(4, 4)) for _ in range(3))

def attention(X):
    Q, K, V = X @ Wq, X @ Wk, X @ Wv
    scores = Q @ K.T / np.sqrt(K.shape[1])
    w = np.exp(scores - scores.max(axis=1, keepdims=True))
    w /= w.sum(axis=1, keepdims=True)
    return w @ V

s1 = ["dog", "bites", "man"]
s2 = ["man", "bites", "dog"]          # same three words, reversed

X1 = np.array([vocab[w] for w in s1])
X2 = np.array([vocab[w] for w in s2])

out1, out2 = attention(X1), attention(X2)

print("sentence 1:", " ".join(s1))
print(out1)
print("\nsentence 2:", " ".join(s2))
print(out2)

# Row 0 of sentence 2 is "man". Row 2 of sentence 1 is also "man".
print("\nis sentence 2 only sentence 1's rows reordered? ",
      np.allclose(out2, out1[[2, 1, 0]]))
print("max difference:", np.abs(out2 - out1[[2, 1, 0]]).max())

# Now stamp a position marker onto each word and repeat.
stamp = np.array([[0.0, 0.0, 0.0, 0.0],
                  [0.0, 0.0, 0.0, 1.0],
                  [0.0, 0.0, 0.0, 2.0]])

p1, p2 = attention(X1 + stamp), attention(X2 + stamp)
print("\nwith a position stamp added:")
print("still only reordered? ", np.allclose(p2, p1[[2, 1, 0]]))
print("max difference:", np.abs(p2 - p1[[2, 1, 0]]).max())
Output
sentence 1: dog bites man
[[-0.613  0.639  0.698  0.79 ]
 [-0.391  0.44   0.47   0.706]
 [-0.601  0.709  0.681  0.716]]

sentence 2: man bites dog
[[-0.601  0.709  0.681  0.716]
 [-0.391  0.44   0.47   0.706]
 [-0.613  0.639  0.698  0.79 ]]

is sentence 2 only sentence 1's rows reordered?  True
max difference: 1.1102230246251565e-16

with a position stamp added:
still only reordered?  False
max difference: 2.2606616142462395

Reading that output

Look at the two matrices before the stamp. They contain the same three rows. Row 0 of the second is row 2 of the first, digit for digit. The middle row is untouched, because "bites" stayed in the middle.

The max difference is 1.11e-16. That is one unit in the last place of a 64-bit float. It is not "close". It is exact, and the tiny number reflects floating-point addition order, nothing more.

After adding the stamp, the difference is 2.26. The same three words in a different order now produce genuinely different vectors. That is all a positional encoding does.

The stamp used here is deliberately crude. It writes 0, 1, 2 into the last channel. It works for a three-word toy and fails badly at scale, for reasons the next lessons work through.

The formal statement, in code

Permutation equivariance means attention(P @ X) == P @ attention(X) for any permutation matrix P. You can check that directly:

equivariance.py
import numpy as np

rng = np.random.default_rng(1)
X = rng.normal(size=(6, 8))
Wq, Wk, Wv = (rng.normal(size=(8, 8)) for _ in range(3))

def attention(X):
    s = (X @ Wq) @ (X @ Wk).T / np.sqrt(8)
    w = np.exp(s - s.max(axis=1, keepdims=True))
    return (w / w.sum(axis=1, keepdims=True)) @ (X @ Wv)

perm = rng.permutation(6)
P = np.eye(6)[perm]                       # a permutation matrix

print("attention(P @ X) == P @ attention(X):",
      np.allclose(attention(P @ X), P @ attention(X)))
print("max difference:", np.abs(attention(P @ X) - P @ attention(X)).max())
Output
attention(P @ X) == P @ attention(X): True
max difference: 8.881784197001252e-16

This holds for any width, any number of heads and any number of layers, as long as no positional information enters. Feed-forward layers and layer norm act on each token on its own, so they do not break it either.

Common mistakes

Thinking the causal mask solves it. A causal mask stops a token attending to its right, which does break full permutation equivariance. It leaks a surprising amount of position information, and there is a whole lesson on that. It is not a replacement for a designed positional encoding in most setups.

Thinking the input list order carries through. The list order decides which row goes where in the matrix. After the first attention step, only row content matters, and content carries no order.

Confusing equivariance with invariance. Invariance would mean the output does not change at all. Equivariance means it changes in exactly the same way the input did. Attention is equivariant.

Adding a position stamp with an unbounded magnitude. The crude 0, 1, 2 stamp above grows without limit. At position 4000 it swamps the word meaning entirely. Real schemes are bounded, or they avoid addition altogether.

Try it yourself

Change the stamp so every row gets the same value, for example all 1.0. Predict what np.allclose prints before you run it. Then try a stamp that repeats every two positions, and work out which pairs of sentences it can no longer tell apart.

What to learn next

Researcher — Mathematics and papers.

The statement

Let $X \in \mathbb{R}^{n \times d}$ be the token matrix and $P \in {0,1}^{n \times n}$ a permutation matrix. Single-head self-attention is

$$ \mathrm{Attn}(X) = \mathrm{softmax}!\left( \frac{XW_Q (XW_K)^\top}{\sqrt{d_k}} \right) X W_V $$

Here $W_Q, W_K \in \mathbb{R}^{d \times d_k}$ and $W_V \in \mathbb{R}^{d \times d_v}$ are learned projections, $d_k$ is the head width, and softmax is applied row-wise.

Substituting $PX$:

$$ \frac{(PXW_Q)(PXW_K)^\top}{\sqrt{d_k}} = P \left( \frac{XW_Q(XW_K)^\top}{\sqrt{d_k}} \right) P^\top $$

Row-wise softmax commutes with a symmetric permutation of rows and columns, and $P^\top P = I$, so

$$ \mathrm{Attn}(PX) = P\,\mathrm{Attn}(X) $$

The result extends to multi-head attention (each head is equivariant, and concatenation is per-token), to the position-wise feed-forward block, to layer normalisation, and to residual connections. Every one of those acts row-wise. A stack of $L$ such layers is therefore equivariant, so a transformer encoder with no positional information computes a set function, not a sequence function.

Consequences for expressivity

Such a model cannot compute any function that distinguishes orderings of the same multiset. That rules out language modelling in any serious sense, since even the bigram distribution requires order.

Yun et al. (2020), Are Transformers universal approximators of sequence-to-sequence functions?, make the boundary precise. Transformers are universal approximators of continuous permutation-equivariant sequence-to-sequence functions on a compact domain; adding positional encodings extends this to arbitrary continuous sequence-to-sequence functions. The positional encoding buys the second class.

The two entry points

Position can enter in exactly two places. Every scheme in this section is one of them, or a hybrid.

Additive, at the input. Modify $X$ before attention: $X' = X + E$ where $E \in \mathbb{R}^{n \times d}$ depends on position. Sinusoidal and learned absolute embeddings live here. The change propagates through $W_Q$ and $W_K$, so the attention logit becomes

$$ (x_m + e_m)W_Q W_K^\top (x_n + e_n)^\top $$

which expands into four terms: content-content, content-position, position-content and position-position. Only the last carries pure positional signal, and the cross terms mix meaning with place. That interference is a documented weakness of the additive family.

Inside the attention logit. Modify the score directly:

$$ a_{mn} = \frac{q_m^\top k_n}{\sqrt{d_k}} + b(m, n) $$

or apply a position-dependent transformation to $q$ and $k$ before the dot product. T5's relative bias, ALiBi and RoPE all live here. This family keeps the value path clean and expresses position as a relation between two tokens rather than a property of one.

Why the field moved to the second family

Three reasons, in order of practical weight.

  • Length behaviour. A bias or rotation depending on $m - n$ is defined for any $m, n$. An absolute table of size $L$ is defined only up to $L$.
  • Translation structure. Language is close to translation-invariant at the sentence level. A relative parameterisation encodes that prior directly instead of spending capacity learning it.
  • Cache compatibility. With KV caching, keys are computed once and reused. A scheme applied to keys once, consistently, composes cleanly with cache reuse and with sliding windows.

Causal masking is not a substitute

Applying a lower-triangular mask $M$ with $M_{mn} = 0$ for $n \le m$ and $-\infty$ otherwise destroys equivariance, because $M$ is not permutation-symmetric. The output at position $m$ depends on the multiset of the first $m+1$ tokens, and those multisets differ in size across positions.

Haviv et al. (2022), Transformer Language Models without Positional Encodings Still Learn Positional Information, show decoder-only models exploit this and recover position internally. Kazemnejad et al. (2023), NeurIPS, extend the analysis and show NoPE can express both absolute and relative schemes. That has its own lesson. The point here is that the mask leaks position as a side effect; it is not a designed encoding.

Papers

What to learn next