Models with no positional encoding at all
Remove every positional mechanism from a decoder and it still learns word order, because the causal mask leaks position. NoPE works, and sometimes generalises better.
- 11 min read
- 3 reading levels
- Updated
Read these first
On this page 6
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Take out every positional mechanism from a text-generating model and it still learns word order. The rule that stops it looking ahead gives the order away.
Think of a relay race where each runner can see everyone behind them but nobody in front. The first runner sees an empty track. The last runner sees a long line of people.
Nobody told any runner their number. Each one can count the people behind them and work it out.
That counting is what a text model does. It cannot see the future, so how much past it can see tells it where it is.
Why this is surprising
An earlier lesson showed that attention treats a sentence as a bag of words and loses the order. That is true, and it applies to models that read the whole input at once.
A text-generating model does not read the whole input at once. When it processes word five, it is forbidden from looking at words six onward. That rule is called causal masking — blocking each position from attending to anything on its right.
Causal masking is there to stop cheating during training. Position information leaks out of it as a side effect.
How the leak works
Word 0 can attend to exactly one token: itself. Word 1 can attend to two. Word 31 can attend to thirty-two.
Attention averages over whatever it can see. Averaging over one item gives a noisy result. Averaging over thirty-two gives a smooth one.
So the amount of noise left in the output tells you roughly how many tokens were visible, which tells you the position.
position 0: average of 1 token -> noisy
position 3: average of 4 tokens -> less noisy
position 31: average of 32 tokens -> smooth
how smooth the output is -> roughly what position we are atThis is not a metaphor. The code below measures it and reads positions back out of a model that was given none.
Why anyone would want this
Three reasons.
Nothing to run out of. No table with a last row, no dial spinning into unseen angles. There is nothing whose range can be exceeded.
It costs nothing. No parameters, no extra tensor, no work in the attention kernel.
It sometimes generalises better. One careful comparison tested five schemes at longer lengths than they were trained on. The decoder with no positional encoding won.
The honest limits
This only works for models that generate text one token at a time. A model that reads the whole input at once, like BERT, is genuinely order-blind and needs a real encoding.
It is also not what large production models do. Almost every model you can download uses rotary encodings. The evidence for the no-encoding approach comes from careful small-scale comparisons, not from frontier-scale training runs.
Recent large models do something in between: they use rotary encodings in most layers and leave a few layers with none. That mix is currently the practical use of this idea.
Remember this
- A causal mask leaks position, because each token can see a different number of predecessors.
- Text-generating models can learn order with no positional mechanism at all.
- Models that read a whole input at once cannot, and still need one.
What to learn next
- What breaks when you exceed the trained length — all five schemes measured side by side.
- Rotary position embeddings (RoPE) — what production models use instead.
- How LLMs work — where causal masking fits in the bigger picture.
Developer — Code and libraries.
Setup
pip install numpyMeasuring the leak
import numpy as np
np.set_printoptions(precision=3, suppress=True)
rng = np.random.default_rng(0)
SEQ, DIM = 32, 16
def attend(X, causal):
scores = X @ X.T / np.sqrt(DIM)
if causal:
scores = scores + np.triu(np.full((SEQ, SEQ), -np.inf), 1)
w = np.exp(scores - scores.max(axis=1, keepdims=True))
return (w / w.sum(axis=1, keepdims=True)) @ X
# Step 1: does a causal mask break the bag-of-words problem on its own?
X = rng.normal(size=(SEQ, DIM))
perm = rng.permutation(SEQ)
for causal in [False, True]:
a = attend(X, causal)[perm] # attend, then reorder the rows
b = attend(X[perm], causal) # reorder the tokens, then attend
print(f"causal={causal!s:<5} reordering input == reordering output:", np.allclose(a, b))
# Step 2: with a causal mask, how much position information leaks out?
base = rng.normal(size=DIM)
out = attend(base + 0.5 * rng.normal(size=(SEQ, DIM)), causal=True)
print("\ndistance of each output row from the shared base vector:")
for i in [0, 1, 3, 7, 15, 31]:
print(f" position {i:>2}: {np.linalg.norm(out[i] - base):.4f}")
# Step 3: can a probe read the position back out?
def probe(causal):
rows = [attend(base + 0.5 * np.random.default_rng(1000 + t).normal(size=(SEQ, DIM)), causal)
for t in range(400)]
A = np.concatenate(rows)
n = np.linalg.norm(A - base, axis=1, keepdims=True)
F = np.hstack([A, n, n ** 2, np.ones((len(A), 1))]) # let the probe see the length too
y = np.tile(np.arange(SEQ), 400).astype(float)
w, *_ = np.linalg.lstsq(F, y, rcond=None)
return np.abs(F @ w - y).mean()
print("\nmean error of a probe guessing the position (0-31), no position encoding:")
print(f" causal attention : {probe(True):.2f} positions")
print(f" bidirectional attention: {probe(False):.2f} positions")
print(f" always guessing 15.5 : {np.abs(np.arange(SEQ) - 15.5).mean():.2f} positions")causal=False reordering input == reordering output: True causal=True reordering input == reordering output: False distance of each output row from the shared base vector: position 0: 1.6259 position 1: 1.2553 position 3: 0.9547 position 7: 0.9701 position 15: 0.6686 position 31: 0.4240 mean error of a probe guessing the position (0-31), no position encoding: causal attention : 4.63 positions bidirectional attention: 8.00 positions always guessing 15.5 : 8.00 positions
Reading that output
Step 1 is the crux. Without the mask, reordering the input is identical to reordering the output — the bag-of-words property. With the mask, that equality fails. The mask alone destroys permutation equivariance, before any positional encoding is involved.
The distances shrink from 1.63 to 0.42. Position 0 averages a single noisy token, so its output sits far from the shared base. Position 31 averages thirty-two, and the noise partly cancels. The trend is not perfectly smooth — position 7 is marginally above position 3 — because these are finite random draws, not an expectation.
The probe gets to 4.63 positions of error against a 8.00 baseline. The baseline is what you achieve by always guessing the middle. Cutting the error nearly in half, from an untrained network with no positional mechanism, is the leak measured directly.
The bidirectional number is 8.00, matching the baseline to two decimals. With no mask, there is no signal, and the probe recovers nothing. That contrast is the point of running both.
Two honest caveats. The probe is given the output vector plus its length and squared length as features, because the signal lives in the magnitude and a purely linear probe on the raw values recovers almost nothing. And this is untrained attention with hand-made inputs, so it demonstrates that the information is present, not that a trained model uses it well.
What real models do with it
Haviv et al. (2022) trained language models with no positional encoding and probed intermediate layers. The models recover position, and the recovery is strongest in early layers. Kazemnejad et al. (2023) followed with a controlled comparison of five schemes on downstream tasks tested beyond the training length, and NoPE came out ahead.
A tiny reproduction of that comparison is in the next lesson, with the numbers from an actual training run rather than a probe.
Common mistakes
Applying this to an encoder. BERT-style models attend both ways and are genuinely permutation-equivariant. Remove the positional embeddings from an encoder and it becomes a bag-of-words model. The causal=False row above is the proof.
Assuming NoPE is free at scale. The published evidence is at small to moderate scale on controlled tasks. Frontier decoders overwhelmingly use RoPE. Treat NoPE as a well-supported research result, not as a drop-in production recommendation.
Confusing NoPE with removing the causal mask. The mask is what makes NoPE possible. Remove it and there is nothing left.
Expecting a NoPE model to extrapolate for free. It generalises better than the alternatives in controlled comparisons. It still degrades past the training length, as the next lesson measures.
Try it yourself
Reduce the noise scale from 0.5 to 0.05 and rerun the probe. Then raise it to 2.0. The leak depends on there being variation between tokens to average over. Finding the noise level where the probe stops beating the baseline tells you exactly what the mechanism depends on.
What to learn next
- What breaks when you exceed the trained length — all five schemes measured side by side.
- Rotary position embeddings (RoPE) — what production models use instead.
- How LLMs work — where causal masking fits in the bigger picture.
Researcher — Mathematics and papers.
The mechanism
Under a causal mask, position $m$ attends over ${0, \dots, m}$. The attention output is a convex combination of $m+1$ value vectors. If those vectors are drawn independently with mean $\mu$ and covariance $\Sigma$, and attention weights are approximately uniform, the output has
$$ \mathbb{E}[o_m] = \mu, \qquad \mathrm{Cov}[o_m] \approx \frac{\Sigma}{m+1} $$
so $\mathbb{E}\lVert o_m - \mu \rVert^2 \approx \mathrm{tr}(\Sigma)/(m+1)$, decaying as $1/(m+1)$. Any downstream layer that is sensitive to magnitude — layer norm, a ReLU threshold, a quadratic feature in the MLP — can read this. The signal is second-order, which is why the developer probe needs the norm as a feature.
The uniform-weight assumption is a simplification. The general statement is weaker and still sufficient: the number of terms in the convex combination differs by position, and softmax over a differing number of logits produces distributionally different outputs.
Haviv et al. (2022)
Transformer Language Models without Positional Encodings Still Learn Positional Information. Key findings:
- Decoder-only models trained with no positional encoding reach perplexity comparable to models with learned or sinusoidal encodings.
- Probing intermediate representations recovers absolute position with high accuracy, strongest in early layers.
- Removing the causal mask removes the ability entirely, confirming the mask as the source.
Kazemnejad et al. (2023)
The Impact of Positional Encoding on Length Generalization in Transformers, NeurIPS 2023. A controlled comparison of APE, T5 relative bias, ALiBi, RoPE and NoPE on reasoning and algorithmic tasks, trained short and tested long.
Two results matter.
Theoretical. NoPE can express both absolute and relative positional schemes. The construction shows a causal transformer can implement a counting mechanism in the first layer that recovers absolute position, and subsequent layers can compute differences. Expressivity is therefore not the obstacle.
Empirical. Trained with SGD, NoPE's learned attention patterns resemble T5's relative bias rather than an absolute scheme. On their benchmark, NoPE outperformed all four explicit encodings on length generalisation, with absolute position embeddings weakest.
The scope is important. These are decoder-only models at moderate scale on synthetic and semi-synthetic reasoning tasks. They are not frontier-scale natural-language pretraining runs.
Why production models still use RoPE
Several reasons, none of them a refutation.
- The comparison evidence is at small scale. Frontier training runs are expensive and conservative.
- RoPE gives an explicit, controllable handle on context extension. YaRN and its relatives operate on RoPE frequencies. There is no equivalent lever for NoPE.
- The NoPE signal is a side effect and is weaker than a designed encoding at short range, where local syntax lives.
- Ecosystem lock-in: kernels, caches and serving stacks are built around RoPE.
The hybrid: NoPE layers inside a RoPE model
The current practical use is partial. Several recent large models interleave a minority of attention layers with no positional encoding among a majority using RoPE. The reported rationale is that RoPE layers handle local structure while NoPE layers carry an unmodulated long-range path, improving behaviour at long context without giving up short-range precision.
A related and older idea is partial rotary: apply RoPE to a fraction of each head's dimensions and leave the rest untouched. GPT-NeoX used 25 percent, and the transformers config exposes partial_rotary_factor. That is NoPE inside a subspace of every head rather than in whole layers.
Open questions
- Does the NoPE advantage survive at frontier scale and on natural language, or is it specific to controlled algorithmic tasks?
- What is the optimal ratio and placement of NoPE layers in a hybrid stack, and does it depend on target context length?
- The leak is second-order and scales as $1/(m+1)$, so the resolution between positions $m$ and $m+1$ falls off quadratically. What does that imply for precise positional reasoning at very long context?
Papers
- Haviv et al., Transformer Language Models without Positional Encodings Still Learn Positional Information, EMNLP Findings 2022 — arxiv.org/abs/2203.16634
- Kazemnejad et al., The Impact of Positional Encoding on Length Generalization in Transformers, NeurIPS 2023 — arxiv.org/abs/2305.19466
- Chi et al., Latent Positional Information is in the Self-Attention Variance of Transformer Language Models Without Positional Embeddings, ACL 2023 — arxiv.org/abs/2305.13571
What to learn next
- What breaks when you exceed the trained length — all five schemes measured side by side.
- Rotary position embeddings (RoPE) — what production models use instead.
- How LLMs work — where causal masking fits in the bigger picture.