How Models Know Word Order

Stretching RoPE for longer context

A model trained at 4k tokens can be stretched to 128k by slowing down RoPE's dials. Linear interpolation, NTK scaling and YaRN are three ways to do it.

On this page 6
  1. What breaks without it
  2. The three fixes, in plain terms
  3. The part nobody advertises
  4. Where you have seen this
  5. Remember this
  6. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

RoPE scaling slows down the dials, so a model trained on short text can read long text without retraining from scratch.

Think of a measuring tape marked from 0 to 100 centimetres. You need to measure a two-metre wall. One option is to reprint the tape from scratch. The cheaper option is to squash the wall's measurements onto the tape you already own, so every reading is halved.

Both numbers are on a scale you already understand. Nothing is off the end of the tape. You have lost fine detail, and you have gained reach.

That squashing is what RoPE scaling does to positions.

What breaks without it

RoPE turns each token by an angle set by its position. Some dials turn fast, some crawl.

During training on 4,000 tokens, the fast dials spin round hundreds of times. The model sees every possible angle on them, many times over.

The slowest dials never complete a single turn in 4,000 tokens. So the model has only ever seen a narrow slice of their possible angles.

Now feed it 40,000 tokens. Those slow dials swing into angles the model has never encountered in its life. It has no idea what they mean, and quality collapses.

The three fixes, in plain terms

Squash everything. Divide every position by the stretch factor. Position 8000 is treated as position 1000. Every dial now stays inside the range the model knows. The cost lands on nearby tokens. They are now a fraction of a step apart, so they become hard to tell apart.

Slow the dials instead. Rather than squashing positions, turn down every dial's speed by changing one base number. Nearby tokens stay distinguishable. Far-apart tokens get compressed. This is often called NTK-aware scaling.

Treat each dial differently. The fast dials are fine as they are, since the model has seen all their angles. The slow dials are the problem. So leave the fast ones alone, squash the slow ones, and blend smoothly in between. That is YaRN, and it is the method most long-context models use today.

   dial speed:   fast  ------------------------->  slow

   squash all:   [squashed][squashed][squashed][squashed]   loses nearby detail
   slow dials:   [ slower ][ slower ][ slower ][ slower ]   uniform compromise
   YaRN      :   [ leave  ][ blend  ][squashed][squashed]   targeted

The part nobody advertises

None of these are free. Every method trades fine-grained detail near a token for reach across the document.

And every one of them works far better with a short finetuning pass afterwards, on a few billion tokens of long text. Applied with no finetuning at all, the model runs and degrades. The published context number on a model card usually includes that finetuning step.

Where you have seen this

  • A model card saying "4k base, extended to 128k with YaRN".
  • Qwen models where you switch YaRN on in the config for long inputs.
  • Llama 3.1's 128k window, built with a frequency-dependent scaling rule.
  • Any "long context" version of a model released a few months after the original.

Remember this

  • RoPE's slowest dials never complete a turn during training, so long inputs show the model angles it has never seen.
  • Scaling squashes positions, slows the dials, or does both selectively.
  • YaRN treats fast and slow dials differently, and it is the current default choice.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install "transformers==5.6.2" torch numpy

Version matters here. In transformers v4 the config field was rope_scaling. In v5 it is rope_parameters, and config.rope_theta no longer exists as a top-level attribute. Code written against v4 fails on v5 in a way that is easy to misread as a model problem.

Comparing every scheme on real numbers

rope_scaling.py
import numpy as np
import torch
from transformers import LlamaConfig
from transformers.modeling_rope_utils import ROPE_INIT_FUNCTIONS

np.set_printoptions(precision=1, suppress=True)
HEAD_DIM, BASE, TRAINED, FACTOR = 16, 10000.0, 2048, 8.0

def cfg(rope_parameters):
    c = LlamaConfig(hidden_size=HEAD_DIM * 4, num_attention_heads=4,
                    num_key_value_heads=4, num_hidden_layers=2, vocab_size=32,
                    intermediate_size=64, max_position_embeddings=TRAINED)
    c.rope_parameters = rope_parameters       # v5 name. In v4 this was rope_scaling.
    return c

def show(name, rope_parameters, seq_len=None):
    conf = cfg(rope_parameters)
    inv_freq, attn_scale = ROPE_INIT_FUNCTIONS[rope_parameters["rope_type"]](
        conf, device=torch.device("cpu"), seq_len=seq_len)
    lam = 2 * np.pi / inv_freq.numpy()        # wavelength: tokens per full turn
    print(f"{name:<28} attention scale {attn_scale:.4f}")
    print(f"  wavelengths: {lam}")
    return lam

print("trained context:", TRAINED, "  target context:", int(TRAINED * FACTOR))
default = 1.0 / (BASE ** (np.arange(0, HEAD_DIM, 2) / HEAD_DIM))
print(f"{'default (no scaling)':<28} attention scale 1.0000")
print(f"  wavelengths: {2 * np.pi / default}")

show("linear (interpolation)", {"rope_type": "linear", "rope_theta": BASE, "factor": FACTOR})
show("dynamic NTK at 2048", {"rope_type": "dynamic", "rope_theta": BASE, "factor": FACTOR,
                             "original_max_position_embeddings": TRAINED}, seq_len=2048)
show("dynamic NTK at 16384", {"rope_type": "dynamic", "rope_theta": BASE, "factor": FACTOR,
                              "original_max_position_embeddings": TRAINED}, seq_len=16384)
lam_yarn = show("yarn", {"rope_type": "yarn", "rope_theta": BASE, "factor": FACTOR,
                         "original_max_position_embeddings": TRAINED,
                         "beta_fast": 32, "beta_slow": 1})

print("\nfull turns each pair makes inside the ORIGINAL 2048-token window:")
print("  default:", (TRAINED / (2 * np.pi / default)).round(2))
print("  yarn   :", (TRAINED / lam_yarn).round(2))
Output
trained context: 2048   target context: 16384
default (no scaling)         attention scale 1.0000
  wavelengths: [    6.3    19.9    62.8   198.7   628.3  1986.9  6283.2 19869.2]
linear (interpolation)       attention scale 1.0000
  wavelengths: [    50.3    159.     502.7   1589.5   5026.5  15895.3  50265.5 158953.4]
dynamic NTK at 2048          attention scale 1.0000
  wavelengths: [    6.3    19.9    62.8   198.7   628.3  1986.9  6283.2 19869.2]
dynamic NTK at 16384         attention scale 1.0000
  wavelengths: [      6.3      35.4     199.5    1123.8    6331.9   35676.   201009.
 1132543.1]
yarn                         attention scale 1.2079
  wavelengths: [     6.3     19.9     62.8    254.3   1117.    5780.1  50265.5 158953.4]

full turns each pair makes inside the ORIGINAL 2048-token window:
  default: [325.9 103.1  32.6  10.3   3.3   1.    0.3   0.1]
  yarn   : [326.  103.1  32.6   8.1   1.8   0.3   0.    0. ]

Reading that output, line by line

The last row of the default block is the whole problem. Inside 2048 tokens the fastest pair turns 325.9 times and the slowest turns 0.1 times. The model has seen every angle of the first pair and one tenth of a turn of the last. Past 2048 tokens, that last pair enters territory it has never observed.

Linear scaling multiplied every wavelength by exactly 8. The fastest pair went from 6.3 to 50.3. That fixes the slow pairs, and it wrecks the fast ones: adjacent tokens are now only one eighth of a step apart on a dial that used to separate them cleanly. This is why pure linear interpolation needs finetuning to recover.

Dynamic NTK does nothing at 2048 and a great deal at 16384. The two rows are identical to the default at the trained length, then change once the sequence exceeds it. That is the "dynamic" part: it scales only when it has to, so short prompts are untouched. Note it changes the base, so the fastest pair stays at 6.3 while the slowest jumps to over a million.

YaRN's row is the interesting one. The first three wavelengths — 6.3, 19.9, 62.8 — are byte-for-byte the default. The last two — 50265.5 and 158953.4 — are byte-for-byte the linear-scaled values. The middle three are blended. YaRN left the high-frequency pairs completely alone and fully interpolated the low-frequency ones.

The attention scale of 1.2079. YaRN also divides the attention logits by a temperature. That value is not arbitrary: the paper's empirical fit is sqrt(1/t) = 0.1 * ln(s) + 1, and 0.1 * ln(8) + 1 = 1.2079. It compensates for entropy growth when more tokens compete in one softmax.

Choosing between them

MethodChangesNeeds finetuningShort prompts affected
linear (PI)all positions divided by factoryes, stronglyyes
dynamic NTKbase, only past the trained lengthworks without, better withno
yarnper-pair, plus attention temperatureyes, but far less datano
llama3per-pair, low and high frequency cutoffsdone at pretrainingno

Common mistakes

Using rope_scaling on transformers v5. The field is rope_parameters now. Setting the old name silently does nothing, and you get an unscaled model that produces degraded long-context output with no error message.

Leaving YaRN enabled for short prompts. Static YaRN compresses positions at every length, including a 50-token prompt. Qwen's own documentation warns about this. Enable it when you need the length, or use the dynamic variant.

Editing max_position_embeddings and nothing else. That number gates checks and buffer sizes. Raising it without changing rope_parameters gives you a model that accepts long inputs and reasons badly over them.

Expecting a benchmark number without finetuning. Every published context extension includes continued pretraining on long documents. YaRN's headline claim is that it needs 10 times fewer tokens and 2.5 times fewer steps than earlier methods, not that it needs none.

Testing only with needle-in-a-haystack. Retrieving one planted fact is far easier than reasoning over a whole long document. A model can pass the needle test at 128k and still be unusable for summarising a 100k-token report.

Try it yourself

Set FACTOR to 32 and rerun. Watch how many of YaRN's leading wavelengths stay untouched. Then change beta_fast from 32 to 8 and see the boundary move. Those two knobs are the entire tuning surface of YaRN, and printing wavelengths is a faster way to understand them than reading the equations cold.

What to learn next

Researcher — Mathematics and papers.

The failure being repaired

For RoPE with base $b$ and head dimension $d$, pair $i$ has angular frequency $\theta_i = b^{-2i/d}$ and wavelength $\lambda_i = 2\pi b^{2i/d}$. Define the number of full rotations completed within training context $L$:

$$ r(i) = \frac{L}{\lambda_i} $$

Pairs with $r(i) \gg 1$ are fully exercised during training. Pairs with $r(i) < 1$ present, at inference beyond $L$, phases outside the training support. Direct extrapolation therefore fails not because RoPE is undefined past $L$ — it is perfectly well defined — but because the low-frequency subspace is out of distribution.

Position Interpolation

Chen et al. (2023) rescale positions rather than extrapolate:

$$ f'(x, m) = f!\left(x, \frac{mL}{L'}\right), \qquad s = \frac{L'}{L} $$

Equivalently $\theta_i \mapsto \theta_i / s$, so every wavelength is multiplied by $s$. All phases remain inside the trained range. The paper extends Llama to 32k with 1000 finetuning steps.

The cost is uniform. High-frequency pairs, which encode local ordering, are compressed by $s$ as well, so adjacent tokens become $s$ times closer in phase. This crowds exactly the signal the model relies on for local syntax.

NTK-aware scaling

Proposed by the pseudonymous bloc97 (2023) on the LocalLLaMA forum and analysed in the YaRN paper. Instead of scaling positions, scale the base:

$$ b' = b \cdot s^{\frac{d}{d-2}} $$

This choice makes the lowest-frequency pair receive scaling $s$ while the highest-frequency pair is left essentially untouched, with a smooth geometric progression in between. It does not require finetuning to give usable results, which is why it appeared first in community deployments.

The "dynamic" variant recomputes $b'$ from the current sequence length, using $s = \max(1, \ell / L)$, so short sequences run with the original base. That is exactly what the developer output shows: identical wavelengths at $\ell = 2048$, changed wavelengths at $\ell = 16384$.

YaRN

Peng, Quesnelle, Fan and Shippole (2023) combine two ideas.

NTK-by-parts interpolation. Define the ratio

$$ r(i) = \frac{L}{\lambda_i} = \frac{L}{2\pi b^{2i/d}} $$

and a ramp

$$ \gamma(r) = \begin{cases} 0 & r < \alpha \ 1 & r > \beta \ \dfrac{r - \alpha}{\beta - \alpha} & \text{otherwise} \end{cases} $$

Then interpolate per pair:

$$ h(\theta_i) = \big(1 - \gamma(r(i))\big)\frac{\theta_i}{s} + \gamma(r(i))\,\theta_i $$

Pairs completing many rotations ($r > \beta$) keep $\theta_i$ unchanged. Pairs completing fewer than $\alpha$ rotations get full interpolation $\theta_i / s$. Between them the ramp blends linearly. For Llama models the paper recommends $\alpha = 1$, $\beta = 32$; these are beta_slow and beta_fast in the transformers config, defaulting to 1 and 32.

This is directly visible in the developer output: three wavelengths equal to the default, two equal to the fully interpolated values, three intermediate.

Attention temperature. Divide the pre-softmax logits by $t$, with the empirical fit

$$ \sqrt{1/t} = 0.1 \ln(s) + 1 $$

At $s = 8$ this gives $1.2079$, matching the printed attention scale. The motivation is entropy: with $s$ times more tokens competing in one softmax, attention flattens, and a temperature restores sharpness. It costs nothing at runtime because it can be folded into the precomputed $\cos$ and $\sin$ tables.

The paper reports YaRN reaching the same perplexity as prior methods with roughly 10 times fewer tokens and 2.5 times fewer training steps.

Llama 3's rule

Llama 3.1 uses a frequency-cutoff scheme, exposed as rope_type: "llama3" with low_freq_factor and high_freq_factor. Wavelengths shorter than a high-frequency threshold are untouched, those longer than a low-frequency threshold are divided by factor, and the band between is smoothly interpolated. Structurally this is the same idea as NTK-by-parts, parameterised by wavelength thresholds rather than by rotation counts. It was applied during pretraining, alongside a base of 500000, rather than as a post-hoc extension.

LongRoPE

Ding et al. (2024) search for a per-dimension rescaling vector with an evolutionary algorithm rather than deriving it, and use a two-stage schedule to reach 2M tokens. Transformers exposes it as rope_type: "longrope" with explicit short_factor and long_factor lists, one entry per pair. It is used in the Phi model family.

Evaluation honestly

Perplexity on long documents is the standard headline and is a weak proxy. Retrieval-style probes such as needle-in-a-haystack measure a narrow skill. Hsieh et al. (2024), RULER, construct tasks with controllable difficulty and find that models advertising 128k or more frequently sustain far shorter effective context, with several dropping below their claimed length by a wide margin. Treat an advertised context length as an input limit, not a capability claim.

Papers

What to learn next