Fast Attention and Long Context

Advertised context vs usable context

A model's headline context length is what it will accept, not what it can use, and the gap between the two is large, measurable and rarely mentioned on the model card.

On this page 9
  1. The short answer
  2. The phone number in the thick notebook
  3. What the number on the page means
  4. What actually happens
  5. Why the advertised number is often optimistic
  6. How to check for yourself
  7. Why this matters for money too
  8. Remember this
  9. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

The context length on a model's page is how much text it will accept. It is not how much text it can actually use.

The phone number in the thick notebook

You wrote a phone number in a notebook six months ago. The number is definitely in there. The notebook has not lost a page.

Whether you find it is another matter. If you wrote it on the last page you used, you will find it in seconds. If it is somewhere in the middle of three hundred pages, you might not find it at all.

The notebook holding the number and you being able to retrieve it are two separate facts. Models are the same.

What the number on the page means

A model card says 128,000 tokens. That means the software will accept an input of that size without an error.

It is a capacity claim, like a lift saying "eight persons". Eight people fit. Whether eight people are comfortable is a different question that the sign does not answer.

What actually happens

Put a fact near the end of a long document and the model uses it easily. Push the same fact towards the middle and performance drops, sharply and measurably.

The information has not been deleted. Every word is still in the input. The model's ability to reach in and pick out the right one is what fades.

There is a shape to it that has been documented repeatedly. Models are strong at the beginning of the input, strong at the end, and weakest in the middle. It is sometimes called "lost in the middle".

Why the advertised number is often optimistic

Many long-context models were not trained on long text throughout. They were trained on shorter text, then stretched afterwards with a mathematical adjustment to how positions are counted.

The adjustment lets a model accept far more text than it ever practised on. Sometimes that works well. Often the model behaves noticeably worse past the length it actually trained on.

You can read this off the settings file of open models. Several popular ones declare a much shorter original length, a stretching factor, and the large advertised figure.

How to check for yourself

Do not trust the number and do not trust a single simple test. The old "hide a sentence and ask for it back" test is passed by nearly everything now. Passing it means little.

Test what you actually need:

  • Find several facts, not one.
  • Combine two facts from far apart.
  • Count or summarise across the whole document.
  • Answer when the fact is absent, without inventing it.

Run each at several lengths and several positions. The point where your accuracy falls below what you can accept is your real context length, for your task.

   advertised    |=================================| 128,000
   trained on    |==========|                        8,000
   your task     |================|                 30,000 (measured)
                             ^ where accuracy started falling

Why this matters for money too

Long inputs cost more and run slower. Paying for a hundred thousand tokens the model then uses poorly is the most expensive mistake here.

Often, finding the right five thousand tokens and sending only those works better and costs a twentieth. That is what RAG is for, and it is why a long context window did not make retrieval obsolete.

Remember this

  • Advertised context is what the model accepts. Usable context is what it can find things in.
  • Accuracy falls with distance, and the middle of a long input is the weakest place.
  • Measure it on your own task. Do not accept the number on the page.

What to learn next

  • What is RAG? — finding the right passage instead of hoping the model does.
  • Context window — what the window is and how tokens are counted.
  • Hallucination — what a model does when it cannot find what it needs.

Developer — Code and libraries.

Setup

bash
pip install torch transformers

Written against PyTorch 2.5.1 and transformers 5.6.2, Python 3.10. Runs on CPU in about a minute after GPT-2 downloads (roughly 500 MB).

GPT-2 has a 1024-token context, which is small enough to test end to end on a laptop, and it shows the effect plainly. The same script works on a long-context model by changing one string, at proportionally more compute.

Measuring retrieval, not vibes

needle.py
import math
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

name = "gpt2"
tok = AutoTokenizer.from_pretrained(name)
model = AutoModelForCausalLM.from_pretrained(name).eval()
MAXLEN = model.config.n_positions          # gpt2's hard limit
print(f"model: {name} | trained context: {MAXLEN} tokens | torch: {torch.__version__}")

needle = " The secret code is 4729."
filler = (" The road was quiet and the shops were closed. A dog slept under a parked van."
          " Rain had washed the dust from the leaves. Someone was frying onions upstairs.")
question = " The secret code is"

filler_ids = tok(filler, return_tensors="pt").input_ids[0]
needle_ids = tok(needle, return_tensors="pt").input_ids[0]
q_ids      = tok(question, return_tensors="pt").input_ids[0]
answer_ids = tok(" 4729", return_tensors="pt").input_ids[0]
print(f"needle tokens: {len(needle_ids)} | answer tokens: {len(answer_ids)} "
      f"| filler block: {len(filler_ids)}")

def build(total_filler_blocks, depth_fraction):
    blocks = [filler_ids] * total_filler_blocks
    pos = int(depth_fraction * total_filler_blocks)
    seq = blocks[:pos] + [needle_ids] + blocks[pos:] + [q_ids]
    return torch.cat(seq).unsqueeze(0)

def answer_logprob(ids):
    with torch.no_grad():
        full = torch.cat([ids, answer_ids.unsqueeze(0)], dim=1)
        logits = model(full).logits[0, ids.shape[1] - 1:-1]
        lp = torch.log_softmax(logits.float(), -1)
        return lp[torch.arange(len(answer_ids)), answer_ids].sum().item()

control = torch.cat([filler_ids] * 8 + [q_ids]).unsqueeze(0)   # same text, needle removed
cl = answer_logprob(control)
print(f"\ncontrol, no needle anywhere: log P = {cl:.3f}   P = {math.exp(cl):.6f}")

print(f"\n{'blocks':>7}{'ctx len':>9}{'depth':>8}{'distance':>10}{'log P':>9}{'P':>9}")
for blocks in (1, 4, 8, 16, 24):
    for depth in (0.0, 0.5, 1.0):
        ids = build(blocks, depth)
        if ids.shape[1] + len(answer_ids) > MAXLEN:
            print(f"{blocks:>7}{ids.shape[1]:>9}{depth:>8.1f}{'-':>10}{'past the limit':>18}")
            continue
        dist = ids.shape[1] - (int(depth * blocks) * len(filler_ids) + len(needle_ids))
        lp = answer_logprob(ids)
        print(f"{blocks:>7}{ids.shape[1]:>9}{depth:>8.1f}{dist:>10}{lp:>9.3f}{math.exp(lp):>9.4f}")
Output
model: gpt2 | trained context: 1024 tokens | torch: 2.5.1+cu121
needle tokens: 7 | answer tokens: 2 | filler block: 33

control, no needle anywhere: log P = -16.337   P = 0.000000

 blocks  ctx len   depth  distance    log P        P
      1       44     0.0        37   -2.768   0.0628
      1       44     0.5        37   -2.768   0.0628
      1       44     1.0         4   -2.021   0.1325
      4      143     0.0       136   -1.755   0.1728
      4      143     0.5        70   -0.778   0.4591
      4      143     1.0         4   -0.134   0.8743
      8      275     0.0       268   -2.415   0.0894
      8      275     0.5       136   -1.021   0.3601
      8      275     1.0         4   -0.121   0.8860
     16      539     0.0       532   -3.050   0.0474
     16      539     0.5       268   -1.024   0.3591
     16      539     1.0         4   -0.105   0.9005
     24      803     0.0       796   -2.763   0.0631
     24      803     0.5       400   -1.204   0.3001
     24      803     1.0         4   -0.158   0.8541

Reading the output

The control is the row that makes the rest meaningful. With no needle anywhere, the model assigns probability 0.000000 to "4729". Everything above that is genuine retrieval, not a lucky guess.

Distance 4: probability around 0.87 to 0.90. When the fact is immediately before the question, the model is near certain. That is the number a demo would quote.

Distance 796: probability 0.063. Same fact, same wording, same model, fully inside the context window. Retrieval has fallen by a factor of about fourteen, purely because of distance.

Sort the whole table by distance and it is close to monotonic. 4 tokens gives 0.85 to 0.90; 70 to 140 gives 0.17 to 0.46; 268 to 796 gives 0.05 to 0.09. Context length is not the variable. Distance is.

Row one is the honest anomaly. At 44 tokens with depth 0.0 the probability is 0.0628, worse than longer contexts at the same depth. With one filler block, depths 0.0 and 0.5 build the same sequence, and the surrounding text is too short for the model to settle. Small-context artefacts are real; report them rather than deleting the row.

Reading the advertised number off a config

The stretching is visible in plain text. These are live values from the Hub:

Modelmax_position_embeddingsoriginal_max_position_embeddingsscaling
Llama-3.1-8B1310728192llama3, factor 8.0
gpt-oss-20b1310724096yarn, factor 32.0
DeepSeek-V31638404096yarn, factor 40
Qwen3-30B-A3B40960—none declared
python
from transformers import AutoConfig
c = AutoConfig.from_pretrained("openai/gpt-oss-20b")
print(c.max_position_embeddings, c.rope_parameters)

Three of these four reach their headline figure by scaling rotary positions well past the length in original_max_position_embeddings. That field is not a promise about training data, and long-context continued pre-training usually follows the scaling. It is a strong hint about where to concentrate your testing.

Qwen3-30B-A3B declares no scaling and a smaller window. A smaller honest number can be worth more than a larger stretched one.

Building a real evaluation

Four task families, in rising difficulty, matching what RULER measures:

  1. Single retrieval. One fact, asked back. Nearly every current model passes. Use it as a smoke test, never as evidence.
  2. Multi-key or multi-value retrieval. Several facts hidden, several asked for. This is where models start to separate.
  3. Multi-hop tracing. Fact A points at fact B, which points at the answer. Distance between hops is the difficulty knob.
  4. Aggregation. Count occurrences, or list every item matching a condition. Needs attention over the whole input rather than one spot.

Add a refusal control: run the same prompt with the fact removed and check the model says it does not know. A model that scores well on retrieval and also confidently invents an answer when the fact is absent is not usable.

Report accuracy as a grid of length by depth. A single averaged number hides the shape, and the shape is the finding.

Common mistakes

Reporting only single-needle results. The easiest test in the family, saturated for years, and still the one in most marketing.

Testing at one depth. Depth effects are as large as length effects. Sample at least beginning, quarter, middle, three-quarter and end.

Using a synthetic haystack that a real task will not resemble. Repeated filler is easy to attend past. Real documents are full of near-duplicates of the needle, which is much harder.

Forgetting cost. At 128k tokens, prefill cost and time-to-first-token both grow. A method that keeps accuracy while doubling latency may lose on the metric your users care about.

Concluding retrieval is obsolete because context is long. These measurements are the strongest argument for RAG. Finding the right 5k tokens beats hoping the model finds them inside 128k.

Try it yourself

Change filler to text that resembles the needle — several other sentences containing numbers. Re-measure. The drop is much steeper, and that is the gap between a synthetic benchmark and a real document.

What to learn next

  • What is RAG? — finding the right passage instead of hoping the model does.
  • Context window — what the window is and how tokens are counted.
  • Hallucination — what a model does when it cannot find what it needs.

Researcher — Mathematics and papers.

The benchmarks

Hsieh et al. (2024), RULER: What's the Real Context Size of Your Long-Context Language Models? (arXiv:2404.06654) is the standard reference. It generates synthetic tasks at configurable length and complexity across four categories: retrieval (multi-needle variants), multi-hop tracing, aggregation, and question answering, with 2,600 examples per length at 4k, 8k, 16k, 32k, 64k and 128k.

Its headline finding is the one to internalise: despite near-perfect accuracy on vanilla needle-in-a-haystack, "almost all models exhibit large performance drops as the context length increases". RULER defines effective context as the largest length at which a model still beats a strong short-context baseline, and that figure is routinely a fraction of the advertised one.

Liu et al. (2023), Lost in the Middle (arXiv:2307.03172) documented the U-shaped position curve: accuracy is highest when relevant information sits at the start or end of the input, and measurably lowest in the middle. The practical consequences are direct — order retrieved passages so the most relevant sit at the edges, and treat "increase k" as a change that can reduce accuracy.

Kuratov et al. (2024), BABILong (arXiv:2406.10149) embeds bAbI reasoning tasks in long natural distractor text, separating retrieval from reasoning-over-retrieved-content. Models that retrieve well can still fail to reason over what they retrieved.

An et al. (2024), Why Does the Effective Context Length of LLMs Fall Short? (arXiv:2410.18745) attributes part of the gap to a left-skewed distribution of relative positions in training data: long-range position pairs are rare during pre-training, so the corresponding rotary frequencies are undertrained.

Where the gap comes from

Four contributing mechanisms, each with a different remedy.

Position extrapolation. RoPE scaling (Position Interpolation, NTK-aware scaling, YaRN) makes long positions representable. Representable is not trained. gpt-oss-20b declares original_max_position_embeddings: 4096 with a YaRN factor of 32; DeepSeek-V3 declares 4096 with factor 40. The frequency bands that only activate beyond the original length see far less data.

Attention dilution. Softmax distributes a fixed unit of weight over $N$ keys. As $N$ grows, the weight available to the correct key falls unless logits sharpen proportionally. Related to why attention sinks form.

Training data. Documents long enough to require 128k of genuine dependency are rare. Much long-context training data is concatenated shorter documents, which teaches length tolerance rather than long-range dependency.

Architectural shortcuts. Sliding windows, sparse patterns and fixed-size recurrent states all trade long-range fidelity for cost. A hybrid's effective context can be well below its advertised context by construction, and aggregate perplexity will not reveal it.

Evaluation protocol

  • Grid, not scalar. Report length x depth. Averaging destroys the U-shape you were looking for.
  • Include a no-needle control. Distinguishes retrieval from prior. The GPT-2 experiment above shows why: without the control, 0.06 looks like failure rather than a fourteen-fold degradation from 0.87.
  • Use log-probability of the gold continuation, not exact-match generation. Exact match is a threshold on a continuous quantity and throws away the gradient of degradation.
  • Adversarial distractors. Needles that resemble the haystack are the realistic case and the discriminating one.
  • Report cost. Effective context that costs 40 seconds to first token is a different product from one that costs 4.
  • Test the deployed configuration. Quantised KV cache, cache eviction policies and sliding-window enforcement all move this curve, and none of them appear in the model card.

The standing conclusion

Context lengths advertised in 2026 are two to three orders of magnitude above 2020, and effective context has not kept pace. The honest position is that headline context is a compatibility statement — the model will not error — and that usable context is a per-model, per-task empirical quantity that you must measure.

The practical corollary is unglamorous and stable: retrieval remains the right default. Long context is a tool for cases where the relevant material genuinely cannot be narrowed, not a replacement for narrowing it.

What to learn next

  • What is RAG? — finding the right passage instead of hoping the model does.
  • Context window — what the window is and how tokens are counted.
  • Hallucination — what a model does when it cannot find what it needs.