The BERT Family

Longformer and long-context encoders

Longformer reads up to 4,096 tokens by making most words only attend to their nearby neighbours instead of every other word, trading a little global awareness for a context window eight times larger than BERT's.

Read these first

On this page 5
  1. Why it exists
  2. How it works
  3. Where you have already seen it
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Longformer reads up to 4,096 tokens. Most words look only at their close neighbours, instead of every other word in the document. That trade buys a much longer reading window.

Think about proofreading a long report one paragraph at a time. You mostly check each sentence against the ones right next to it. Occasionally you flip back to the executive summary at the top, to keep the document's point in mind. You are not comparing every sentence against every other sentence — that would take forever. Reading locally, with occasional checks against a few anchor points, gets most of the value at a fraction of the effort.

Longformer's attention works the same way. Most words attend only to a nearby window of neighbours. A small number of specially marked words get full, unrestricted attention to the whole document. They play the role of that executive summary you keep checking back against.

Why it exists

BERT's full attention lets every word look at every other word. The previous lesson covered why that hits a hard wall at 512 tokens — the position table runs out entirely. Even setting the token-limit problem aside, full attention has a deeper cost problem. Comparing every word to every other word costs compute that grows with the square of the document length. Doubling the document length quadruples the attention cost. This alone would make a naive 4,096-token BERT sixty-four times more expensive per layer than the 512-token original.

Longformer, released in 2020, tackles both problems together. It extends the position table far past 512. It restricts most attention to a local window, so cost grows roughly linearly with document length instead of quadratically. That makes a much longer context window computationally realistic.

How it works

  BERT:  every word attends to every other word
         (n words -> n x n comparisons -> cost grows with the SQUARE of length)

  Longformer:  most words attend only to a nearby window
               (n words -> roughly n x window comparisons -> cost grows
                roughly LINEARLY with length)

               A few specially marked "global attention" words --
               often only the [CLS] token, or a question in a QA task --
               still attend to, and are attended to by, EVERY other word.

  Local window:        [-----window-----] word [-----window-----]
  Global attention:    [CLS] can see every word, every word can see [CLS]

The result is a middle ground. It is not quite as thorough as BERT's full attention at every layer. But it reads documents eight times longer than BERT's hard limit, at a computational cost that scales far more gently as documents grow.

Where you have already seen it

  • Legal and academic document search. Full contracts or research papers need to be read as one unit, not chopped into six separate BERT-sized chunks that lose track of each other.
  • Long-document summarisation systems, where the model genuinely needs to see the beginning and end of a document at once, to summarise it coherently.
  • Question answering over long documents, where the answer might depend on a detail from several pages earlier than the question-relevant passage.

Remember this

  • Longformer trades full attention (every word sees every word) for local attention (most words see only nearby neighbours) plus a small set of globally-attending words.
  • This buys a much longer context window — 4,096 tokens versus BERT's 512 — at a computational cost that scales roughly linearly instead of quadratically.
  • It is one of several answers to the same problem the previous lesson introduced; ModernBERT is a more recent one, using a different mechanism.

What to learn next

  • The 512-token wall — the exact limitation this lesson's architecture was built to remove.
  • ModernBERT — a newer, differently-engineered answer to the same long-context problem.
  • Fast attention and long context — the broader family of techniques Longformer's local-plus-global attention belongs to.

Developer — Code and libraries.

Longformer is loaded through the standard transformers API. The checkpoint used here, allenai/longformer-base-4096, is a genuinely large download (roughly 600 MB) — worth knowing before running this yourself.

Setup

bash
pip install transformers torch

Minimal runnable code: config and tokenizer

longformer_config.py
from transformers import AutoConfig, AutoTokenizer

cfg = AutoConfig.from_pretrained("allenai/longformer-base-4096")
print("max_position_embeddings:", cfg.max_position_embeddings)
print("attention_window:", cfg.attention_window[:3], "... (one value per layer)")
print("num_hidden_layers:", cfg.num_hidden_layers)

tok = AutoTokenizer.from_pretrained("allenai/longformer-base-4096")
long_text = "The Chennai Metro Rail network connects the airport to the city. " * 60
enc = tok(long_text, return_tensors="pt")
print("\nTokens in the same long text from the previous lesson:", enc["input_ids"].shape[1])
print("(BERT's hard limit on this exact text was 512 -- this text has no trouble here)")
Output
max_position_embeddings: 4098
attention_window: [512, 512, 512] ... (one value per layer)
num_hidden_layers: 12

Tokens in the same long text from the previous lesson: 723
(BERT's hard limit on this exact text was 512 -- this text has no trouble here)

The exact same repeated sentence that produced a hard RuntimeError against BERT in the previous lesson tokenizes and fits comfortably here — 723 tokens against a 4,098-token limit, no truncation needed.

An honest note on loading the full model weights

longformer_forward.py
from transformers import AutoModel
import torch

try:
    model = AutoModel.from_pretrained("allenai/longformer-base-4096")
    with torch.no_grad():
        out = model(**enc)
    print("Forward pass output shape:", out.last_hidden_state.shape)
except Exception as e:
    print(f"Could not load weights here: {type(e).__name__}: {e}")
Output
Could not load weights here: ValueError: Due to a serious vulnerability issue in
`torch.load`, even with `weights_only=True`, we now require users to upgrade
torch to at least v2.6 in order to use the function. This version restriction
does not apply when loading files with safetensors.
See the vulnerability report here https://nvd.nist.gov/vuln/detail/CVE-2025-32434

The exact wording of this error is tied to your installed torch/transformers versions and may change slightly over time — the underlying cause (an old, non-safetensors checkpoint plus a version-gated security check) is what matters, not this precise sentence.

This is a genuinely useful failure to see. allenai/longformer-base-4096 is an older checkpoint on the Hugging Face Hub that predates the now-standard .safetensors weight format, shipping only an older-style pytorch_model.bin file instead. Recent transformers versions refuse to load such files unless PyTorch is at least version 2.6, for a real security reason explained in the error message — arbitrary code execution risk in older PyTorch checkpoint loading. Config and tokenizer loading, shown above, are unaffected by this check; only full model-weight loading hits it. The fix is pip install -U torch, or choosing a checkpoint mirror that ships .safetensors weights instead. This is the same underlying issue covered in the DeBERTa lesson's common mistakes — worth recognising as one recurring class of problem, not two unrelated ones, since it affects any older, pre-safetensors checkpoint on the Hub.

Line by line

attention_window is a list with one entry per transformer layer, here 512 at every layer shown — this is the local window size from the beginner block, in practice: each token attends to 512 tokens total (256 on either side, roughly), not the full 4,096-token sequence, at most layers.

max_position_embeddings: 4098, not exactly 4096 — Longformer, like the original BERT, reserves a couple of extra position slots for special tokens (<s> and </s>), following the same convention its RoBERTa-based initialisation inherited.

Common mistakes

Assuming "4,096-token context" means every layer sees the full 4,096 tokens for every word. As the config shows directly, most attention stays local — a real trade-off, not a limitation-free upgrade over BERT's full attention.

Not budgeting for the download size. At roughly 600 MB, allenai/longformer-base-4096 is a meaningfully bigger download than the distilbert-base-uncased checkpoint used throughout the rest of this section — worth checking available disk space and bandwidth before running this on a constrained machine.

Forgetting to mark global attention positions for tasks that need them. For question answering specifically, the question tokens should typically be marked for global attention (via a global_attention_mask argument), so every other token in the document can see the question directly — without this, the model defaults to local-only attention everywhere, and may miss connections between the question and a distant relevant passage.

Try it yourself

Compare cfg.max_position_embeddings across every model this section has covered — bert-base-uncased (512), roberta-base (514), microsoft/deberta-v3-xsmall (512), answerdotai/ModernBERT-base (8,192), and allenai/longformer-base-4096 (4,098). The spread across models released between 2018 and 2024 is a compact history of this one specific problem being tackled from several different angles.

What to learn next

  • ModernBERT — a more recent, differently-engineered answer to the same long-context problem, with a longer window still.
  • The 512-token wall — the specific limitation this lesson's architecture addresses.
  • Fast attention and long context — the wider family of sparse and windowed attention techniques Longformer belongs to.

Researcher — Mathematics and papers.

The attention pattern, formally

Longformer (Beltagy, Peters & Cohan, 2020) combines two attention patterns within the same model:

text
Local (sliding window) attention:
  token i attends to tokens { i - w/2, ..., i + w/2 }    for window size w

Global attention (a small, task-chosen set G of positions):
  token i in G attends to ALL n tokens, and is attended to by all n tokens
  • w is the attention window size, configurable per layer (the checkpoint inspected in the developer block uses w = 512 at every layer).
  • G is typically small relative to n — the [CLS]/<s> token by default, plus task-specific positions (all question tokens, for extractive QA) chosen based on which task Longformer is fine-tuned for.

Complexity for local attention alone is O(n * w), linear in sequence length n for fixed window w — a direct improvement over full attention's O(n^2). Global attention adds O(|G| * n), which stays small in practice because |G| is kept small by design; total complexity is O(n * w + |G| * n) = O(n * (w + |G|)), still linear in n.

Why global attention is necessary, not optional

Pure local attention with no global component has a real limitation: information can only propagate w/2 positions per layer, since that is the furthest any single attention operation reaches. A fact at token 100 and a fact at token 4,000 can only interact after enough layers for local attention windows to chain together across the full distance — at 12 layers and w = 512, the effective receptive field is on the order of 12 * 256 = 3,072 positions, which may or may not cover a 4,000-token gap depending on layer count and window size. Global attention tokens sidestep this entirely: any token marked global reaches every other token in a single layer, and every other token reaches it in a single layer too, giving the network at least one direct, single-hop path for connecting distant information, regardless of local window size or layer depth.

Longformer's initialisation strategy

Rather than pretraining from scratch — expensive, and unnecessary if most of the network's learned behaviour transfers — Longformer initialises its weights from a pretrained RoBERTa checkpoint. Its local attention window uses the same attention mechanism RoBERTa already learned; only the position embedding table needs extending, from RoBERTa's 512 positions to 4,096, done by copying the existing 512-position embeddings repeatedly to fill the extended table, then continuing MLM pretraining briefly on long documents to let the model adapt the newly-extended positions. This is a considerably cheaper path to a long-context model than pretraining an equivalent model from random initialisation, and the same "initialise from a shorter-context checkpoint, then extend" strategy reappears in later long-context adaptation work across the field, including several LLM context-extension techniques.

Comparison with other sparse-attention long-context designs

Sparse Transformers (Child et al., 2019) and BigBird (Zaheer et al., 2020) combine similar local-window attention with different supplementary patterns — BigBird adds both global tokens (as Longformer does) and a set of random attention connections, arguing the combination provides a stronger theoretical guarantee of information flow across the whole sequence than local-plus-global alone. Empirically, the gap between these closely related sparse-attention designs on downstream benchmarks tends to be small relative to the gap between any of them and full attention on short sequences — the family shares its core trade-off, and differs mainly in secondary design choices layered on top of it.

Key references

  • Beltagy, I., Peters, M. & Cohan, A. (2020). Longformer: The Long-Document Transformer. arXiv:2004.05150
  • Zaheer, M. et al. (2020). Big Bird: Transformers for Longer Sequences. arXiv:2007.14062
  • Child, R. et al. (2019). Generating Long Sequences with Sparse Transformers. arXiv:1904.10509
  • Liu, Y. et al. (2019). RoBERTa: A Robustly Optimized BERT Pretraining Approach. arXiv:1907.11692 — the checkpoint Longformer initialises from.

Current state and open problems

Sparse and windowed attention, as a family, has been substantially displaced at the very largest scale by exact, hardware-optimised full attention — FlashAttention made full O(n^2) attention fast enough in practice, via better memory access patterns rather than a smaller asymptotic complexity, that many current long-context models use full attention directly at large context lengths instead of approximating it away. Longformer's specific approach nonetheless remains a genuinely useful, well-understood technique, and its local-plus-global attention pattern reappears as a component inside newer architectures — including ModernBERT's alternating local/global layer design, which borrows the same core idea (most layers local, some layers global) while combining it with rotary position embeddings rather than Longformer's extended absolute position table. Whether sparse attention patterns like Longformer's or exact-but-optimised full attention will dominate long-context model design longer term remains genuinely unsettled, with both approaches actively used across different current architectures.

What to learn next

  • ModernBERT — a current architecture combining a similar local/global pattern with rotary position embeddings instead.
  • Fast attention and long context — the fuller landscape of techniques, including FlashAttention, this lesson's approach sits alongside.
  • The 512-token wall — the problem this entire lesson exists to solve.

What to learn next

These follow on from what you just read.

  • Sequence Labelling and Structure

    Part-of-speech tagging

    Part-of-speech tagging labels every word with its grammatical role in that exact sentence, which is the first structural clue almost every later NLP step depends on.

  • Sequence Labelling and Structure

    BIO and BILOU tagging schemes

    BIO and BILOU are ways of writing down where an entity starts, continues and ends using one tag per token, which is what turns entity spans into something a per-token classifier can learn.

  • Sequence Labelling and Structure

    Conditional random fields

    A conditional random field tags a whole sequence at once instead of one word at a time, so it can rule out impossible tag combinations like a sentence ending mid-entity.