How Models Know Word Order

ALiBi

ALiBi adds no vectors at all. It subtracts a penalty from the attention score that grows with distance, giving each head its own sense of how far back to look.

On this page 7
  1. Why anyone would do it this way
  2. How it works
  3. The clever part: every head gets a different rate
  4. Where this shows up
  5. What is honestly hard here
  6. Remember this
  7. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

ALiBi does not encode position at all. It penalises attention that reaches far back, and the penalty grows with distance.

Think of standing in a noisy railway platform crowd. You hear the person beside you without effort. Someone ten metres away is faint. Someone at the far end is inaudible.

Nobody handed you a distance chart. Sound fades with distance on its own, and you use that fade to work out who is near.

ALiBi does the same to attention. Nothing gets stamped, added or rotated. Distant tokens are turned down, and the model reads distance from how loud each one is.

Why anyone would do it this way

Every other method builds a representation of position and hopes attention learns to use it. ALiBi skips the representation and edits attention directly.

The payoff is length. There is no table to run out of and no dial that spins into unseen angles. Distance 5000 gets a penalty of five thousand times the rate, computed by one multiplication.

The name is short for attention with linear biases — a straight-line penalty applied to the attention scores.

How it works

Before attention decides who to listen to, it produces a score for every pair of tokens. ALiBi subtracts an amount proportional to the distance from each score.

   token 10 looking back:

   at token  9  (1 back)   score - 1 x rate
   at token  8  (2 back)   score - 2 x rate
   at token  5  (5 back)   score - 5 x rate
   at token  0 (10 back)   score - 10 x rate

   then softmax turns those into attention weights

One number per head, called the slope, sets the rate. That is the whole method.

The clever part: every head gets a different rate

If all heads had the same penalty, the model would have one fixed attention span. Instead the slopes are spread out geometrically.

The first head gets a steep penalty and can barely see past a handful of tokens. The last head gets a penalty so gentle it is almost ignoring distance entirely.

So one head handles "the word right before this one" and another handles "somewhere much earlier in the document". The model gets a range of spans for free, without learning any of them.

Where this shows up

  • BLOOM, the 176-billion-parameter multilingual model, uses ALiBi throughout.
  • MPT models from MosaicML, marketed on their ability to accept very long inputs.
  • Baichuan 2's larger model.
  • Several speech and audio transformers, where inputs vary hugely in length.

What is honestly hard here

ALiBi's headline claim is length extrapolation, and the claim is narrower than it sounds. Perplexity stays low on much longer inputs. That is a measure of how well the model predicts the next token, and a recency-biased model does well on it.

Tasks needing a fact from far back are different. The penalty that makes ALiBi extrapolate also makes distant tokens quiet. Later work found ALiBi weak at long-range retrieval for exactly that reason. The design has a built-in preference and it costs something.

Remember this

  • ALiBi subtracts a penalty from attention scores, proportional to distance.
  • Each head gets its own penalty rate, so the model gets many attention spans for free.
  • It handles longer inputs without changes, but it is biased toward nearby tokens by construction.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install numpy

The whole method

alibi.py
import math
import numpy as np
np.set_printoptions(precision=3, suppress=True)

def alibi_slopes(n_heads):
    """One fixed penalty rate per head. Never trained."""
    def geometric(n):
        start = 2 ** (-8 / n)                       # ratio and first term are both 2^(-8/n)
        return [start ** (i + 1) for i in range(n)]
    if math.log2(n_heads).is_integer():
        return geometric(n_heads)
    closest = 2 ** math.floor(math.log2(n_heads))   # the ragged case from the paper
    return geometric(closest) + geometric(2 * closest)[0::2][: n_heads - closest]

for h in [8, 12]:
    print(f"{h} heads:", np.array(alibi_slopes(h)))

def alibi_bias(seq_len, slope):
    """A number added to every attention score, before the softmax."""
    q = np.arange(seq_len)[:, None]
    k = np.arange(seq_len)[None, :]
    dist = k - q                                    # negative means to the left
    bias = slope * dist.astype(float)               # further left, more negative
    bias[dist > 0] = -np.inf                        # causal mask: no peeking right
    return bias

print("\nbias matrix for a head with slope 1/4, 5 tokens:")
print(alibi_bias(5, 0.25))

SCORES = np.random.default_rng(1).normal(scale=0.3, size=(16, 16))   # one fixed set of q.k

def last_row(slope):
    s = SCORES + alibi_bias(16, slope)
    w = np.exp(s - s.max(axis=1, keepdims=True))
    return (w / w.sum(axis=1, keepdims=True))[-1]

print("\nwhere the last token puts its attention, same scores, different slopes:")
for slope in [0.0, 0.5, 0.25, 1 / 64, 1 / 256]:
    w = last_row(slope)
    print(f"  slope {slope:.6f}  nearest 4 {w[-4:].round(3)}"
          f"  furthest 4 {w[:4].round(3)}  beyond 8 back {w[:8].sum():.4f}")

print("\nthe same code at lengths never trained on:")
for n in [16, 64, 4096]:
    print(f"  seq_len {n:>5}: bias matrix shape {alibi_bias(n, 0.25).shape},"
          f" no table lookup anywhere")
Output
8 heads: [0.5   0.25  0.125 0.062 0.031 0.016 0.008 0.004]
12 heads: [0.5   0.25  0.125 0.062 0.031 0.016 0.008 0.004 0.707 0.354 0.177 0.088]

bias matrix for a head with slope 1/4, 5 tokens:
[[ 0.    -inf  -inf  -inf  -inf]
 [-0.25  0.    -inf  -inf  -inf]
 [-0.5  -0.25  0.    -inf  -inf]
 [-0.75 -0.5  -0.25  0.    -inf]
 [-1.   -0.75 -0.5  -0.25  0.  ]]

where the last token puts its attention, same scores, different slopes:
  slope 0.000000  nearest 4 [0.068 0.05  0.077 0.055]  furthest 4 [0.065 0.06  0.067 0.039]  beyond 8 back 0.5060
  slope 0.500000  nearest 4 [0.097 0.118 0.3   0.35 ]  furthest 4 [0.    0.    0.001 0.001]  beyond 8 back 0.0186
  slope 0.250000  nearest 4 [0.116 0.111 0.218 0.199]  furthest 4 [0.006 0.007 0.009 0.007]  beyond 8 back 0.1219
  slope 0.015625  nearest 4 [0.072 0.055 0.085 0.061]  furthest 4 [0.057 0.054 0.061 0.036]  beyond 8 back 0.4748
  slope 0.003906  nearest 4 [0.069 0.051 0.079 0.056]  furthest 4 [0.063 0.058 0.065 0.038]  beyond 8 back 0.4982

the same code at lengths never trained on:
  seq_len    16: bias matrix shape (16, 16), no table lookup anywhere
  seq_len    64: bias matrix shape (64, 64), no table lookup anywhere
  seq_len  4096: bias matrix shape (4096, 4096), no table lookup anywhere

Reading that output

The 8-head slopes halve each time: 1/2, 1/4, 1/8 down to 1/256. That is a geometric sequence with ratio 2^(-8/n), which for n = 8 is exactly 1/2.

The 12-head list looks wrong, and it is not. After the eight powers of two it appends 0.707, 0.354, 0.177, 0.088 — larger values, out of order. When the head count is not a power of two, the reference implementation falls back to interleaving a finer geometric sequence, and does not re-sort. Every published ALiBi model with 12 heads has this ordering, so reproduce it rather than tidying it.

The bias matrix is the entire method. Zero on the diagonal, -slope one step left, -2 * slope two steps left. Nothing is learned, nothing is looked up.

The five attention rows share identical raw scores. Only the slope changes. At slope 0 the attention spreads out, with 0.506 of the weight landing more than eight tokens back. At slope 0.5 that collapses to 0.019, and the two nearest tokens take 0.65 between them. At slope 1/256 the row is almost indistinguishable from no penalty at all.

That spread is the argument for per-head slopes. In one layer you have a head that sees three tokens and a head that sees everything.

The final block is the extrapolation claim, stated modestly. No table, no lookup, no maximum. alibi_bias(4096, ...) builds without special handling because it is arithmetic on indices.

Making it cheap in real code

The naive version materialises an (n_heads, seq, seq) tensor, which is large and mostly redundant. Two standard optimisations:

alibi_efficient.py
import torch

def alibi_bias_rows(seq_len, slopes):
    """(heads, 1, seq): one row that broadcasts over all queries."""
    dist = torch.arange(seq_len) - (seq_len - 1)      # 0 at the last position
    return slopes[:, None, None] * dist[None, None, :]

slopes = torch.tensor([0.5, 0.25, 0.125, 0.0625])
print(alibi_bias_rows(6, slopes).shape)
print(alibi_bias_rows(6, slopes)[0, 0])
Output
torch.Size([4, 1, 6])
tensor([-2.5000, -2.0000, -1.5000, -1.0000, -0.5000,  0.0000])

Because the bias depends only on k - q, one row broadcasts across queries during single-token decoding. During prefill you still need the triangular form, but it can be built once per length and cached.

Common mistakes

Adding the bias after the softmax. It goes on the raw scores, before normalisation. Applied afterwards it stops being a soft penalty and becomes an arbitrary rescaling that does not sum to one.

Sorting the 12-head slopes. Tempting and wrong if you are loading published weights. The head order is baked into the trained parameters.

Combining ALiBi with a positional encoding. ALiBi replaces sinusoidal, learned and rotary encodings. Papers using it add no positional vectors at all. Stacking two schemes gives conflicting signals.

Assuming any attention kernel supports it. FlashAttention added ALiBi support, but not every fused kernel or serving stack does. Check before choosing ALiBi for a deployment, since the fallback is a much slower attention path.

Try it yourself

Set every slope to the same value and rerun the attention comparison. You now have a model with exactly one attention span. Then make the slopes span a much narrower range, say 0.2 to 0.3, and see how much of the head diversity disappears. That experiment is the fastest way to feel why the geometric spread was chosen.

What to learn next

Researcher — Mathematics and papers.

Definition

Press, Smith and Lewis (2022), Train Short, Test Long, modify the pre-softmax logit for head $h$:

$$ a^{(h)}_{ij} = \frac{q_i^\top k_j}{\sqrt{d_k}} - m_h \,\lvert i - j \rvert $$

for $j \le i$ under causal masking, with $m_h > 0$ a fixed per-head slope. No embedding is added anywhere and there are no learned positional parameters.

For $n$ heads with $n$ a power of two, the slopes are the geometric sequence

$$ m_h = 2^{-8h/n}, \qquad h = 1, \dots, n $$

so the ratio and the first term are both $2^{-8/n}$, and the sequence spans $2^{-8/n}$ down to $2^{-8}$ regardless of $n$. For non-powers of two the reference implementation concatenates the sequence for the nearest lower power of two with every second element of the sequence for twice that value — the ordering artefact visible in the developer output.

Why the slopes are not learned

The paper reports that making $m_h$ trainable did not improve results and hurt extrapolation. The fixed geometric spread is a strong prior: it hard-codes a range of attention spans across heads without spending gradient on discovering them.

The extrapolation result, and its limits

The headline is that a model trained at length $L$ evaluates at $\gg L$ with perplexity that does not blow up, whereas sinusoidal and learned embeddings degrade sharply. The paper trains at 1024 and evaluates to 2048 and beyond on WikiText-103.

Two honest qualifications.

The mechanism is boundedness, not understanding. As $|i-j|$ grows, $-m_h|i-j| \to -\infty$, so distant tokens receive vanishing weight. Every head has an effective window of order $1/m_h$. Beyond that window, extending the input adds almost nothing, and perplexity is stable because the model is effectively ignoring the extra context.

Long-range retrieval is where it costs. Kazemnejad et al. (2023) find ALiBi below T5's relative bias and NoPE on downstream length generalisation. Work on long-context retrieval finds the exponential decay actively harmful for tasks requiring a fact from thousands of tokens back. Perplexity stability and long-range capability are different properties, and ALiBi buys the first.

Relation to a sliding window

With slope $m_h$, the bias reaches $-\ln(10^k)$ at distance $|i-j| = k\ln(10)/m_h$. Beyond a few multiples of $1/m_h$ the contribution is numerically irrelevant. ALiBi is therefore close to a soft, per-head sliding window, with the softness giving a smooth gradient instead of a hard cutoff.

This connection is why ALiBi composes cleanly with windowed attention kernels, and why comparisons with sliding-window models often show similar behaviour.

Cost

Zero parameters. The bias term is $O(n^2)$ additions per head in the naive form, but the dependence on $i-j$ alone means it can be applied as a broadcast row, or folded into a fused attention kernel with negligible overhead. Memory is $O(n)$ per head rather than $O(n^2)$ when broadcast.

Adoption and current standing

BLOOM (BigScience, 2022), MPT (MosaicML, 2023) and Baichuan 2's 13B model all use ALiBi. Several speech models adopted it for variable-length inputs.

New large decoder-only language models have largely converged on RoPE with a scaling scheme instead. The reasons given in practice are long-range retrieval quality and the ecosystem effect: kernels, serving stacks and context-extension recipes are all built around RoPE. ALiBi remains a strong choice when inputs vary widely in length and the task is genuinely local.

  • Kerple (Chi et al., 2022) generalises the linear penalty to a family of conditionally positive definite kernels, including logarithmic variants, and reports better extrapolation.
  • Sandwich (Chi et al., 2023) derives a bias from the inner product of sinusoidal encodings, connecting the additive and bias families.
  • FIRE (Li et al., 2024) learns the bias as a function of a normalised progressive distance, and reports strong length generalisation.

Papers

What to learn next