ALiBi
ALiBi adds no vectors at all. It subtracts a penalty from the attention score that grows with distance, giving each head its own sense of how far back to look.
- 13 min read
- 3 reading levels
- Updated
Read these first
On this page 7
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
ALiBi does not encode position at all. It penalises attention that reaches far back, and the penalty grows with distance.
Think of standing in a noisy railway platform crowd. You hear the person beside you without effort. Someone ten metres away is faint. Someone at the far end is inaudible.
Nobody handed you a distance chart. Sound fades with distance on its own, and you use that fade to work out who is near.
ALiBi does the same to attention. Nothing gets stamped, added or rotated. Distant tokens are turned down, and the model reads distance from how loud each one is.
Why anyone would do it this way
Every other method builds a representation of position and hopes attention learns to use it. ALiBi skips the representation and edits attention directly.
The payoff is length. There is no table to run out of and no dial that spins into unseen angles. Distance 5000 gets a penalty of five thousand times the rate, computed by one multiplication.
The name is short for attention with linear biases — a straight-line penalty applied to the attention scores.
How it works
Before attention decides who to listen to, it produces a score for every pair of tokens. ALiBi subtracts an amount proportional to the distance from each score.
token 10 looking back:
at token 9 (1 back) score - 1 x rate
at token 8 (2 back) score - 2 x rate
at token 5 (5 back) score - 5 x rate
at token 0 (10 back) score - 10 x rate
then softmax turns those into attention weightsOne number per head, called the slope, sets the rate. That is the whole method.
The clever part: every head gets a different rate
If all heads had the same penalty, the model would have one fixed attention span. Instead the slopes are spread out geometrically.
The first head gets a steep penalty and can barely see past a handful of tokens. The last head gets a penalty so gentle it is almost ignoring distance entirely.
So one head handles "the word right before this one" and another handles "somewhere much earlier in the document". The model gets a range of spans for free, without learning any of them.
Where this shows up
- BLOOM, the 176-billion-parameter multilingual model, uses ALiBi throughout.
- MPT models from MosaicML, marketed on their ability to accept very long inputs.
- Baichuan 2's larger model.
- Several speech and audio transformers, where inputs vary hugely in length.
What is honestly hard here
ALiBi's headline claim is length extrapolation, and the claim is narrower than it sounds. Perplexity stays low on much longer inputs. That is a measure of how well the model predicts the next token, and a recency-biased model does well on it.
Tasks needing a fact from far back are different. The penalty that makes ALiBi extrapolate also makes distant tokens quiet. Later work found ALiBi weak at long-range retrieval for exactly that reason. The design has a built-in preference and it costs something.
Remember this
- ALiBi subtracts a penalty from attention scores, proportional to distance.
- Each head gets its own penalty rate, so the model gets many attention spans for free.
- It handles longer inputs without changes, but it is biased toward nearby tokens by construction.
What to learn next
- Relative position embeddings — the learned version of the same bias idea.
- What breaks when you exceed the trained length — the claim measured directly.
- Attention — the scores ALiBi is editing.
Developer — Code and libraries.
Setup
pip install numpyThe whole method
import math
import numpy as np
np.set_printoptions(precision=3, suppress=True)
def alibi_slopes(n_heads):
"""One fixed penalty rate per head. Never trained."""
def geometric(n):
start = 2 ** (-8 / n) # ratio and first term are both 2^(-8/n)
return [start ** (i + 1) for i in range(n)]
if math.log2(n_heads).is_integer():
return geometric(n_heads)
closest = 2 ** math.floor(math.log2(n_heads)) # the ragged case from the paper
return geometric(closest) + geometric(2 * closest)[0::2][: n_heads - closest]
for h in [8, 12]:
print(f"{h} heads:", np.array(alibi_slopes(h)))
def alibi_bias(seq_len, slope):
"""A number added to every attention score, before the softmax."""
q = np.arange(seq_len)[:, None]
k = np.arange(seq_len)[None, :]
dist = k - q # negative means to the left
bias = slope * dist.astype(float) # further left, more negative
bias[dist > 0] = -np.inf # causal mask: no peeking right
return bias
print("\nbias matrix for a head with slope 1/4, 5 tokens:")
print(alibi_bias(5, 0.25))
SCORES = np.random.default_rng(1).normal(scale=0.3, size=(16, 16)) # one fixed set of q.k
def last_row(slope):
s = SCORES + alibi_bias(16, slope)
w = np.exp(s - s.max(axis=1, keepdims=True))
return (w / w.sum(axis=1, keepdims=True))[-1]
print("\nwhere the last token puts its attention, same scores, different slopes:")
for slope in [0.0, 0.5, 0.25, 1 / 64, 1 / 256]:
w = last_row(slope)
print(f" slope {slope:.6f} nearest 4 {w[-4:].round(3)}"
f" furthest 4 {w[:4].round(3)} beyond 8 back {w[:8].sum():.4f}")
print("\nthe same code at lengths never trained on:")
for n in [16, 64, 4096]:
print(f" seq_len {n:>5}: bias matrix shape {alibi_bias(n, 0.25).shape},"
f" no table lookup anywhere")8 heads: [0.5 0.25 0.125 0.062 0.031 0.016 0.008 0.004] 12 heads: [0.5 0.25 0.125 0.062 0.031 0.016 0.008 0.004 0.707 0.354 0.177 0.088] bias matrix for a head with slope 1/4, 5 tokens: [[ 0. -inf -inf -inf -inf] [-0.25 0. -inf -inf -inf] [-0.5 -0.25 0. -inf -inf] [-0.75 -0.5 -0.25 0. -inf] [-1. -0.75 -0.5 -0.25 0. ]] where the last token puts its attention, same scores, different slopes: slope 0.000000 nearest 4 [0.068 0.05 0.077 0.055] furthest 4 [0.065 0.06 0.067 0.039] beyond 8 back 0.5060 slope 0.500000 nearest 4 [0.097 0.118 0.3 0.35 ] furthest 4 [0. 0. 0.001 0.001] beyond 8 back 0.0186 slope 0.250000 nearest 4 [0.116 0.111 0.218 0.199] furthest 4 [0.006 0.007 0.009 0.007] beyond 8 back 0.1219 slope 0.015625 nearest 4 [0.072 0.055 0.085 0.061] furthest 4 [0.057 0.054 0.061 0.036] beyond 8 back 0.4748 slope 0.003906 nearest 4 [0.069 0.051 0.079 0.056] furthest 4 [0.063 0.058 0.065 0.038] beyond 8 back 0.4982 the same code at lengths never trained on: seq_len 16: bias matrix shape (16, 16), no table lookup anywhere seq_len 64: bias matrix shape (64, 64), no table lookup anywhere seq_len 4096: bias matrix shape (4096, 4096), no table lookup anywhere
Reading that output
The 8-head slopes halve each time: 1/2, 1/4, 1/8 down to 1/256. That is a geometric sequence with ratio 2^(-8/n), which for n = 8 is exactly 1/2.
The 12-head list looks wrong, and it is not. After the eight powers of two it appends 0.707, 0.354, 0.177, 0.088 — larger values, out of order. When the head count is not a power of two, the reference implementation falls back to interleaving a finer geometric sequence, and does not re-sort. Every published ALiBi model with 12 heads has this ordering, so reproduce it rather than tidying it.
The bias matrix is the entire method. Zero on the diagonal, -slope one step left, -2 * slope two steps left. Nothing is learned, nothing is looked up.
The five attention rows share identical raw scores. Only the slope changes. At slope 0 the attention spreads out, with 0.506 of the weight landing more than eight tokens back. At slope 0.5 that collapses to 0.019, and the two nearest tokens take 0.65 between them. At slope 1/256 the row is almost indistinguishable from no penalty at all.
That spread is the argument for per-head slopes. In one layer you have a head that sees three tokens and a head that sees everything.
The final block is the extrapolation claim, stated modestly. No table, no lookup, no maximum. alibi_bias(4096, ...) builds without special handling because it is arithmetic on indices.
Making it cheap in real code
The naive version materialises an (n_heads, seq, seq) tensor, which is large and mostly redundant. Two standard optimisations:
import torch
def alibi_bias_rows(seq_len, slopes):
"""(heads, 1, seq): one row that broadcasts over all queries."""
dist = torch.arange(seq_len) - (seq_len - 1) # 0 at the last position
return slopes[:, None, None] * dist[None, None, :]
slopes = torch.tensor([0.5, 0.25, 0.125, 0.0625])
print(alibi_bias_rows(6, slopes).shape)
print(alibi_bias_rows(6, slopes)[0, 0])torch.Size([4, 1, 6]) tensor([-2.5000, -2.0000, -1.5000, -1.0000, -0.5000, 0.0000])
Because the bias depends only on k - q, one row broadcasts across queries during single-token decoding. During prefill you still need the triangular form, but it can be built once per length and cached.
Common mistakes
Adding the bias after the softmax. It goes on the raw scores, before normalisation. Applied afterwards it stops being a soft penalty and becomes an arbitrary rescaling that does not sum to one.
Sorting the 12-head slopes. Tempting and wrong if you are loading published weights. The head order is baked into the trained parameters.
Combining ALiBi with a positional encoding. ALiBi replaces sinusoidal, learned and rotary encodings. Papers using it add no positional vectors at all. Stacking two schemes gives conflicting signals.
Assuming any attention kernel supports it. FlashAttention added ALiBi support, but not every fused kernel or serving stack does. Check before choosing ALiBi for a deployment, since the fallback is a much slower attention path.
Try it yourself
Set every slope to the same value and rerun the attention comparison. You now have a model with exactly one attention span. Then make the slopes span a much narrower range, say 0.2 to 0.3, and see how much of the head diversity disappears. That experiment is the fastest way to feel why the geometric spread was chosen.
What to learn next
- Relative position embeddings — the learned version of the same bias idea.
- What breaks when you exceed the trained length — the claim measured directly.
- Attention — the scores ALiBi is editing.
Researcher — Mathematics and papers.
Definition
Press, Smith and Lewis (2022), Train Short, Test Long, modify the pre-softmax logit for head $h$:
$$ a^{(h)}_{ij} = \frac{q_i^\top k_j}{\sqrt{d_k}} - m_h \,\lvert i - j \rvert $$
for $j \le i$ under causal masking, with $m_h > 0$ a fixed per-head slope. No embedding is added anywhere and there are no learned positional parameters.
For $n$ heads with $n$ a power of two, the slopes are the geometric sequence
$$ m_h = 2^{-8h/n}, \qquad h = 1, \dots, n $$
so the ratio and the first term are both $2^{-8/n}$, and the sequence spans $2^{-8/n}$ down to $2^{-8}$ regardless of $n$. For non-powers of two the reference implementation concatenates the sequence for the nearest lower power of two with every second element of the sequence for twice that value — the ordering artefact visible in the developer output.
Why the slopes are not learned
The paper reports that making $m_h$ trainable did not improve results and hurt extrapolation. The fixed geometric spread is a strong prior: it hard-codes a range of attention spans across heads without spending gradient on discovering them.
The extrapolation result, and its limits
The headline is that a model trained at length $L$ evaluates at $\gg L$ with perplexity that does not blow up, whereas sinusoidal and learned embeddings degrade sharply. The paper trains at 1024 and evaluates to 2048 and beyond on WikiText-103.
Two honest qualifications.
The mechanism is boundedness, not understanding. As $|i-j|$ grows, $-m_h|i-j| \to -\infty$, so distant tokens receive vanishing weight. Every head has an effective window of order $1/m_h$. Beyond that window, extending the input adds almost nothing, and perplexity is stable because the model is effectively ignoring the extra context.
Long-range retrieval is where it costs. Kazemnejad et al. (2023) find ALiBi below T5's relative bias and NoPE on downstream length generalisation. Work on long-context retrieval finds the exponential decay actively harmful for tasks requiring a fact from thousands of tokens back. Perplexity stability and long-range capability are different properties, and ALiBi buys the first.
Relation to a sliding window
With slope $m_h$, the bias reaches $-\ln(10^k)$ at distance $|i-j| = k\ln(10)/m_h$. Beyond a few multiples of $1/m_h$ the contribution is numerically irrelevant. ALiBi is therefore close to a soft, per-head sliding window, with the softness giving a smooth gradient instead of a hard cutoff.
This connection is why ALiBi composes cleanly with windowed attention kernels, and why comparisons with sliding-window models often show similar behaviour.
Cost
Zero parameters. The bias term is $O(n^2)$ additions per head in the naive form, but the dependence on $i-j$ alone means it can be applied as a broadcast row, or folded into a fused attention kernel with negligible overhead. Memory is $O(n)$ per head rather than $O(n^2)$ when broadcast.
Adoption and current standing
BLOOM (BigScience, 2022), MPT (MosaicML, 2023) and Baichuan 2's 13B model all use ALiBi. Several speech models adopted it for variable-length inputs.
New large decoder-only language models have largely converged on RoPE with a scaling scheme instead. The reasons given in practice are long-range retrieval quality and the ecosystem effect: kernels, serving stacks and context-extension recipes are all built around RoPE. ALiBi remains a strong choice when inputs vary widely in length and the task is genuinely local.
Related bias designs
- Kerple (Chi et al., 2022) generalises the linear penalty to a family of conditionally positive definite kernels, including logarithmic variants, and reports better extrapolation.
- Sandwich (Chi et al., 2023) derives a bias from the inner product of sinusoidal encodings, connecting the additive and bias families.
- FIRE (Li et al., 2024) learns the bias as a function of a normalised progressive distance, and reports strong length generalisation.
Papers
- Press, Smith and Lewis, Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation, ICLR 2022 — arxiv.org/abs/2108.12409
- Chi et al., Kerple: Kernelized Relative Positional Embedding for Length Extrapolation, NeurIPS 2022 — arxiv.org/abs/2205.09921
- Li et al., Functional Interpolation for Relative Positions Improves Long Context Transformers (FIRE), ICLR 2024 — arxiv.org/abs/2310.04418
- Kazemnejad et al., The Impact of Positional Encoding on Length Generalization in Transformers, NeurIPS 2023 — arxiv.org/abs/2305.19466
What to learn next
- Relative position embeddings — the learned version of the same bias idea.
- What breaks when you exceed the trained length — the claim measured directly.
- Attention — the scores ALiBi is editing.