Walking through one transformer block
A transformer is one small block copied dozens of times, and that block does exactly two things - look around at other tokens, then think about each token alone.
- 11 min read
- 3 reading levels
- Updated
Read these first
On this page 7
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
A transformer block does two things. Every token looks around at the other tokens, then thinks by itself. That block is copied dozens of times.
Think about getting a form processed at a government office. You go to counter one, then counter two, then counter three. At every counter the clerk does the same pair of things. They glance at the other papers in your file. Then they read your form on its own and add a stamp.
Nothing about the form changes shape. It arrives at counter one as an A4 sheet. It leaves counter forty as an A4 sheet, carrying more stamps.
That is a transformer. One counter, repeated. And the reason it can be repeated is that the paper comes out the same size it went in.
The two halves
Half one: look around. This is the attention step. Every token collects information from the other tokens in the sequence. A pronoun finds its noun. A verb finds its subject.
Half two: think alone. This is the feedforward step. Every token is processed on its own, with no reference to any other token. This is where a lot of the model's stored knowledge is applied.
The alternation matters. Gather, then digest. Gather, then digest. Forty times over.
The two rules that make stacking possible
Rule one: nothing is replaced, only added. Each half computes a small adjustment. It adds that to what was already there. The token's existing information is never thrown away. This addition is called a residual connection — a path that carries the original straight through, untouched.
Rule two: tidy the numbers before each half. The numbers are rescaled to a consistent size first. This step is called normalisation, meaning bringing values back to a standard range. Without it, forty layers of additions produce numbers far too big or far too small to work with.
The shape of one block
token comes in
│
├──────────────────────────┐ keep a copy (this is the residual path)
│ │
tidy the numbers │
│ │
LOOK AROUND (attention) │
│ │
└────────► add ◄───────────┘
│
┌───────────┴──────────────┐ keep a copy again
│ │
tidy the numbers │
│ │
THINK ALONE (feedforward) │
│ │
└────────► add ◄───────────┘
│
token goes out
(same size as it came in)Two halves. Two copies kept. Two additions. That is the whole block.
Why the same block over and over
You might expect a model to have different kinds of layers doing different jobs. A factory has different machines, after all. Transformers do not work that way.
Every block has the same shape. What differs is the numbers inside, which are learned separately for each block. Early blocks tend to handle surface things like spelling and grammar. Later blocks handle meaning. Nobody designed that split — it emerges from training.
This uniformity is a large part of why transformers took over. One block, easy to write and easy to optimise. To make a model bigger, you make the block wider or you add more of them.
What a block is not
A block does not decide anything. It has no memory of previous sentences. It does not loop or branch.
It is a fixed sequence of multiplications and additions, identical for every input. All of a model's apparent cleverness lives in the numbers, not in the structure.
Remember this
- A block is attention (look around) followed by feedforward (think alone).
- Each half adds its result to what was already there rather than replacing it.
- Input and output are the same size, which is why blocks stack without any glue.
What to learn next
- The residual stream — the path those two additions write into.
- Multi-head attention — the first half of the block, in detail.
- Transformers — the wider picture this block sits inside.
Developer — Code and libraries.
Setup
pip install torchThe whole block, with every intermediate shape printed
import torch
import torch.nn as nn
import torch.nn.functional as F
torch.manual_seed(0)
class Block(nn.Module):
"""One pre-norm transformer block, written so every step is visible."""
def __init__(self, d_model=32, n_heads=4, mult=4):
super().__init__()
self.n_heads = n_heads
self.norm1 = nn.LayerNorm(d_model)
self.qkv = nn.Linear(d_model, 3 * d_model, bias=False)
self.proj = nn.Linear(d_model, d_model, bias=False)
self.norm2 = nn.LayerNorm(d_model)
self.fc1 = nn.Linear(d_model, mult * d_model)
self.fc2 = nn.Linear(mult * d_model, d_model)
def forward(self, x, trace=False):
B, T, C = x.shape
H, hd = self.n_heads, C // self.n_heads
log = (lambda n, t: print(f" {n:26s} {tuple(t.shape)}")) if trace else (lambda n, t: None)
log("input x", x)
h = self.norm1(x); log("norm1(x)", h)
q, k, v = self.qkv(h).chunk(3, dim=-1); log("q (and k, v)", q)
q, k, v = (t.view(B, T, H, hd).transpose(1, 2) for t in (q, k, v))
log("q split into heads", q)
a = F.scaled_dot_product_attention(q, k, v, is_causal=True)
log("attention output", a)
a = a.transpose(1, 2).reshape(B, T, C); log("heads merged", a)
a = self.proj(a); log("out projection", a)
x = x + a; log("x + attention <- add", x)
h = self.norm2(x); log("norm2(x)", h)
h = self.fc2(F.gelu(self.fc1(h))); log("feedforward", h)
x = x + h; log("x + feedforward <- add", x)
return x
block = Block()
x = torch.randn(2, 5, 32) # 2 sequences, 5 tokens each, width 32
print("shape at every step inside one block:")
y = block(x, trace=True)
print("\nthe block preserves shape exactly:", tuple(x.shape), "->", tuple(y.shape))
print("so blocks stack without any glue between them.\n")
print("where the parameters sit:")
for name, p in block.named_parameters():
print(f" {name:12s} {str(tuple(p.shape)):12s} {p.numel():>6,}")
print(f" {'TOTAL':12s} {'':12s} {sum(p.numel() for p in block.parameters()):>6,}")
deep = nn.Sequential(*[Block() for _ in range(6)])
print("\nsix blocks stacked:", tuple(deep(x).shape),
f"| {sum(p.numel() for p in deep.parameters()):,} parameters")shape at every step inside one block: input x (2, 5, 32) norm1(x) (2, 5, 32) q (and k, v) (2, 5, 32) q split into heads (2, 4, 5, 8) attention output (2, 4, 5, 8) heads merged (2, 5, 32) out projection (2, 5, 32) x + attention <- add (2, 5, 32) norm2(x) (2, 5, 32) feedforward (2, 5, 32) x + feedforward <- add (2, 5, 32) the block preserves shape exactly: (2, 5, 32) -> (2, 5, 32) so blocks stack without any glue between them. where the parameters sit: norm1.weight (32,) 32 norm1.bias (32,) 32 qkv.weight (96, 32) 3,072 proj.weight (32, 32) 1,024 norm2.weight (32,) 32 norm2.bias (32,) 32 fc1.weight (128, 32) 4,096 fc1.bias (128,) 128 fc2.weight (32, 128) 4,096 fc2.bias (32,) 32 TOTAL 12,576 six blocks stacked: (2, 5, 32) | 75,456 parameters
Reading the trace
Only one line changes shape, and it changes back. (2, 5, 32) becomes (2, 4, 5, 8) for the attention step and returns immediately. Four heads times eight slots is thirty-two. Nothing is created or lost; the width is regrouped.
Two lines are marked <- add, and they are the only lines that touch x. Everything between them computes a proposal. Those two additions are what the block actually does to the token. Delete either one and a deep stack becomes untrainable.
The feedforward layer holds two thirds of the parameters. fc1 plus fc2 is 8,192 of the block's 12,576. Attention is 4,096. Most people assume the attention is where the weights live. It is not. See the feedforward layer.
The norms are 128 parameters out of 12,576. One percent, and removing them makes deep training fall apart. Parameter count is a poor guide to importance.
The two things that vary between real models
Where the norm sits. This version normalises before each half, called pre-norm. The 2017 original normalised after each half. That choice decides whether a deep model trains without a warm-up period — see pre-norm vs post-norm.
Which norm. LayerNorm here; most models since Llama use RMSNorm, which drops the mean-subtraction step and half the parameters.
Everything else — the two halves, the two additions, the shape preservation — is common to essentially every transformer in production.
Reading a real config file
A Hugging Face config.json maps onto this block directly:
| Field | What it sets |
|---|---|
hidden_size | the width C carried between blocks |
num_hidden_layers | how many copies of this block |
num_attention_heads | how the width is split for attention |
intermediate_size | the inner width of the feedforward half |
num_key_value_heads | key and value heads, if fewer than query heads |
rms_norm_eps | the small constant inside the normalisation |
Llama 3 8B: hidden_size 4096, num_hidden_layers 32, num_attention_heads 32, num_key_value_heads 8, intermediate_size 14336. Once you can read those five numbers, you can read any transformer config.
Common mistakes
Writing x = self.norm1(x) instead of h = self.norm1(x). Normalising x in place destroys the residual path, so the block replaces rather than adds. It trains, badly, and the bug is invisible in the shapes.
Forgetting is_causal=True in a decoder. Loss drops implausibly fast because every position can read its own answer. If your training loss looks too good, check this first.
Using one nn.LayerNorm object for both positions. They must be separate modules with separate weights. Sharing them silently ties two unrelated parameters together.
Assuming Block can be dropped into nn.Sequential with extra arguments. nn.Sequential passes one tensor. The trace flag above defaults to False for exactly that reason.
Try it yourself
Comment out x = x + a and replace it with x = a. Then stack twelve blocks and compare output magnitudes with the original. Then put it back and change mult from 4 to 1. Watch the parameter share flip from two thirds to one third.
What to learn next
- The residual stream — the path those two additions write into.
- Multi-head attention — the first half of the block, in detail.
- Transformers — the wider picture this block sits inside.
Researcher — Mathematics and papers.
The block as a function
For input $X \in \mathbb{R}^{T \times d}$, the pre-norm block is:
$$ X' = X + \operatorname{MHA}(\operatorname{Norm}(X)) $$ $$ X'' = X' + \operatorname{FFN}(\operatorname{Norm}(X')) $$
Both sublayers map $\mathbb{R}^{T \times d} \to \mathbb{R}^{T \times d}$, so the block is an endomorphism and composition is unconstrained. That is the structural reason depth is a free parameter in this architecture.
Attention mixes across the $T$ axis and acts identically across the $d$ axis within a head. The feedforward layer mixes across the $d$ axis and acts identically across the $T$ axis. The block therefore alternates mixing along the two axes of the input. That is the structural motif of the MLP-Mixer (Tolstikhin et al., 2021, arXiv:2105.01601). It replaces attention with a second MLP and still works reasonably, which suggests the alternation itself carries some value.
Parameter accounting
Per block, ignoring norms and biases, with $d_{\text{ff}} = 4d$:
$$ P_{\text{attn}} = 4d^2, \qquad P_{\text{ffn}} = 8d^2, \qquad P_{\text{block}} = 12d^2 $$
The $12d^2$ per layer is the constant behind the standard $N \approx 12 L d^2$ estimate for non-embedding parameters. It is therefore also behind the $6ND$ training-compute rule. Gated feedforward layers change $P_{\text{ffn}}$ to $3 d\, d_{\text{ff}}$; grouped-query attention changes $P_{\text{attn}}$ to $2d^2 + 2 d\, g\, d_h$. See counting a model's parameters by hand.
Where the design has actually moved since 2017
The block's skeleton is unchanged. Five substitutions account for most of the difference between a 2017 encoder and a 2025 decoder:
| Component | 2017 | Common now |
|---|---|---|
| Norm placement | post-norm | pre-norm |
| Norm | LayerNorm | RMSNorm |
| Feedforward | ReLU, two matrices | SwiGLU, three matrices |
| Position | added sinusoids | rotary, applied inside attention |
| Key/value heads | one per query head | grouped |
Narang et al. (2021), arXiv:2102.11972, is the necessary counterweight to any list like this. Sweeping dozens of published modifications under one codebase, they found most did not reproduce their reported gains. Gated feedforward layers and RMSNorm were among the few that did.
Interpretability consequences of the additive form
Both sublayers write additively into a shared stream. The model output therefore decomposes exactly into a sum of per-component contributions:
$$ X^{(L)} = X^{(0)} + \sum_{\ell=1}^{L} \left[ \operatorname{MHA}^{(\ell)}(\cdot) + \operatorname{FFN}^{(\ell)}(\cdot) \right] $$
Elhage et al. (2021) build the residual-stream framework on this identity. The logit-lens family (nostalgebraist, 2020) exploits it by applying the output head to intermediate stream states. The decomposition is exact. What is not exact is treating the terms as independent, since each term's input depends on all preceding terms. See the residual stream.
Papers
- Vaswani et al., Attention Is All You Need, 2017 — arxiv.org/abs/1706.03762
- Xiong et al., On Layer Normalization in the Transformer Architecture, 2020 — arxiv.org/abs/2002.04745
- Narang et al., Do Transformer Modifications Transfer?, 2021 — arxiv.org/abs/2102.11972
- Tolstikhin et al., MLP-Mixer, 2021 — arxiv.org/abs/2105.01601
- Elhage et al., A Mathematical Framework for Transformer Circuits, 2021 — transformer-circuits.pub/2021/framework
What to learn next
- The residual stream — the path those two additions write into.
- Multi-head attention — the first half of the block, in detail.
- Transformers — the wider picture this block sits inside.