Inside a Transformer Block

Walking through one transformer block

A transformer is one small block copied dozens of times, and that block does exactly two things - look around at other tokens, then think about each token alone.

On this page 7
  1. The two halves
  2. The two rules that make stacking possible
  3. The shape of one block
  4. Why the same block over and over
  5. What a block is not
  6. Remember this
  7. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

A transformer block does two things. Every token looks around at the other tokens, then thinks by itself. That block is copied dozens of times.

Think about getting a form processed at a government office. You go to counter one, then counter two, then counter three. At every counter the clerk does the same pair of things. They glance at the other papers in your file. Then they read your form on its own and add a stamp.

Nothing about the form changes shape. It arrives at counter one as an A4 sheet. It leaves counter forty as an A4 sheet, carrying more stamps.

That is a transformer. One counter, repeated. And the reason it can be repeated is that the paper comes out the same size it went in.

The two halves

Half one: look around. This is the attention step. Every token collects information from the other tokens in the sequence. A pronoun finds its noun. A verb finds its subject.

Half two: think alone. This is the feedforward step. Every token is processed on its own, with no reference to any other token. This is where a lot of the model's stored knowledge is applied.

The alternation matters. Gather, then digest. Gather, then digest. Forty times over.

The two rules that make stacking possible

Rule one: nothing is replaced, only added. Each half computes a small adjustment. It adds that to what was already there. The token's existing information is never thrown away. This addition is called a residual connection — a path that carries the original straight through, untouched.

Rule two: tidy the numbers before each half. The numbers are rescaled to a consistent size first. This step is called normalisation, meaning bringing values back to a standard range. Without it, forty layers of additions produce numbers far too big or far too small to work with.

The shape of one block

   token comes in
        │
        ├──────────────────────────┐   keep a copy (this is the residual path)
        │                          │
   tidy the numbers                │
        │                          │
   LOOK AROUND (attention)         │
        │                          │
        └────────► add ◄───────────┘
                    │
        ┌───────────┴──────────────┐   keep a copy again
        │                          │
   tidy the numbers                │
        │                          │
   THINK ALONE (feedforward)       │
        │                          │
        └────────► add ◄───────────┘
                    │
              token goes out
             (same size as it came in)

Two halves. Two copies kept. Two additions. That is the whole block.

Why the same block over and over

You might expect a model to have different kinds of layers doing different jobs. A factory has different machines, after all. Transformers do not work that way.

Every block has the same shape. What differs is the numbers inside, which are learned separately for each block. Early blocks tend to handle surface things like spelling and grammar. Later blocks handle meaning. Nobody designed that split — it emerges from training.

This uniformity is a large part of why transformers took over. One block, easy to write and easy to optimise. To make a model bigger, you make the block wider or you add more of them.

What a block is not

A block does not decide anything. It has no memory of previous sentences. It does not loop or branch.

It is a fixed sequence of multiplications and additions, identical for every input. All of a model's apparent cleverness lives in the numbers, not in the structure.

Remember this

  • A block is attention (look around) followed by feedforward (think alone).
  • Each half adds its result to what was already there rather than replacing it.
  • Input and output are the same size, which is why blocks stack without any glue.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install torch

The whole block, with every intermediate shape printed

block.py
import torch
import torch.nn as nn
import torch.nn.functional as F

torch.manual_seed(0)

class Block(nn.Module):
    """One pre-norm transformer block, written so every step is visible."""
    def __init__(self, d_model=32, n_heads=4, mult=4):
        super().__init__()
        self.n_heads = n_heads
        self.norm1 = nn.LayerNorm(d_model)
        self.qkv   = nn.Linear(d_model, 3 * d_model, bias=False)
        self.proj  = nn.Linear(d_model, d_model, bias=False)
        self.norm2 = nn.LayerNorm(d_model)
        self.fc1   = nn.Linear(d_model, mult * d_model)
        self.fc2   = nn.Linear(mult * d_model, d_model)

    def forward(self, x, trace=False):
        B, T, C = x.shape
        H, hd = self.n_heads, C // self.n_heads
        log = (lambda n, t: print(f"  {n:26s} {tuple(t.shape)}")) if trace else (lambda n, t: None)

        log("input x", x)
        h = self.norm1(x);                                    log("norm1(x)", h)
        q, k, v = self.qkv(h).chunk(3, dim=-1);               log("q (and k, v)", q)
        q, k, v = (t.view(B, T, H, hd).transpose(1, 2) for t in (q, k, v))
        log("q split into heads", q)
        a = F.scaled_dot_product_attention(q, k, v, is_causal=True)
        log("attention output", a)
        a = a.transpose(1, 2).reshape(B, T, C);               log("heads merged", a)
        a = self.proj(a);                                     log("out projection", a)
        x = x + a;                                            log("x + attention  <- add", x)
        h = self.norm2(x);                                    log("norm2(x)", h)
        h = self.fc2(F.gelu(self.fc1(h)));                    log("feedforward", h)
        x = x + h;                                            log("x + feedforward <- add", x)
        return x

block = Block()
x = torch.randn(2, 5, 32)                  # 2 sequences, 5 tokens each, width 32
print("shape at every step inside one block:")
y = block(x, trace=True)

print("\nthe block preserves shape exactly:", tuple(x.shape), "->", tuple(y.shape))
print("so blocks stack without any glue between them.\n")

print("where the parameters sit:")
for name, p in block.named_parameters():
    print(f"  {name:12s} {str(tuple(p.shape)):12s} {p.numel():>6,}")
print(f"  {'TOTAL':12s} {'':12s} {sum(p.numel() for p in block.parameters()):>6,}")

deep = nn.Sequential(*[Block() for _ in range(6)])
print("\nsix blocks stacked:", tuple(deep(x).shape),
      f"| {sum(p.numel() for p in deep.parameters()):,} parameters")
Output
shape at every step inside one block:
  input x                    (2, 5, 32)
  norm1(x)                   (2, 5, 32)
  q (and k, v)               (2, 5, 32)
  q split into heads         (2, 4, 5, 8)
  attention output           (2, 4, 5, 8)
  heads merged               (2, 5, 32)
  out projection             (2, 5, 32)
  x + attention  <- add      (2, 5, 32)
  norm2(x)                   (2, 5, 32)
  feedforward                (2, 5, 32)
  x + feedforward <- add     (2, 5, 32)

the block preserves shape exactly: (2, 5, 32) -> (2, 5, 32)
so blocks stack without any glue between them.

where the parameters sit:
  norm1.weight (32,)            32
  norm1.bias   (32,)            32
  qkv.weight   (96, 32)      3,072
  proj.weight  (32, 32)      1,024
  norm2.weight (32,)            32
  norm2.bias   (32,)            32
  fc1.weight   (128, 32)     4,096
  fc1.bias     (128,)          128
  fc2.weight   (32, 128)     4,096
  fc2.bias     (32,)            32
  TOTAL                     12,576

six blocks stacked: (2, 5, 32) | 75,456 parameters

Reading the trace

Only one line changes shape, and it changes back. (2, 5, 32) becomes (2, 4, 5, 8) for the attention step and returns immediately. Four heads times eight slots is thirty-two. Nothing is created or lost; the width is regrouped.

Two lines are marked <- add, and they are the only lines that touch x. Everything between them computes a proposal. Those two additions are what the block actually does to the token. Delete either one and a deep stack becomes untrainable.

The feedforward layer holds two thirds of the parameters. fc1 plus fc2 is 8,192 of the block's 12,576. Attention is 4,096. Most people assume the attention is where the weights live. It is not. See the feedforward layer.

The norms are 128 parameters out of 12,576. One percent, and removing them makes deep training fall apart. Parameter count is a poor guide to importance.

The two things that vary between real models

Where the norm sits. This version normalises before each half, called pre-norm. The 2017 original normalised after each half. That choice decides whether a deep model trains without a warm-up period — see pre-norm vs post-norm.

Which norm. LayerNorm here; most models since Llama use RMSNorm, which drops the mean-subtraction step and half the parameters.

Everything else — the two halves, the two additions, the shape preservation — is common to essentially every transformer in production.

Reading a real config file

A Hugging Face config.json maps onto this block directly:

FieldWhat it sets
hidden_sizethe width C carried between blocks
num_hidden_layershow many copies of this block
num_attention_headshow the width is split for attention
intermediate_sizethe inner width of the feedforward half
num_key_value_headskey and value heads, if fewer than query heads
rms_norm_epsthe small constant inside the normalisation

Llama 3 8B: hidden_size 4096, num_hidden_layers 32, num_attention_heads 32, num_key_value_heads 8, intermediate_size 14336. Once you can read those five numbers, you can read any transformer config.

Common mistakes

Writing x = self.norm1(x) instead of h = self.norm1(x). Normalising x in place destroys the residual path, so the block replaces rather than adds. It trains, badly, and the bug is invisible in the shapes.

Forgetting is_causal=True in a decoder. Loss drops implausibly fast because every position can read its own answer. If your training loss looks too good, check this first.

Using one nn.LayerNorm object for both positions. They must be separate modules with separate weights. Sharing them silently ties two unrelated parameters together.

Assuming Block can be dropped into nn.Sequential with extra arguments. nn.Sequential passes one tensor. The trace flag above defaults to False for exactly that reason.

Try it yourself

Comment out x = x + a and replace it with x = a. Then stack twelve blocks and compare output magnitudes with the original. Then put it back and change mult from 4 to 1. Watch the parameter share flip from two thirds to one third.

What to learn next

Researcher — Mathematics and papers.

The block as a function

For input $X \in \mathbb{R}^{T \times d}$, the pre-norm block is:

$$ X' = X + \operatorname{MHA}(\operatorname{Norm}(X)) $$ $$ X'' = X' + \operatorname{FFN}(\operatorname{Norm}(X')) $$

Both sublayers map $\mathbb{R}^{T \times d} \to \mathbb{R}^{T \times d}$, so the block is an endomorphism and composition is unconstrained. That is the structural reason depth is a free parameter in this architecture.

Attention mixes across the $T$ axis and acts identically across the $d$ axis within a head. The feedforward layer mixes across the $d$ axis and acts identically across the $T$ axis. The block therefore alternates mixing along the two axes of the input. That is the structural motif of the MLP-Mixer (Tolstikhin et al., 2021, arXiv:2105.01601). It replaces attention with a second MLP and still works reasonably, which suggests the alternation itself carries some value.

Parameter accounting

Per block, ignoring norms and biases, with $d_{\text{ff}} = 4d$:

$$ P_{\text{attn}} = 4d^2, \qquad P_{\text{ffn}} = 8d^2, \qquad P_{\text{block}} = 12d^2 $$

The $12d^2$ per layer is the constant behind the standard $N \approx 12 L d^2$ estimate for non-embedding parameters. It is therefore also behind the $6ND$ training-compute rule. Gated feedforward layers change $P_{\text{ffn}}$ to $3 d\, d_{\text{ff}}$; grouped-query attention changes $P_{\text{attn}}$ to $2d^2 + 2 d\, g\, d_h$. See counting a model's parameters by hand.

Where the design has actually moved since 2017

The block's skeleton is unchanged. Five substitutions account for most of the difference between a 2017 encoder and a 2025 decoder:

Component2017Common now
Norm placementpost-normpre-norm
NormLayerNormRMSNorm
FeedforwardReLU, two matricesSwiGLU, three matrices
Positionadded sinusoidsrotary, applied inside attention
Key/value headsone per query headgrouped

Narang et al. (2021), arXiv:2102.11972, is the necessary counterweight to any list like this. Sweeping dozens of published modifications under one codebase, they found most did not reproduce their reported gains. Gated feedforward layers and RMSNorm were among the few that did.

Interpretability consequences of the additive form

Both sublayers write additively into a shared stream. The model output therefore decomposes exactly into a sum of per-component contributions:

$$ X^{(L)} = X^{(0)} + \sum_{\ell=1}^{L} \left[ \operatorname{MHA}^{(\ell)}(\cdot) + \operatorname{FFN}^{(\ell)}(\cdot) \right] $$

Elhage et al. (2021) build the residual-stream framework on this identity. The logit-lens family (nostalgebraist, 2020) exploits it by applying the output head to intermediate stream states. The decomposition is exact. What is not exact is treating the terms as independent, since each term's input depends on all preceding terms. See the residual stream.

Papers

What to learn next

What to learn next

These follow on from what you just read.

  • Inside a Transformer Block

    The residual stream

    Every layer writes into one shared running total rather than replacing it, which is why deep transformers train at all and why their internals can be pulled apart afterwards.

  • Inside a Transformer Block

    Layer normalisation

    Rescale each token's own numbers to a standard size, using only that token and nothing else in the batch, which is what keeps a deep stack of layers from drifting apart.

  • Inside a Transformer Block

    Pre-norm vs post-norm

    Normalising before a sublayer leaves the residual path untouched; normalising after it rescales the whole running total every layer - and that one choice decides whether a deep model trains easily.