Inside a Transformer Block

The residual stream

Every layer writes into one shared running total rather than replacing it, which is why deep transformers train at all and why their internals can be pulled apart afterwards.

On this page 6
  1. Why adding beats replacing
  2. What each layer really does
  3. The part that makes research possible
  4. The honest part
  5. Remember this
  6. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Every layer in a transformer adds to one shared running total instead of overwriting it.

Think of the khata at your local shop — the running account book. You buy milk, the shopkeeper writes it down. You buy bread, that gets added underneath. Nobody rubs out yesterday's entry to write today's.

At any moment the total is the sum of everything written so far. And because it is a sum, you can look back and see exactly what each purchase contributed.

A transformer works the same way. That shared running total has a name: the residual stream. It is one vector per token, which every layer reads from and adds to.

Why adding beats replacing

Imagine the other design, where each layer replaces the token's vector completely. Layer one produces something new, layer two throws it away and produces something else again.

Two things go wrong.

Information dies. Whatever the token originally was gets overwritten early and can never be recovered. A model forty layers deep would have no idea, by the end, which word it started with.

Learning stops. A model learns by tracing how a change at the end connects back to each layer. Through forty replacements, that trace gets multiplied down to nothing. Early layers receive no useful signal.

Addition fixes both. The original is always still in the total. And the trace back has a clean, unbroken path to every layer, because a sum passes signal straight through.

What each layer really does

Once you see the stream as a running total, the job of a layer changes shape in your head.

A layer does not compute the token's new meaning. It computes an adjustment, and posts it to the account.

   start:      the token's embedding
      + block 1's note
      + block 2's note
      + block 3's note
      ...
      + block 32's note
   ────────────────────────
   end:        what the model finally uses to pick the next word

Each note is usually small compared to the running total. A layer nudges. It does not rewrite.

The part that makes research possible

Because the end result is a sum, you can subtract.

Want to know what block seven contributed? Remove its note from the total and see what changes. This is called an ablation — deleting one part to measure what it was doing.

You cannot do this with a design that overwrites. There, removing a layer breaks everything after it, and you learn nothing. With a sum, each contribution can be inspected on its own.

Almost all of the research on what happens inside a language model relies on this one property.

The honest part

Calling it a "stream" makes it sound like a channel with a direction. It is not. It is one list of numbers per token, and the layers take turns adding to it.

Also, the notes are not independent. Block seven reads the total that blocks one to six produced. Remove block three and block seven would have written something different. So an ablation tells you what one layer added in that particular run. Change the earlier layers, and it would have added something else.

That distinction gets glossed over constantly, and it matters.

Remember this

  • Every layer adds to a shared running total; nothing is overwritten.
  • This keeps information alive and keeps the learning signal reaching early layers.
  • Because the result is a sum, individual layers can be removed and measured.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install torch

The stream as an explicit running total

residual_stream.py
import torch
import torch.nn as nn
import torch.nn.functional as F

torch.manual_seed(0)
d_model, n_layers = 64, 8

class Sublayer(nn.Module):
    """Stands in for 'attention' or 'feedforward'. What matters is the +."""
    def __init__(self):
        super().__init__()
        self.norm = nn.LayerNorm(d_model)
        self.fc1 = nn.Linear(d_model, 4 * d_model)
        self.fc2 = nn.Linear(4 * d_model, d_model)
    def delta(self, x):                       # what this layer WRITES to the stream
        return self.fc2(F.gelu(self.fc1(self.norm(x))))

layers = nn.ModuleList(Sublayer() for _ in range(n_layers))

stream = torch.randn(1, 6, d_model)           # the embeddings start the stream off
start = stream.clone()
deltas = []

print("the stream is a running total. each layer adds to it.")
print(f"{'layer':>6} {'||what it wrote||':>18} {'||stream after||':>18} {'write share':>12}")
for i, layer in enumerate(layers):
    d = layer.delta(stream)
    deltas.append(d)
    stream = stream + d                        # THE residual connection
    dn, sn = d.norm().item(), stream.norm().item()
    print(f"{i:>6} {dn:>18.3f} {sn:>18.3f} {dn/sn:>11.1%}")

total = start + torch.stack(deltas).sum(0)
print("\nis the final stream exactly the start plus every write?")
print("  max difference:", float((total - stream).abs().max()))

print("\nablation: delete one layer's write and see how much the end moves.")
base = stream
for i in range(n_layers):
    without = start + torch.stack([d for j, d in enumerate(deltas) if j != i]).sum(0)
    print(f"  drop layer {i}: relative change at the output "
          f"{float((without-base).norm()/base.norm()):.4f}")

print("\nhow much of the original token survives to the end?")
def cos(a, b):
    return F.cosine_similarity(a.flatten(), b.flatten(), dim=0).item()
print(f"  with the residual connection: cos(input, output) = {cos(start, stream):+.4f}")

h = start.clone()
for layer in layers:
    h = layer.delta(h)                         # replace the '+' with 'overwrite'
print(f"  with the '+' removed        : cos(input, output) = {cos(start, h):+.4f}")
Output
the stream is a running total. each layer adds to it.
 layer  ||what it wrote||   ||stream after||  write share
     0              3.733             20.475       18.2%
     1              4.127             20.806       19.8%
     2              4.322             21.283       20.3%
     3              3.993             21.675       18.4%
     4              3.851             21.938       17.6%
     5              4.018             22.216       18.1%
     6              4.019             22.707       17.7%
     7              3.854             23.255       16.6%

is the final stream exactly the start plus every write?
  max difference: 4.76837158203125e-07

ablation: delete one layer's write and see how much the end moves.
  drop layer 0: relative change at the output 0.1605
  drop layer 1: relative change at the output 0.1775
  drop layer 2: relative change at the output 0.1859
  drop layer 3: relative change at the output 0.1717
  drop layer 4: relative change at the output 0.1656
  drop layer 5: relative change at the output 0.1728
  drop layer 6: relative change at the output 0.1728
  drop layer 7: relative change at the output 0.1657

how much of the original token survives to the end?
  with the residual connection: cos(input, output) = +0.8741
  with the '+' removed        : cos(input, output) = -0.0203

The four results, in order of importance

+0.8741 against -0.0203. This is the headline. With the residual connection, the output still points substantially in the same direction as the input after eight layers. Remove the addition and the correlation is zero — the network has forgotten what it was given. Depth without a residual path is depth that destroys its own input.

4.77e-07 confirms the sum is exact. The final stream equals the starting vector plus every delta, to float32 precision. This identity is not an approximation or a helpful picture. It is what the arithmetic does, and it is the licence for everything that follows.

Each layer writes about 18 percent of the stream's size. Layers nudge. Suppose a layer wrote something the same size as the whole stream. It would be overwriting in all but name.

Every ablation moves the output by about 0.17, and no single layer is critical. In an untrained network, that uniformity is expected. In a trained model the pattern is very different. A small number of layers matter enormously, and most matter little. Run this experiment on a trained model and the flat profile becomes spiky.

Reading the stream mid-flight

The stream lives in the same space at every depth. So the model's own output head can be applied to an intermediate state. This is the logit lens:

python
# after block i, with `head` the model's output projection and `final_norm` its last norm
probs = head(final_norm(stream)).softmax(-1)

It shows what the model would predict if it stopped at that layer. In real models the prediction is usually vague in early layers and sharpens through the second half.

Treat this as a rough probe rather than a measurement. The head was trained on the final stream. Applying it earlier is out of distribution, and the results can mislead. Tuned-lens variants fit a small correction per layer for exactly this reason.

Practical consequences

The stream width is the model's bandwidth. Every layer competes to write into the same d_model numbers. This is why width and depth are scaled together rather than depth alone.

Gradient checkpointing is cheap here. The residual structure makes recomputing a block's activations during the backward pass straightforward. That is how large models fit into memory. See gradient accumulation for the related memory trade.

Hooks let you watch it live. Register a forward hook on each block to capture the stream at every depth without editing the model. See forward and backward hooks.

Common mistakes

Writing x = norm(x) before the sublayer. This overwrites the residual path with the normalised version, so the block replaces instead of adding. It is the single most common transformer implementation bug and produces no error.

In-place operations on the stream. x += delta inside a module can break autograd when the pre-addition value is needed for a gradient. Use x = x + delta. See in-place operations and autograd.

Reading an ablation as "this layer is unnecessary". It shows the effect of removing that layer while every other layer keeps behaving as it did. It says nothing about a model retrained without it.

Assuming the stream's scale is constant. It grows with depth in pre-norm models: 20.5 to 23.3 above. A real 32-layer model grows much more. Anything comparing raw magnitudes across layers must account for that.

Try it yourself

Scale each delta by 0.1 before adding it, mimicking the small-initialisation schemes used in very deep models. Watch the final cosine similarity climb toward 1.0. Then scale by 5.0 and watch it collapse. The size of what layers write, relative to the stream, is a real design knob.

What to learn next

Researcher — Mathematics and papers.

The identity

For a pre-norm stack of $L$ sublayers with functions $f_\ell$:

$$ x^{(L)} = x^{(0)} + \sum_{\ell=1}^{L} f_\ell!\left( \operatorname{Norm}(x^{(\ell-1)}) \right) $$

The stream is a fixed $d$-dimensional space that every component reads from and writes to. Elhage et al. (2021) describe it as a communication channel with no privileged basis. Components claim subspaces to write into, and later components read from those subspaces.

Two properties follow directly:

Linearity of contributions. Each term enters the final state additively, so contributions can be separated and projected onto the output head. Logit attribution methods depend on exactly this.

The gradient identity. Differentiating the recurrence $x^{(\ell)} = x^{(\ell-1)} + f_\ell(\cdot)$:

$$ \frac{\partial x^{(L)}}{\partial x^{(\ell)}} = \prod_{j=\ell+1}^{L} \left( I + J_j \right) $$

where $J_j$ is the Jacobian of sublayer $j$ with respect to the stream. The identity term guarantees a path with derivative exactly one. The gradient at layer $\ell$ cannot vanish through the depth of the stack alone. Without the residual, the product is $\prod J_j$, whose spectral norm generically decays or explodes geometrically. He et al. (2016), Deep Residual Learning, arXiv:1512.03385, made this argument for vision; the transformer inherits it unchanged.

Why the confound in ablation is not fixable by better ablation

The additive decomposition is exact, but the terms are not independent. $f_\ell$ takes $x^{(\ell-1)}$, which already contains every earlier write. Zero-ablating $f_k$ for $k < \ell$ therefore moves $f_\ell$ off distribution.

This is why the interpretability literature distinguishes:

  • Zero ablation — set the component's output to zero. Off distribution.
  • Mean ablation — replace with its dataset mean. Removes the component's variance while keeping its typical contribution.
  • Resample or activation patching — substitute the activation from a different input, so the network stays on distribution. Meng et al. (2022), arXiv:2202.05262, build on this, as does the wider causal-tracing literature.
  • Path patching patches a specific edge in the computational graph rather than a node. That isolates one route through the stream (Wang et al., 2022, arXiv:2211.00593).

Any claim of the form "head X does Y" should say which of these was used.

Norm growth and its consequences

In pre-norm models, $|x^{(\ell)}|$ grows roughly with $\sqrt{\ell}$. That holds when the writes are near-orthogonal and the delta scale is stable. Since each sublayer's input is normalised, a growing stream means each new write is a smaller relative perturbation. Later layers therefore have progressively less influence. The stream is sometimes described as becoming harder to write into with depth.

Several deep-model recipes counter this directly. DeepNorm (Wang et al., 2022, arXiv:2203.00555) scales the residual branch by a depth-dependent constant. It shows stable training to 1,000 layers. ReZero (Bachlechner et al., 2020, arXiv:2003.04887) multiplies each branch by a learned scalar initialised at zero. The network then begins as the identity.

Superposition and the width constraint

Every component writes into the same $d$ dimensions. The number of distinguishable features the stream can carry is bounded by geometry, not parameter count. Elhage et al. (2022), Toy Models of Superposition, argue that networks represent more features than dimensions. Features are placed in near-orthogonal directions, and interference is tolerated. Sparsity is what favours this.

This gives a concrete reading of the stream width as a bandwidth budget. It also explains why depth alone scales poorly against joint depth-width scaling.

Papers

What to learn next