Looking Inside a Trained Model
Superposition and polysemantic neurons
Superposition is a model cramming more features than it has room for, by sharing directions between features that rarely appear at the same time.
- 9 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Superposition is a model squeezing more concepts into its numbers than it has room for. It shares space between concepts that rarely show up together.
Think of a Swiss Army knife. It has far fewer tools than jobs it can do, because one blade folds out to cut, saw, and scrape. It shares one slot between jobs that rarely happen at once.
A model's neurons do the same trick. One neuron can respond to several unrelated things at once. There are always far more concepts than neurons to hold them.
Why it exists
A model has a fixed number of neurons. Language has an enormous number of concepts: countries, emotions, verb tenses, code syntax, and millions more. There are not nearly enough neurons for one each.
If two concepts almost never appear in the same input, a model can safely store them in overlapping space. It only causes confusion on the rare input where both concepts show up together.
This is not a design choice anyone made on purpose. It is a byproduct of squeezing a huge number of rare, sparse concepts into a much smaller number of neurons.
How it works
Neuron #482 fires strongly for:
- the word "bark" in "dog bark"
- the number "7" in a maths problem
- the colour blue in an unrelated sentence
These three concepts almost never appear in the same
sentence. So the neuron can represent all three,
usually without confusing them.A neuron that responds to several unrelated things like this is called polysemantic, meaning "many meanings". Superposition is the underlying reason polysemantic neurons exist at all.
Where you have already seen it
- Confusing AI explanations. When a tool tries to show "what a neuron detects" and the examples look unrelated, superposition is usually why.
- Unexpected model behaviour. A model that misbehaves on an unusual topic combination may be hitting two overlapping concepts colliding.
- Why interpretability is hard. This single phenomenon is a major reason "look at the neurons directly" does not work for understanding a model.
Remember this
- Superposition means a model stores more concepts than it has neurons, by sharing space between rarely co-occurring ones.
- A neuron that fires for several unrelated concepts is called polysemantic.
- This happens naturally from training, not from any deliberate design choice.
What to learn next
- Sparse autoencoders for feature discovery — a technique built specifically to untangle this overlap.
- Where a model stores a fact — the part of the network where this crowding happens most.
- Probing hidden states — why a probe can succeed even when a property is sharing space like this.
Developer — Code and libraries.
This rebuilds the classic toy experiment that demonstrates superposition directly: compress more features than you have room for, and watch what sparsity does to the result.
Setup
pip install torchSparsity is what makes superposition possible
import torch
torch.manual_seed(0)
n_features, n_hidden = 20, 5 # 20 things to represent, only 5 numbers to store them in
def train(sparsity, steps=2000):
W = torch.randn(n_hidden, n_features, requires_grad=True)
opt = torch.optim.Adam([W], lr=0.02)
for _ in range(steps):
x = torch.rand(512, n_features)
mask = (torch.rand(512, n_features) > sparsity).float() # most entries are 0
x = x * mask
hidden = x @ W.T # compress 20 features into 5 numbers
x_hat = torch.relu(hidden @ W) # try to reconstruct all 20 from those 5
loss = ((x - x_hat) ** 2).mean()
opt.zero_grad(); loss.backward(); opt.step()
return loss.item()
for sparsity in [0.0, 0.5, 0.9, 0.95]:
final_loss = train(sparsity)
print(f"sparsity={sparsity:.2f} (active {100*(1-sparsity):.0f}% of the time): final loss = {final_loss:.4f}")sparsity=0.00 (active 100% of the time): final loss = 0.0654 sparsity=0.50 (active 50% of the time): final loss = 0.0714 sparsity=0.90 (active 10% of the time): final loss = 0.0129 sparsity=0.95 (active 5% of the time): final loss = 0.0059
Same task throughout: cram 20 features into 5 numbers. The only thing that changed is how often each feature is actually present. As features get sparser, reconstruction loss drops more than ten times over. Sparsity is what makes cramming this many features in possible at all.
The exact loss numbers above can shift on a different machine, even with the seed fixed — floating-point summation order is not always identical across CPUs and library versions. The pattern holds regardless: loss should fall sharply, and by roughly the same order of magnitude, as sparsity climbs from 0 toward 0.95.
Line by line
W has shape (5, 20). Twenty feature directions, squeezed into only 5 numbers. There is no way to fit 20 genuinely independent directions into a 5-dimensional space; something has to overlap.
mask = (torch.rand(...) > sparsity) controls how often each feature is actually present in a given input. Higher sparsity means features are rarely active at the same time as each other.
Reconstruction loss measures how well the original 20 numbers can be recovered from the compressed 5. Low loss, despite the compression, is the signature of successful superposition.
Common mistakes
Assuming more neurons always means less superposition. Superposition depends on the ratio of features to neurons and how sparse those features are, not on neuron count alone. A bigger model with far more concepts to represent can show as much superposition as a smaller one does.
Treating one active neuron as one clean concept. As this experiment shows, a single dimension can carry several unrelated features cleanly, as long as they rarely overlap in the same input.
Expecting the reconstruction loss to reach exactly zero. Some interference is mathematically unavoidable once you compress more independent directions than the space allows, however sparse the features are.
Try it yourself
Add a line that measures the cosine similarity between different columns of the trained W, at sparsity=0.9. A value near 0 means two features got fully separate directions; a value further from 0 means they share space.
Expect a meaningfully non-zero average similarity, direct evidence that some features are genuinely sharing directions rather than each getting one to itself.
What to learn next
- Sparse autoencoders for feature discovery — training a second, wider network to pull these overlapping features back apart.
- Where a model stores a fact — the specific sublayer where this crowding is most studied.
- Probing hidden states — testing whether a specific feature survives this crowding well enough to be linearly readable.
Researcher — Mathematics and papers.
The toy model
Elhage et al. (2022), Toy Models of Superposition, define a minimal setting reproduced in the developer block: n sparse features x in R^n, compressed to m < n dimensions via W in R^{m x n}, and reconstructed through a ReLU output:
x_hat = ReLU(W^T W x + b)- Each feature
x_iis active with probability1 - S(its sparsityS) and drawn from[0, 1]when active,0otherwise. - The model is trained to minimise reconstruction loss, optionally weighted per feature by an importance term.
W^T Wis the key object: its off-diagonal entries measure how much two features interfere with each other's reconstruction.
Why sparsity enables superposition
At S = 0 (every feature always active), any two features sharing a direction interfere on every single input, since they are always both present. Loss from that interference cannot be avoided.
At high S, two features are almost never both active at once. Interference only bites on the rare input where both happen to fire together, so the expected cost of sharing a direction falls sharply, matching the loss curve reproduced in the developer block.
Elhage et al. show this trade-off produces a genuine phase transition as sparsity, feature importance and the dimensionality ratio n/m vary: below a critical sparsity, the model prefers a small number of orthogonal features, each with dedicated capacity. Above it, the model prefers superposition, packing many features into overlapping, non-orthogonal directions, exactly the geometry probed by the try-it-yourself cosine similarity exercise.
Consequences for interpretability
If a model represents n_features > n_neurons through superposition, no amount of staring at individual neurons in isolation can cleanly separate every feature, because the features are not axis-aligned. A neuron's activation is a mixture, and there is no guarantee the mixing coefficients are simple or few.
This directly motivates the shift, covered in the next lesson, away from studying neurons directly. Instead, a sparse autoencoder learns a wider basis that better matches the model's true feature directions.
Complexity
The toy model itself trains in seconds: O(m * n) parameters, gradient descent over a few thousand sparse random batches. The much harder problem, and the one under active research, is doing the analogous decomposition on a real trained model's activations, at far higher dimensionality and with no known ground-truth feature set to validate against.
Key references
- Elhage, N. et al. (2022). Toy Models of Superposition. Anthropic. transformer-circuits.pub
- Olah, C. et al. (2020). Zoom In: An Introduction to Circuits. Distill — the earlier observation of polysemantic neurons that motivated this line of work.
Current state and open problems
The toy model reproduced here is deliberately minimal. Real trained models are not guaranteed to organise their features the same way. Even so, empirical evidence for real superposition, and for the specific geometric patterns predicted here, has been found in genuine trained networks too.
The open problem is establishing, for any real model, exactly which features exist and how many there genuinely are. The sparse-dictionary techniques used to find them, the subject of the next lesson, might recover the model's real structure, or only a plausible-looking approximation of it. That question remains central and unresolved in this literature.
What to learn next
- Sparse autoencoders for feature discovery — the leading current technique for pulling superposed features back apart.
- Where a model stores a fact — the sublayer where this crowding is most concretely studied.
- Steering a model with activation vectors — using a found feature direction to change model behaviour on purpose.