Looking Inside a Trained Model

Probing hidden states

A probe is a small classifier trained on a model's hidden states, testing whether specific information is present at a given layer, before the model even finishes reading.

Read these first

On this page 5
  1. Why it exists
  2. How it works
  3. Where you have already seen it
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

A probe is a small test. It checks if specific information is present inside a model's numbers, at a chosen layer.

A doctor presses on your stomach to check for a specific problem. She never opens you up to look directly. The press tells her something is there, or is not, without seeing it herself.

A probe does the same thing to a model. It never reads the model's numbers directly. Instead, it trains a small classifier to guess a property from those numbers. Then it checks if that guess is any good.

Why it exists

A model's hidden numbers, its hidden states, are long lists of decimals. Staring at them tells a person nothing. Useful information might be encoded in there somewhere, but not in a form anyone can read by eye.

Probing sidesteps that problem entirely. Instead of reading the numbers, you ask a simple question. Can a small classifier, trained only on these numbers, correctly guess some property of the text?

If it can, that property is present in the numbers, in some usable form. If it cannot, either the property is missing, or it is encoded in a way the simple classifier cannot find.

How it works

  Sentence:  "the food was not good"    (true meaning: negative)

  Take the model's hidden numbers for this sentence, at layer 4.
  Train a tiny classifier: numbers -> positive or negative?

  If the classifier gets it right on sentences it never trained on,
  that layer's numbers really do carry the true sentiment,
  including the flip caused by the word "not".

Repeating this at every layer shows exactly where a property first becomes readable, layer by layer.

Where you have already seen it

  • Bias research. Checking if a model's hidden numbers encode a person's gender or ethnicity, even when that was never an explicit input.
  • Model auditing. Testing whether a model actually represents a concept it claims to understand, like grammatical tense.
  • Choosing which layer to reuse. When building a new tool on an existing model's hidden states, probing shows which layer holds the property you need.

Remember this

  • A probe is a small classifier trained to detect a property from a model's hidden numbers.
  • If the probe succeeds, the property is present in a usable form at that layer.
  • Running a probe at every layer shows where a property first becomes readable.

What to learn next

Developer — Code and libraries.

This probes for a genuinely hard property: negation. "Good" and "not good" share almost every word. This is a real test of whether a layer understands meaning, not only word presence.

Setup

bash
pip install transformers torch scikit-learn

The first run downloads distilbert-base-uncased, roughly 260 MB.

Probing every layer for negated sentiment

probe_demo.py
import random
import torch
from transformers import AutoTokenizer, AutoModel
from sklearn.linear_model import LogisticRegression

tok = AutoTokenizer.from_pretrained("distilbert-base-uncased")
model = AutoModel.from_pretrained("distilbert-base-uncased", output_hidden_states=True)
model.eval()

# The label is the TRUE sentiment after negation, not just "does a positive
# word appear". "not bad" is a positive sentence built from a negative word.
pos_words = ["good", "great", "wonderful", "excellent", "amazing", "fantastic"]
neg_words = ["bad", "terrible", "awful", "poor", "boring", "disappointing"]

sentences, labels = [], []
for w in pos_words:
    sentences += [f"the food was {w}", f"the food was not {w}"]
    labels += [1, 0]
for w in neg_words:
    sentences += [f"the food was {w}", f"the food was not {w}"]
    labels += [0, 1]

random.seed(0)
idx = list(range(len(sentences)))
random.shuffle(idx)
train_idx, test_idx = idx[:18], idx[18:]

inputs = tok(sentences, return_tensors="pt", padding=True)
with torch.no_grad():
    out = model(**inputs)
mask = inputs["attention_mask"].unsqueeze(-1)
per_layer = [(h * mask).sum(1) / mask.sum(1) for h in out.hidden_states]  # mean-pool each layer

for layer_idx, layer_vecs in enumerate(per_layer):
    X = layer_vecs.numpy()
    X_train, y_train = X[train_idx], [labels[i] for i in train_idx]
    X_test, y_test = X[test_idx], [labels[i] for i in test_idx]
    probe = LogisticRegression(max_iter=1000).fit(X_train, y_train)
    acc = probe.score(X_test, y_test)
    print(f"layer {layer_idx}: held-out probe accuracy = {acc:.2f}")
Output
layer 0: held-out probe accuracy = 0.17
layer 1: held-out probe accuracy = 0.17
layer 2: held-out probe accuracy = 0.17
layer 3: held-out probe accuracy = 0.17
layer 4: held-out probe accuracy = 0.83
layer 5: held-out probe accuracy = 0.83
layer 6: held-out probe accuracy = 0.67

This is a small, single run on 24 hand-written sentences, so treat the exact numbers as illustrative rather than precise. The shape is still striking: layers 0 through 3 score below chance (0.5), while layers 4 and 5 jump to 0.83.

Line by line

Layer 0 is the embedding layer, before any self-attention. Each word gets a vector, but words have not yet exchanged information with each other. A probe there is closer to a bag-of-words test than a meaning test.

Below-chance accuracy in early layers is a real, informative result, not a bug. It suggests the raw presence of the word "not" is actively misleading a linear probe at that stage, before the model has combined it correctly with the sentiment word.

mean-pool each layer averages every token's vector into one sentence-level vector, ignoring padding. This is a common simplification; a probe could instead use only the final token's vector, which would test a different, related question.

Common mistakes

Testing the probe on the same sentences it trained on. That measures memorisation, not whether the property is genuinely encoded. Always hold out separate test sentences, as done above.

Using a positive word as a stand-in label for genuine sentiment. Without the negated examples specifically included, a probe could succeed by detecting the presence of "good" alone, never actually testing whether negation is understood.

Drawing a firm conclusion from 24 sentences. This dataset is small enough that a couple of unlucky test examples can swing the accuracy considerably. Treat it as a demonstration of the method, not a benchmark result.

Try it yourself

Add a probe at the mean of all layers stacked together, instead of one layer at a time, and compare its accuracy against the best single layer.

This tests whether combining information across layers helps, or whether one layer alone already carries everything a linear probe can extract.

What to learn next

Researcher — Mathematics and papers.

Formal setup

Given a frozen, pretrained model producing hidden states h_l(x) at layer l for input x, and a labelled dataset {(x_i, y_i)} for some property y, a probe is a classifier p_theta trained to minimise:

text
min_theta  sum_i  Loss( p_theta(h_l(x_i)), y_i )
  • The model's own weights stay frozen throughout; only theta, the probe's parameters, are trained.
  • p_theta is deliberately kept simple, usually linear or a shallow MLP, so a high score reflects the representation's structure, not the probe's own capacity to learn.

The selectivity problem

Hewitt & Liang (2019), Designing and Interpreting Probes with Control Tasks, raised a foundational concern. A sufficiently powerful probe can achieve high accuracy on any property, even one assigned to inputs completely at random, by memorising input-to-label mappings alone. High probe accuracy does not, by itself, prove a representation genuinely encodes that property.

Their proposed fix, control tasks, assigns random labels to the same inputs and trains an identical probe on those. Selectivity is defined as the gap between real-task accuracy and control-task accuracy:

text
selectivity = accuracy(real labels) - accuracy(random labels)

A probe with high accuracy but low selectivity is likely exploiting its own capacity, not reading a genuine property of the representation. This is why the developer block deliberately uses a simple linear probe: it lowers the ceiling on what memorisation alone can achieve.

Probing versus causal methods

Probing is correlational: it shows a property is linearly decodable, not that the model uses that property in producing its output. A probe can succeed on information the model computed but never actually acted on downstream.

This is the central limitation that activation patching, covered earlier in this section, was built to address. The two techniques are complementary and commonly used together: probing to locate where a property might live, patching to confirm the model actually depends on it.

What probing has found

Tenney, Das & Pavlick (2019), BERT Rediscovers the Classical NLP Pipeline, probed BERT layer by layer, across part-of-speech, parsing, semantic roles and coreference. They found a rough ordering: shallow properties peak at earlier layers, abstract, longer-range ones peak later. This gave early support for the idea that transformers build structure hierarchically, echoing the classical NLP pipeline, without anyone designing them to do so.

Complexity

Training one probe is cheap: a linear classifier over a fixed-size representation, on a labelled dataset of modest size, trains in seconds on CPU. The cost scales with the number of layers and properties probed, since each combination needs its own probe, not with the size of the underlying frozen model.

Key references

  • Alain, G. & Bengio, Y. (2016). Understanding Intermediate Layers Using Linear Classifier Probes. arXiv:1610.01644
  • Hewitt, J. & Liang, P. (2019). Designing and Interpreting Probes with Control Tasks. arXiv:1909.03368
  • Tenney, I., Das, D. & Pavlick, E. (2019). BERT Rediscovers the Classical NLP Pipeline. arXiv:1905.05950

Current state and open problems

Probing remains a cheap, widely used first step for asking what a representation contains, particularly popular because it needs no access to model internals beyond hidden states, and works on closed-weight models through their exposed embeddings.

The open problem is the same selectivity concern Hewitt & Liang raised. As probes and probing datasets grow more sophisticated, separating "the model genuinely represents this" from "the probe learned to approximate it" remains an active methodological question, without a fully settled answer.

What to learn next