Looking Inside a Trained Model
Steering a model with activation vectors
A steering vector is a direction added to a model's internal numbers while it runs, nudging its output toward a chosen trait without retraining anything.
- 9 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
A steering vector nudges a model toward a chosen trait. It adds one direction to its internal numbers while it runs.
Riding a bicycle, you steer by leaning your body a little to one side. You do not rebuild the bicycle. A small, well-placed push changes where it goes next.
A steering vector does the same thing to a model. It adds a small push, in one specific direction, to the model's numbers while it generates text. Nothing gets retrained.
Why it exists
Changing a model's behaviour usually means retraining it, which is slow and expensive. Sometimes you want a smaller, faster lever. Make this one response a bit more cheerful, or a bit more formal, right now.
Earlier lessons showed that concepts live along specific directions inside a model's numbers. Say a direction represents "happy" or "formal". Nudging the model's numbers along that direction should nudge its behaviour the same way.
Steering vectors turn that idea into a working tool: find the direction, then add it in, at generation time.
How it works
Step 1: find the direction.
Average the internal numbers for a pile of happy sentences.
Average the internal numbers for a pile of sad sentences.
Subtract one average from the other -- the leftover direction
is the steering vector, pointing from "sad" toward "happy".
Step 2: while generating text, add the steering vector
to the model's internal numbers at one layer.
Prompt: "The weather today is"
No steering: "...very good, and we're going to get some good weather"
+happy nudge: "...very nice and I am very happy with the weather!"
-happy nudge: "...bad. It's not good for you."The same prompt, nudged in opposite directions, produces noticeably different tone, without a single weight in the model being changed.
Where you have already seen it
- AI safety research. Testing whether pushing a model along a "deception" direction makes it more likely to produce misleading answers.
- Personality and tone controls. Some chatbot products let you dial a response toward formal or casual, using a technique like this.
- Research into what a model "knows". Steering along a concept direction and watching outputs change is a way to confirm the model actually represents that concept.
Remember this
- A steering vector is one direction, added to a model's numbers while it runs, to nudge behaviour.
- It is found by averaging and comparing activations for contrasting examples, like happy versus sad text.
- Push too hard and the output breaks down into repetition or nonsense, so strength needs tuning.
What to learn next
- Sparse autoencoders for feature discovery — a more precise way to find a clean direction to push along.
- Activation patching — the closely related technique of testing a direction's effect, rather than steering with it on purpose.
- Where a model stores a fact — another example of a targeted intervention on a model's internals.
Developer — Code and libraries.
This finds a "happy minus sad" direction from a handful of contrasting sentences, then adds it into GPT-2's generation at different strengths.
Setup
pip install transformers torchThe first run downloads gpt2, roughly 500 MB.
Building and applying a steering vector
import torch
from transformers import GPT2LMHeadModel, GPT2Tokenizer
tok = GPT2Tokenizer.from_pretrained("gpt2")
model = GPT2LMHeadModel.from_pretrained("gpt2")
model.eval()
LAYER = 6
happy = ["I feel so happy and excited today!", "This is the best day of my life, pure joy!",
"What a wonderful, delightful surprise!"]
sad = ["I feel so sad and hopeless today.", "This is the worst day of my life, pure misery.",
"What a terrible, upsetting disappointment."]
def mean_activation(sentences):
vecs = []
def hook(module, inputs, output):
vecs.append(output[:, -1, :].detach())
handle = model.transformer.h[LAYER].register_forward_hook(hook)
for s in sentences:
with torch.no_grad():
model(**tok(s, return_tensors="pt"))
handle.remove()
return torch.stack(vecs).mean(0)
steering_vector = mean_activation(happy) - mean_activation(sad)
def generate(prompt, strength=0.0):
handle = None
if strength != 0.0:
def add_vector(module, inputs, output):
return output + strength * steering_vector
handle = model.transformer.h[LAYER].register_forward_hook(add_vector)
inputs = tok(prompt, return_tensors="pt")
out = model.generate(**inputs, max_new_tokens=12, do_sample=False, pad_token_id=tok.eos_token_id)
if handle: handle.remove()
return tok.decode(out[0], skip_special_tokens=True).strip()
prompt = "The weather today is"
print("no steering: ", generate(prompt, 0.0))
print("+happy, mild: ", generate(prompt, 1.0))
print("-happy, mild: ", generate(prompt, -1.0))
print("+happy, hard: ", generate(prompt, 3.0))no steering: The weather today is very good, and we're going to be able to get +happy, mild: The weather today is very nice and I am very happy with the weather! -happy, mild: The weather today is bad. It's not good for you. It's not +happy, hard: The weather today is very! and we are very! and! and! and
The mild nudges work well: pushing toward happy makes the tone visibly warmer, pushing away makes it visibly more negative, both while staying fluent. Pushing three times harder breaks the sentence entirely.
Line by line
vecs.append(output[:, -1, :].detach()) grabs the last token's activation at the chosen layer, for each example sentence. Averaging these across several happy sentences gives a stable estimate of "what happy tends to look like here", rather than relying on any single sentence.
steering_vector = mean_activation(happy) - mean_activation(sad) is the entire technique: a single subtraction, producing a direction that points from one concept toward its opposite.
output + strength * steering_vector adds the direction into the residual stream at every position, for the rest of generation, not only the prompt. strength controls how hard the nudge is.
Common mistakes
Pushing too hard. As shown above, strength 3.0 collapses the sentence into exclamation marks. There is a real, model-specific ceiling on how hard a direction can be pushed before fluency breaks.
Building the direction from too few, too similar examples. Three sentences is small. A steering vector built from a wider variety of happy and sad text is a more robust, less example-specific direction.
Steering at the wrong layer. Early layers have not yet built up much semantic structure; very late layers are close to producing the actual output already. Middle layers, as used here, are the usual starting point.
Try it yourself
Try LAYER = 2 and LAYER = 10 instead of 6, at strength 1.0 for both. Compare the results against the middle-layer version above.
Expect a weaker or less coherent effect at one or both of the alternate layers. This is a direct, hands-on way to see why "which layer" is a real design decision in this technique, not an afterthought.
What to learn next
- Sparse autoencoders for feature discovery — finding a cleaner direction than a simple difference-of-means.
- Where a model stores a fact — a different kind of targeted intervention, aimed at facts instead of tone.
- Reward hacking — what happens when a training process, rather than a person, finds and exploits a similar shortcut.
Researcher — Mathematics and papers.
Difference-of-means steering
The technique in the developer block is CAA, contrastive activation addition (Rimsky et al., 2023, building on earlier "activation addition" work by Turner et al., 2023):
v = mean_{x in D_positive} a_l(x) - mean_{x in D_negative} a_l(x)
a_l'(x) = a_l(x) + alpha * va_l(x)is the residual stream activation at layerl, for inputx, typically at the final token position.D_positive,D_negativeare contrastive example sets, differing along the one trait being steered.alphais the steering strength, a free hyperparameter tuned per model, per layer and per trait.
This is a strict subset of the broader activation-engineering literature. The direction found is a difference-of-means direction, not a learned or optimised one. That makes it cheap to compute, but sensitive to the quality and diversity of the contrastive examples.
Why this works: the linear representation hypothesis
Steering by simple vector addition only makes sense if concepts are represented, at least approximately, as linear directions in activation space, an assumption known as the linear representation hypothesis. Park, Choe & Veitch (2023) provide supporting theoretical and empirical analysis, connecting it to earlier observations in static word embeddings (the classic king - man + woman ~= queen arithmetic).
This hypothesis is an approximation, not an exact law. Superposition, covered earlier in this section, predicts that not every concept gets a clean, dedicated, linear direction. That is part of why steering strength has a breaking point. Past some magnitude, the added vector starts interfering with other, unrelated directions sharing nearby space.
Relationship to sparse-autoencoder features
A difference-of-means direction is a coarse tool: it mixes together every way two example sets differ, not only the intended trait. An SAE-derived feature direction, from the previous lesson, is typically sharper, since it was optimised specifically to isolate one concept from the surrounding superposition.
Templeton et al. (2024), Scaling Monosemanticity, demonstrated steering using SAE features found in a production-scale model. They report more targeted behavioural changes than difference-of-means steering achieves at comparable strength, at the cost of needing a trained SAE for that model first.
Complexity
Computing a steering vector costs one forward pass per contrastive example, a one-time cost. Applying it at generation time adds a single vector addition per forward pass, at negligible extra compute over ordinary inference.
Key references
- Turner, A. et al. (2023). Activation Addition: Steering Language Models Without Optimization. arXiv:2308.10248
- Rimsky, N. et al. (2023). Steering Llama 2 via Contrastive Activation Addition. arXiv:2312.06681
- Park, K., Choe, Y. J. & Veitch, V. (2023). The Linear Representation Hypothesis and the Geometry of Large Language Models. arXiv:2311.03658
- Templeton, A. et al. (2024). Scaling Monosemanticity. Anthropic.
Current state and open problems
Steering vectors are a practical, low-cost intervention, used both as a research probe, confirming a model represents a concept by showing it can be pushed, and increasingly as a lightweight production tool for behaviour adjustment without fine-tuning.
The open problem is robustness. A steering vector tuned on one set of prompts can generalise poorly to structurally different ones. The breaking point where fluency collapses is not currently predictable in advance. It is only discoverable empirically, per model, per layer and per trait, as the try-it-yourself exercise above demonstrates directly.
What to learn next
- Sparse autoencoders for feature discovery — the sharper, more targeted alternative to a difference-of-means direction.
- Superposition and polysemantic neurons — the structural reason steering has a breaking point.
- Preference datasets — a heavier, training-time alternative for achieving a similar behavioural shift.