RLHF — learning from human feedback
RLHF turns "which of these two answers do you prefer?" into a reward the model can be trained against — it is how a raw language model became a usable assistant.
- 19 min read
- 3 reading levels
- Updated
Read these first
On this page 10
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
RLHF trains a model on human preferences. People say which of two answers is better, and the model learns to produce more answers like the preferred one.
RLHF stands for reinforcement learning from human feedback.
The analogy you have already lived
Think about an eye test.
The optometrist does not ask you to rate the blurriness of a letter out of ten. Nobody can do that reliably. Instead they flip a lens and ask: "Better with one, or two?"
You always know the answer. Comparing two things you are looking at right now is easy. Scoring one thing on an absolute scale is very hard, and different people would give different numbers anyway.
After twenty comparisons, the optometrist has your prescription — a precise number — built entirely out of easy comparisons.
RLHF is exactly that, applied to a language model. Nobody can score a paragraph out of ten consistently. Everyone can pick which of two paragraphs they prefer.
Why this had to be invented
A language model is trained to predict the next word in text from the internet. That is all it is trained to do.
The internet is full of text that is unhelpful, confidently wrong, aggressive, or dangerous. A model that predicts internet text faithfully will produce all of those, because it is doing its job correctly.
Ask an untuned model how to reset a password. It might reply with another list of questions. On the internet, a question is often followed by more questions. Nothing has gone wrong. The model was never asked to be helpful. It was asked to predict.
Predicting text and being useful are different objectives. RLHF is how the second one gets added.
Why not write the rules yourself
You might think: write a reward. Give points for helpfulness, subtract for rudeness.
Try to write that down. What is the formula for "helpful"? How many points is a well-organised answer worth against a shorter one? What is the exact rudeness threshold?
Nobody can write it. It is the mango problem from the neural network lesson all over again — people recognise a good answer instantly and cannot state the rule they used.
So instead of writing the rule, you collect judgements and let a model learn the rule from them.
The three stages
STAGE 1 Show the model a few thousand good example answers,
written by people. It learns the shape of "an answer".
|
v
STAGE 2 For each question, show people two answers from the model.
They pick the better one. Train a second model - the reward
model - to predict which one people will pick.
|
v
STAGE 3 Now the reward model can score any answer instantly, without
a human. Use it as the reward, and train the language model
to produce answers that score well.Stage 1 is ordinary fine-tuning.
Stage 2 is the clever part. Humans are slow and expensive; you cannot ask a person to rate a million answers. So you ask them a few thousand times, then train a model to imitate their judgement. That model is a reward model — a stand-in for human taste that runs in milliseconds.
Stage 3 is where reinforcement learning finally appears. The language model is the agent. Writing the next word is the action. The reward model provides the reward. PPO is usually the algorithm.
The leash
There is a problem you can probably guess from the reward shaping lesson.
The reward model is a stand-in, not the real thing. Optimise any stand-in hard enough and you find the places where it disagrees with what it was standing in for. The model discovers phrasings that the reward model loves and no human likes — over-enthusiastic openings, hedging in every direction, padding an answer to look thorough.
So a second rule is added: do not drift too far from the model you started with. Every step away from the original wording is penalised. The model may improve, but it may not wander off into strange territory to chase a score.
That restraint is called a KL penalty, and it is not optional. Remove it and the model reliably degenerates into fluent nonsense that scores brilliantly. You will see a miniature version of exactly this in the code below.
What is honestly hard here
Whose preferences? A few hundred paid annotators, following a written guideline, decide what "better" means for a system used by hundreds of millions of people. Their instructions embed judgements about tone, caution and politics. This is a real and unresolved question, not a detail.
People prefer confident answers. Annotators reliably rate a fluent, assured wrong answer above a hedged correct one. Preference training therefore teaches confidence somewhat more than it teaches accuracy, which is one reason hallucination survives it.
People prefer longer answers. This is well documented, and it is why tuned assistants tend to pad.
Labels disagree. Two trained annotators agree only about three quarters of the time on which answer is better. The reward model is learning a noisy signal, and no amount of RL removes the noise.
RLHF made assistants usable. It did not make them correct, and confusing the two is the most common misunderstanding about it.
Where you have already seen it
Every mainstream AI assistant you have used has been through some version of this. The difference between a raw language model and the thing that answers your questions politely is, in large part, this procedure.
Remember this
- People give comparisons, not scores, because comparisons are the thing humans do reliably.
- A reward model learns to imitate those comparisons, so the reward becomes fast and automatic.
- A KL penalty keeps the model near where it started, or it games the reward model into nonsense.
What to learn next
- Fine-tuning — stage one of the pipeline, in detail.
- PPO — the algorithm doing stage three.
- Hallucination — the failure RLHF reduces but does not remove.
Developer — Code and libraries.
Setup
pip install torchStage 3 needs a GPU and a real language model, which is beyond what a lesson can run. Stage 2 is the conceptual heart of RLHF, it is small, and you can run it right now on a CPU in about two seconds.
Training a reward model from preferences
Five candidate replies to one question. Ten human preference pairs. A reward model that learns to score any reply.
import re
import torch
import torch.nn as nn
torch.manual_seed(0)
PROMPT = "how do I reset my password?"
REPLIES = {
"polite_steps": "Sure. Open settings, tap forgot password, then follow the link we email you.",
"curt_steps": "settings, forgot password, email link",
"honest_no": "I am not sure how that works for your account, sorry.",
"rude": "figure it out yourself, it is not hard",
"keyword_spam": "password password reset password settings password email",
}
# Pairs a human labelled. Each pair reads: the first reply is better than the second.
PREFERENCES = [
("polite_steps", "curt_steps"),
("polite_steps", "honest_no"),
("polite_steps", "rude"),
("polite_steps", "keyword_spam"),
("curt_steps", "honest_no"),
("curt_steps", "rude"),
("curt_steps", "keyword_spam"),
("honest_no", "rude"),
("honest_no", "keyword_spam"),
("keyword_spam", "rude"),
]
vocab = sorted({w for text in REPLIES.values() for w in re.findall(r"[a-z]+", text.lower())})
index = {w: i for i, w in enumerate(vocab)}
def features(text):
"""A bag of words. A real reward model reads the text with a transformer instead."""
v = torch.zeros(len(vocab))
for w in re.findall(r"[a-z]+", text.lower()):
if w in index: # a word it never saw carries no opinion at all
v[index[w]] += 1.0
return v / max(1.0, v.sum()) # normalise, so a long reply is not rewarded for length
X = {name: features(text) for name, text in REPLIES.items()}
reward_model = nn.Linear(len(vocab), 1, bias=False) # one number out: "how good is this reply"
opt = torch.optim.Adam(reward_model.parameters(), lr=0.05)
for step in range(600):
better = torch.stack([X[a] for a, _ in PREFERENCES])
worse = torch.stack([X[b] for _, b in PREFERENCES])
margin = reward_model(better) - reward_model(worse)
# Bradley-Terry: the chance a human prefers A over B is sigmoid(r(A) - r(B)).
loss = -torch.nn.functional.logsigmoid(margin).mean()
opt.zero_grad(); loss.backward(); opt.step()
if step % 200 == 0:
agree = (margin > 0).float().mean().item()
print(f"step {step:3d} loss {loss.item():.4f} agrees with the humans on {agree:.0%} of pairs")
print(f"\nvocabulary size: {len(vocab)} words")
print(f"every reply below answers: {PROMPT}\n")
print("learned reward for the replies humans ranked, best first:")
with torch.no_grad():
scored = sorted(((reward_model(X[n]).item(), n) for n in REPLIES), reverse=True)
for score, name in scored:
print(f" {score:+.3f} {name:<13} {REPLIES[name][:52]}")
unseen = {
"new_helpful": "Open settings and choose forgot password, then check your email.",
"new_rude": "not my problem, figure it out",
}
print("\nreplies the reward model has never seen:")
with torch.no_grad():
for name, text in unseen.items():
print(f" {reward_model(features(text)).item():+.3f} {name:<12} {text}")
# What happens if a policy optimises this reward with nothing holding it back?
with torch.no_grad():
top = torch.argsort(reward_model.weight[0], descending=True)[:3].tolist()
gibberish = (" ".join(vocab[i] for i in top) + " ") * 4
print("\nreward-hacking check")
print(f" best real reply scored : {scored[0][0]:+.3f}")
print(f" '{gibberish.strip()}'")
print(f" scores : {reward_model(features(gibberish)).item():+.3f}")step 0 loss 0.7013 agrees with the humans on 40% of pairs step 200 loss 0.0558 agrees with the humans on 100% of pairs step 400 loss 0.0250 agrees with the humans on 100% of pairs vocabulary size: 30 words every reply below answers: how do I reset my password? learned reward for the replies humans ranked, best first: +6.173 polite_steps Sure. Open settings, tap forgot password, then follo +2.866 curt_steps settings, forgot password, email link -0.355 honest_no I am not sure how that works for your account, sorry -3.567 keyword_spam password password reset password settings password e -7.167 rude figure it out yourself, it is not hard replies the reward model has never seen: +3.487 new_helpful Open settings and choose forgot password, then check your email. -7.204 new_rude not my problem, figure it out reward-hacking check best real reply scored : +6.173 'tap the sure tap the sure tap the sure tap the sure' scores : +8.353
What that output shows, in three parts
It recovered a full ranking from comparisons alone. Nobody ever told the model a number. It was given ten "this beats that" pairs and produced a scale: +6.17 down to -7.17, in exactly the order the humans implied. This is the eye test working — precise numbers built out of easy comparisons.
It generalises. The last two replies were never in the training data. A new helpful one scores +3.49; a new rude one scores -7.20. That generalisation is the entire reason a reward model is worth building. If it only scored the replies it had already seen, it would be a lookup table and useless for training anything.
And it can be gamed. The final block takes the three words the model weights most highly and repeats them. "tap the sure tap the sure..." is meaningless. It scores +8.35 — higher than the best genuine reply in the dataset.
That last number is the whole argument for the KL penalty, demonstrated in miniature. Point a policy at this reward with nothing holding it back, and it will produce that string, not a helpful answer. The reward model is not lying; it was trained on five replies and has never seen anything like this. It is being asked about a region it knows nothing about, and it answers with unfounded confidence.
Real reward models fail in exactly this way, on a larger and better-hidden scale.
Line by line
-logsigmoid(r_better - r_worse) — this is the Bradley-Terry loss and it is the whole of reward-model training. It assumes the chance a human prefers A over B is sigmoid(r(A) - r(B)). Minimising the negative log-likelihood pushes the gap between the preferred and rejected reply upward. That is the same loss used on real reward models with billions of parameters; only the feature extractor changes.
Only the difference matters. Add 100 to every reward and the loss is identical. Reward-model scores have no absolute meaning — a +6.17 is not "good", it is "better than the things it was compared against". Comparing raw scores across two separately trained reward models is meaningless.
v / max(1.0, v.sum()) — normalising by length. Remove it and the model can raise a reply's score by making it longer, which is a crude version of a genuine and well-documented RLHF failure: preference-tuned models pad.
nn.Linear(len(vocab), 1, bias=False) — no bias, because a constant shift changes nothing (see above). In production this layer sits on top of a transformer, usually initialised from the same checkpoint being tuned, with the language-model head replaced by a single scalar output.
What stage 3 looks like
The RL step is PPO, with one addition. Sketching the reward it optimises:
# not runnable on its own - it needs a real model, a tokenizer and a GPU
response = policy.generate(prompt)
score = reward_model(prompt, response)
# the leash: how far has the policy drifted from where it started?
kl = (policy.log_prob(response) - frozen_reference.log_prob(response)).sum()
reward = score - BETA * kl # BETA is typically 0.01 to 0.1frozen_reference is a copy of the model as it was after stage 1, never updated. The KL term measures how differently the two models would have written the same text. Set BETA to zero and you get the "tap the sure tap the sure" outcome at full scale.
DPO: the same goal without the RL
Direct preference optimisation (2023) makes an observation: the optimal policy under a KL-constrained reward objective has a closed form, so the reward model can be substituted out entirely. You can train directly on the preference pairs with one loss:
# pi = the model being trained, ref = the frozen stage-1 copy
logits = BETA * ((pi_logp_chosen - ref_logp_chosen) - (pi_logp_rejected - ref_logp_rejected))
loss = -torch.nn.functional.logsigmoid(logits).mean()No reward model, no sampling loop, no PPO. It is far simpler to run and it is now the common choice for smaller projects.
The honest comparison: DPO is easier and often matches PPO-based RLHF at moderate scale. PPO-based pipelines still lead at the largest scale, partly because they can generate fresh samples and score them, while DPO is limited to the preference data it was given. Neither is universally correct.
Common mistakes
Training the reward model to convergence on a small dataset. The output above shows 100% agreement on ten pairs — which means it has memorised them. A memorised reward model is maximally easy to game. Hold out preference pairs and stop when held-out accuracy stops improving. Real reward models plateau around 70 to 75% held-out accuracy, which is close to the rate at which two humans agree with each other.
Removing the KL penalty because scores go up. The reward-model score always goes up when you remove it. That is the symptom, not the success.
Not measuring anything but the reward. The reward model's score is a proxy. Track it alongside held-out preference accuracy and a human spot-check, or you will optimise a number that has quietly stopped meaning anything.
Assuming preferences are consistent. Two annotators agree about 73% of the time in published work. A reward model far above that on held-out data is fitting annotator noise, not human taste.
Try it yourself
Delete the / max(1.0, v.sum()) normalisation and re-run. Then score a reply that repeats a good word ten times. Watch the score climb with length alone. You will have reproduced, in ten seconds, the reason preference-tuned assistants write longer answers than they need to.
What to learn next
- Fine-tuning — stage one of the pipeline, in detail.
- PPO — the algorithm doing stage three.
- Hallucination — the failure RLHF reduces but does not remove.
Researcher — Mathematics and papers.
The pipeline, formally
Stage 1 — supervised fine-tuning. Maximise $\mathbb{E}{(x,y)\sim\mathcal{D}{\text{SFT}}}[\log \pi(y \mid x)]$ on curated demonstrations, giving $\pi^{\text{SFT}}$.
Stage 2 — reward modelling. Under the Bradley-Terry model (1952), for a prompt $x$ and completions $y_w \succ y_l$: $$ p(y_w \succ y_l \mid x) = \frac{\exp r^(x, y_w)}{\exp r^(x,y_w) + \exp r^(x,y_l)} = \sigma!\left(r^(x,y_w) - r^(x,y_l)\right) $$ Fit $r_\psi$ by maximum likelihood: $$ \mathcal{L}R(\psi) = -\mathbb{E}{(x, y_w, y_l) \sim \mathcal{D}}\left[ \log \sigma!\left( r_\psi(x, y_w) - r_\psi(x, y_l) \right) \right] $$ $r^$ is identifiable only up to an additive function of $x$, so scores carry no absolute meaning.
Stage 3 — RL against the reward model. $$ \max_{\pi_\theta} \; \mathbb{E}{x \sim \mathcal{D},\, y \sim \pi\theta(\cdot \mid x)} \left[ r_\psi(x,y) \right] - \beta\, \mathbb{D}{\text{KL}}!\left[ \pi\theta(y\mid x) \,|\, \pi^{\text{SFT}}(y \mid x) \right] $$ Solved with PPO, treating each token as an action, the prefix as the state, and the sequence-level reward assigned at the final token with the per-token KL folded into the reward.
The KL term is not a regulariser
It is part of the objective. The closed-form maximiser is $$ \pi^*(y \mid x) = \frac{1}{Z(x)} \pi^{\text{SFT}}(y \mid x) \exp!\left( \tfrac{1}{\beta} r_\psi(x,y) \right) $$ an exponential tilting of the reference policy. This identity is the foundation of DPO. It also shows $\beta$ selects a point on a frontier: $\beta \to \infty$ recovers $\pi^{\text{SFT}}$, and $\beta \to 0$ collapses onto the reward model's argmax, which is where reward hacking lives.
Reward over-optimisation
Gao, Schulman and Hilton (2023) measured proxy reward against gold reward using a large reward model as ground truth. With $d = \sqrt{\mathrm{KL}(\pi | \pi^{\text{SFT}})}$, gold reward follows $$ R_{\text{gold}}(d) = d\,(\alpha - \beta d) $$ for best-of-$n$ sampling, and a logarithmic variant for RL. It rises, peaks, and then falls. Larger reward models peak later and higher, but every one of them turns over.
The practical consequence is that KL distance is a usable early-warning signal, and that there is an optimal amount of optimisation which more compute does not remove.
DPO and the offline family
Rafailov et al. (2023) reparameterise the reward as $r(x,y) = \beta \log \frac{\pi_\theta(y \mid x)}{\pi^{\text{SFT}}(y\mid x)} + \beta \log Z(x)$. Substituting into the Bradley-Terry likelihood cancels $Z(x)$: $$ \mathcal{L}{\text{DPO}} = -\mathbb{E}\left[ \log \sigma!\left( \beta \log \frac{\pi\theta(y_w\mid x)}{\pi^{\text{SFT}}(y_w \mid x)} - \beta \log \frac{\pi_\theta(y_l \mid x)}{\pi^{\text{SFT}}(y_l\mid x)} \right) \right] $$
No reward model, no sampling, no RL. Known limitations: it is strictly offline, so it cannot correct itself on states it visits after updating; it is prone to reducing the likelihood of both completions in a pair; and the implicit reward is only well-behaved on the support of the preference data. Variants address these — IPO (Azar et al., 2023) removes the Bradley-Terry assumption, KTO (Ethayarajh et al., 2024) needs only binary good/bad labels, and ORPO folds the SFT stage in.
Removing the value network
For reasoning tasks with a verifiable answer, a learned critic over token sequences is expensive and badly conditioned. GRPO (Shao et al., 2024) samples a group of $G$ completions per prompt and uses the group's own statistics as the baseline: $$ \hat{A}_i = \frac{r_i - \operatorname{mean}(r_1 \dots r_G)}{\operatorname{std}(r_1 \dots r_G)} $$ This is REINFORCE with a group baseline plus PPO clipping. Combined with rule-based rewards — does the code compile, is the final answer correct — it removed the reward model from the loop entirely for a large class of tasks, and it underpins much of the recent work on reasoning models.
Scaling the feedback
Human labelling is the bottleneck. Two directions:
- Constitutional AI / RLAIF (Bai et al., 2022; Lee et al., 2023) — an LLM produces the preference labels against a written set of principles. Matches human-labelled RLHF on several benchmarks. It relocates the value judgements into the written constitution, which is arguably an improvement in transparency.
- Process supervision (Lightman et al., 2023) — label each reasoning step rather than the final answer. Substantially outperforms outcome supervision on mathematical reasoning, at much higher labelling cost.
Known failure modes, with citations
- Length bias. Singhal et al. (2023) show reward models correlate strongly with length and that much of the apparent gain from RLHF is explained by it.
- Sycophancy. Sharma et al. (2023) demonstrate that preference training increases agreement with a user's stated view, including when it is wrong, because annotators prefer agreement.
- Annotator disagreement. InstructGPT reports inter-annotator agreement near 73%, which is a ceiling on reward-model accuracy.
- Distributional narrowing. RLHF measurably reduces output diversity. For creative tasks this is a loss, and it is why some products expose a less-tuned mode.
- Open problems. Casper et al. (2023), Open Problems and Fundamental Limitations of RLHF, is the systematic catalogue and the right starting point for anyone treating RLHF as solved.
References
- Christiano et al. (2017), Deep Reinforcement Learning from Human Preferences — arxiv.org/abs/1706.03741.
- Stiennon et al. (2020), Learning to summarize from human feedback — arxiv.org/abs/2009.01325.
- Ouyang et al. (2022), Training language models to follow instructions with human feedback — arxiv.org/abs/2203.02155.
- Bai et al. (2022), Constitutional AI: Harmlessness from AI Feedback — arxiv.org/abs/2212.08073.
- Gao, Schulman and Hilton (2023), Scaling Laws for Reward Model Overoptimization — arxiv.org/abs/2210.10760.
- Rafailov et al. (2023), Direct Preference Optimization — arxiv.org/abs/2305.18290.
- Casper et al. (2023), Open Problems and Fundamental Limitations of RLHF — arxiv.org/abs/2307.15217.
- Shao et al. (2024), DeepSeekMath (GRPO) — arxiv.org/abs/2402.03300.
What to learn next
- Fine-tuning — stage one of the pipeline, in detail.
- PPO — the algorithm doing stage three.
- Hallucination — the failure RLHF reduces but does not remove.