Reward models
A reward model is a second network trained on which answer humans preferred, and it stands in for a human rater at a scale no human team could reach.
- 14 min read
- 3 reading levels
- Updated
Read these first
On this page 8
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
A reward model is a scoring machine. It is trained on pairs of answers, told which one a person liked more.
The analogy you have already lived
You are buying a shirt. The shopkeeper holds up two and asks which you prefer. You answer in a second.
Now imagine he asks you to score a single shirt out of ten instead. You hesitate. Is this a seven or an eight? Ask again tomorrow and you say something different.
People are reliable at comparing and unreliable at scoring. Reward models are built entirely around that fact.
Why it exists
You want a model that gives helpful answers. There is no way to write down a formula for "helpful". You know it when you see it.
So you get humans to see it. You show a person two answers to the same question and ask which is better. You collect tens of thousands of those judgements.
But a model produces millions of answers during training. No human team can judge all of them. So you train a second network to imitate the human's choices.
That second network is the reward model. It has one job. Read an answer and produce a number. A higher number means the humans would have liked it more.
How it works
question + answer A --> [ reward model ] --> 7.2
question + answer B --> [ reward model ] --> 4.1
|
A scores higher, so A winsTraining it is a comparison game. Show it a pair where a human chose A. Nudge A's score up and B's down. Repeat.
Notice what is not being learned. The reward model never learns that a good answer scores 7.2. It only learns to put better answers above worse ones. The absolute numbers mean nothing on their own.
The trap that ruins reward models
Suppose your labellers preferred the longer answer nine times out of ten. Not because length is good. Because the longer answers happened to be the thorough ones.
The reward model cannot tell those two stories apart. It sees "long won" over and over. So it learns "long is good".
Now you train your chat model to score highly. It writes longer and longer answers, saying less and less, and the reward model is delighted.
This is not a rare edge case. It is the single most common failure of the method, and the developer section below reproduces it in twenty lines.
Where you have already seen this
- A cooking competition judged by tasting two dishes side by side.
- Two photos of the same scene, where you can say which looks better but not why.
- A/B tests, where a website shows two versions and counts which one people click.
Remember this
- People compare well and score badly, so reward models are trained on comparisons.
- The scores are only meaningful relative to each other.
- The reward model learns whatever pattern predicted the human's choice, including patterns you did not intend.
What to learn next
- The KL penalty and the reference model — the leash that stops reward over-optimisation.
- Building preference data — where those comparisons come from.
- RLHF — the full loop this model sits inside.
Developer — Code and libraries.
Setup
pip install torchRuns on a CPU in a few seconds.
The reward-model loss, in full
For a preference pair, the loss is
loss = -F.logsigmoid(score(chosen) - score(rejected))That single line is the whole method. It is logistic regression on the difference of two scores.
Training one, and then breaking it
import torch
import torch.nn.functional as F
torch.manual_seed(0)
# Each answer is described by three features the labeller can react to.
FEATURES = ["correct", "polite", "long"]
TRUE_W = torch.tensor([3.0, 1.0, 0.0]) # labellers care about correctness, a little
# about politeness, and nothing about length
def sample(n, length_correlates):
correct = torch.randint(0, 2, (n, 1)).float()
polite = torch.randint(0, 2, (n, 1)).float()
if length_correlates:
# in this dataset the correct answers happen to also be the long ones
long = correct.clone()
else:
long = torch.randint(0, 2, (n, 1)).float()
return torch.cat([correct, polite, long], 1)
def make_pairs(n, length_correlates):
a, b = sample(n, length_correlates), sample(n, length_correlates)
ra, rb = a @ TRUE_W, b @ TRUE_W
# Bradley-Terry: P(a preferred) = sigmoid(r_a - r_b)
p = torch.sigmoid(ra - rb)
a_wins = (torch.rand(n) < p).float()
chosen = torch.where(a_wins[:, None].bool(), a, b)
rejected = torch.where(a_wins[:, None].bool(), b, a)
return chosen, rejected
def train_rm(chosen, rejected, steps=400):
w = torch.zeros(3, requires_grad=True)
opt = torch.optim.Adam([w], lr=0.05)
for _ in range(steps):
margin = chosen @ w - rejected @ w
loss = -F.logsigmoid(margin).mean() # the reward-model loss, in full
opt.zero_grad()
loss.backward()
opt.step()
return w.detach(), loss.item()
for corr in (False, True):
ch, rj = make_pairs(4000, corr)
w, loss = train_rm(ch, rj)
te_ch, te_rj = make_pairs(2000, corr)
acc = ((te_ch @ w) > (te_rj @ w)).float().mean()
tag = "length correlated with correctness" if corr else "length independent"
print(f"\n--- training data: {tag} ---")
print(f" final loss {loss:.4f}, held-out pair accuracy {acc:.1%}")
for name, learned, true in zip(FEATURES, w.tolist(), TRUE_W.tolist()):
print(f" weight on {name:<8} learned {learned:>6.2f} true {true:>5.2f}")
# now score four hand-made answers with the learned reward model
probe = torch.tensor([[1., 1., 0.], # correct, polite, short
[1., 0., 1.], # correct, blunt, long
[0., 1., 1.], # WRONG, polite, long
[0., 0., 0.]]) # wrong, blunt, short
labels = ["correct+polite+short", "correct+blunt+long",
"WRONG+polite+long", "wrong+blunt+short"]
print(" reward scores:")
for lab, r in zip(labels, (probe @ w).tolist()):
print(f" {lab:<22} {r:>6.2f}")--- training data: length independent ---
final loss 0.4237, held-out pair accuracy 72.0%
weight on correct learned 3.06 true 3.00
weight on polite learned 0.93 true 1.00
weight on long learned 0.03 true 0.00
reward scores:
correct+polite+short 3.99
correct+blunt+long 3.09
WRONG+polite+long 0.95
wrong+blunt+short 0.00
--- training data: length correlated with correctness ---
final loss 0.4222, held-out pair accuracy 68.9%
weight on correct learned 1.49 true 3.00
weight on polite learned 1.03 true 1.00
weight on long learned 1.49 true 0.00
reward scores:
correct+polite+short 2.52
correct+blunt+long 2.97
WRONG+polite+long 2.52
wrong+blunt+short 0.00Written against PyTorch 2.5.1, CPU. torch.manual_seed(0) fixes the sampled preferences, so this reproduces on this build; a different build may shift the last decimal.
This output is the whole lesson
With clean data, the reward model recovered the truth. Learned weights 3.06, 0.93, 0.03 against true 3.0, 1.0, 0.0. It was never told the true weights, only which of two answers a person picked.
Held-out accuracy was 72%, not 99%, and that is correct. The labels themselves are noisy — Bradley-Terry says a human picks the better answer with probability sigmoid(margin), not certainty. A reward model scoring 95% on real preference data is memorising annotator quirks, not learning quality. A ceiling near 70–75% is normal on human preference data and is not a bug.
With confounded data, the model split the credit. correct fell from 3.06 to 1.49 and long rose from 0.03 to 1.49. Nothing in the loss can separate two features that always move together.
Look at the last block of scores. WRONG+polite+long scores 2.52, exactly equal to correct+polite+short. The reward model now believes a wrong long answer is as good as a right short one. Optimise a chat model against this and you get reward hacking, on purpose, from a reward model with a perfectly respectable loss curve.
Both configurations had almost identical loss: 0.4237 and 0.4222. The loss cannot tell you which reward model is broken. Only probing it with cases you constructed can.
A real reward model
In practice the linear w becomes a language model with the language-modelling head replaced by a single scalar output, read from the final token's hidden state:
from transformers import AutoModelForSequenceClassification
model = AutoModelForSequenceClassification.from_pretrained(
"Qwen/Qwen3-0.6B", num_labels=1) # one output = a scalar rewardNo output block — this downloads a model and needs a GPU for anything real. What matters is num_labels=1. The loss is exactly the -logsigmoid(chosen - rejected) line above, applied to those scalars. TRL's RewardTrainer wraps it.
The reward model is usually initialised from the SFT model, so it starts already understanding the task's language. Some recipes freeze most layers and train only the head plus the top few blocks.
Common mistakes
Reading the scores as absolute. A reward of 4.1 means nothing on its own. The loss is invariant to adding a constant to every score, so the zero point is arbitrary and drifts between training runs. Compare scores only within one reward model, on one prompt.
Comparing across prompts. Reward models are not calibrated across different questions. A score of 3 on a hard question may be better than a 6 on an easy one.
Training on pairs where both answers are equal. Ties carry no gradient signal and add noise. Drop them, or use a margin-aware loss.
Never probing the model. Build a small set of adversarial pairs — same content, one much longer; same content, one with more bullet points; correct-but-terse against wrong-but-fluent. Score them before you trust the reward model with a training run.
Reusing a stale reward model. As the policy improves, its outputs leave the distribution the reward model was trained on, and the scores degrade. This is why RLHF pipelines refresh preference data between rounds.
Try it yourself
Set TRUE_W = torch.tensor([3.0, 1.0, -1.0]), so labellers actively dislike long answers, and keep length_correlates=True. Predict the learned weights before running it. The answer explains why "the labellers said they wanted short answers" does not save you from length bias.
What to learn next
- The KL penalty and the reference model — the leash that stops reward over-optimisation.
- Building preference data — where those comparisons come from.
- RLHF — the full loop this model sits inside.
Researcher — Mathematics and papers.
Bradley–Terry
The standard reward model assumes preferences follow the Bradley–Terry model (Bradley and Terry, 1952). For prompt $x$ and completions $y_1, y_2$:
$$ p(y_1 \succ y_2 \mid x) = \frac{\exp r^(x,y_1)}{\exp r^(x,y_1) + \exp r^(x,y_2)} = \sigma!\left(r^(x,y_1) - r^*(x,y_2)\right) $$
$r^*$ is the latent reward and $\sigma$ the logistic function. Fitting $r_\phi$ by maximum likelihood over a preference dataset $\mathcal{D} = {(x, y_w, y_l)}$ gives
$$ \mathcal{L}R(\phi) = -\mathbb{E}{(x,y_w,y_l)\sim\mathcal{D}}\left[\log \sigma!\left(r_\phi(x,y_w) - r_\phi(x,y_l)\right)\right] $$
Two structural properties follow immediately and both matter downstream.
Shift invariance. $r_\phi$ and $r_\phi + c(x)$ give identical loss for any function of the prompt alone. The reward scale has no absolute meaning, and per-prompt offsets are unidentifiable. This is why PPO-style RLHF normalises advantages per prompt and why GRPO subtracts a group mean.
Transitivity is assumed, not observed. Bradley–Terry cannot represent intransitive human preferences ($A \succ B \succ C \succ A$), which real annotator populations exhibit. Munos et al., 2024 (Nash Learning from Human Feedback) and the SPPO line drop the BT assumption and solve for a Nash equilibrium of the preference game instead.
Where the signal ceiling comes from
Inter-annotator agreement on open-ended preference data is typically 60–75%. InstructGPT (Ouyang et al., 2022) reported around 73% agreement between held-out human labellers, and their reward model's held-out accuracy was in the same range. A reward model cannot exceed the agreement rate of the labellers it was trained on, so accuracy well above that band is evidence of leakage or annotator-identity memorisation rather than quality.
RewardBench (Lambert et al., 2024) is the standard public benchmark. It reveals wide variance across categories — chat, reasoning, safety — and a persistent weakness on subtly-wrong reasoning, which is precisely the case where the reward signal is most needed.
Length bias, quantified
The confounding demonstrated in the code is documented at scale.
- Singhal et al., 2023 (A Long Way to Go: Investigating Length Correlations in RLHF) found reward-model score correlates with response length at $r > 0.7$ on standard preference datasets, and that a large fraction of RLHF's apparent improvement is reproducible by optimising length alone.
- Dubois et al., 2024 (Length-Controlled AlpacaEval) built a regression-based debiasing of the evaluation itself, and showed length control substantially reorders the leaderboard.
- Park et al., 2024 (Disentangling Length from Quality in Direct Preference Optimization) show the same pathology arises in DPO without any explicit reward model.
Mitigations that are actually used: length-normalising the reward, adding a length penalty, balancing lengths within preference pairs at collection time, and reporting length-controlled win rates. None of them fully removes the effect.
Reward over-optimisation
Gao et al., 2023 (Scaling Laws for Reward Model Overoptimization) is the key empirical result. Using a large "gold" reward model to label data for a smaller proxy reward model, they measured true reward as a function of KL distance $d = \sqrt{D_{\mathrm{KL}}(\pi | \pi_{\text{ref}})}$ from the reference policy, finding:
$$ R(d) = d\,(\alpha - \beta d) $$
for best-of-$n$ sampling, and a logarithmic analogue for RL. True reward rises, peaks, then falls as the policy exploits the proxy. The peak location scales with reward-model size and data; the shape does not go away. This is Goodhart's law with a fitted curve, and it is why the KL penalty exists.
Variants beyond a single scalar
| Approach | Idea | Trade-off |
|---|---|---|
| Ensemble of reward models | average or take a conservative quantile | reduces over-optimisation; $k\times$ cost |
| Multi-objective / multi-head | separate scores for helpfulness, safety, verbosity | interpretable; needs per-attribute labels |
| Process reward model (PRM) | score each reasoning step, not the final answer | far better on maths; step labels are expensive |
| Generative reward model / LLM judge | an LLM outputs a judgement in text | flexible; inherits the judge's biases |
| Rule-based verifier | no learned model at all | exact where applicable — see RLVR |
Lightman et al., 2023 (Let's Verify Step by Step) is the reference for process supervision: a PRM trained on 800k step-level labels solved 78% of a MATH test subset by reranking, substantially beating an outcome-supervised model. The cost is the label collection.
Rule-based verifiers are the most consequential recent shift. Where the answer can be checked mechanically — a unit test, a numeric match — the reward is exact, unhackable in the Bradley–Terry sense, and free. That is the foundation of the 2024–2025 reasoning models.
Papers
- Bradley and Terry, Rank Analysis of Incomplete Block Designs, Biometrika 1952.
- Christiano et al., Deep Reinforcement Learning from Human Preferences, NeurIPS 2017 — arxiv.org/abs/1706.03741
- Stiennon et al., Learning to summarize from human feedback, NeurIPS 2020 — arxiv.org/abs/2009.01325
- Ouyang et al., Training language models to follow instructions with human feedback, 2022 — arxiv.org/abs/2203.02155
- Gao et al., Scaling Laws for Reward Model Overoptimization, ICML 2023 — arxiv.org/abs/2210.10760
- Lightman et al., Let's Verify Step by Step, 2023 — arxiv.org/abs/2305.20050
- Singhal et al., A Long Way to Go: Investigating Length Correlations in RLHF, 2023 — arxiv.org/abs/2310.03716
- Lambert et al., RewardBench, 2024 — arxiv.org/abs/2403.13787
- Munos et al., Nash Learning from Human Feedback, ICML 2024 — arxiv.org/abs/2312.00886
What to learn next
- The KL penalty and the reference model — the leash that stops reward over-optimisation.
- Building preference data — where those comparisons come from.
- RLHF — the full loop this model sits inside.