Post-training and Alignment

Reward models

A reward model is a second network trained on which answer humans preferred, and it stands in for a human rater at a scale no human team could reach.

On this page 8
  1. The short answer
  2. The analogy you have already lived
  3. Why it exists
  4. How it works
  5. The trap that ruins reward models
  6. Where you have already seen this
  7. Remember this
  8. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

A reward model is a scoring machine. It is trained on pairs of answers, told which one a person liked more.

The analogy you have already lived

You are buying a shirt. The shopkeeper holds up two and asks which you prefer. You answer in a second.

Now imagine he asks you to score a single shirt out of ten instead. You hesitate. Is this a seven or an eight? Ask again tomorrow and you say something different.

People are reliable at comparing and unreliable at scoring. Reward models are built entirely around that fact.

Why it exists

You want a model that gives helpful answers. There is no way to write down a formula for "helpful". You know it when you see it.

So you get humans to see it. You show a person two answers to the same question and ask which is better. You collect tens of thousands of those judgements.

But a model produces millions of answers during training. No human team can judge all of them. So you train a second network to imitate the human's choices.

That second network is the reward model. It has one job. Read an answer and produce a number. A higher number means the humans would have liked it more.

How it works

   question + answer A  -->  [ reward model ]  -->  7.2
   question + answer B  -->  [ reward model ]  -->  4.1
                                                     |
                                        A scores higher, so A wins

Training it is a comparison game. Show it a pair where a human chose A. Nudge A's score up and B's down. Repeat.

Notice what is not being learned. The reward model never learns that a good answer scores 7.2. It only learns to put better answers above worse ones. The absolute numbers mean nothing on their own.

The trap that ruins reward models

Suppose your labellers preferred the longer answer nine times out of ten. Not because length is good. Because the longer answers happened to be the thorough ones.

The reward model cannot tell those two stories apart. It sees "long won" over and over. So it learns "long is good".

Now you train your chat model to score highly. It writes longer and longer answers, saying less and less, and the reward model is delighted.

This is not a rare edge case. It is the single most common failure of the method, and the developer section below reproduces it in twenty lines.

Where you have already seen this

  • A cooking competition judged by tasting two dishes side by side.
  • Two photos of the same scene, where you can say which looks better but not why.
  • A/B tests, where a website shows two versions and counts which one people click.

Remember this

  • People compare well and score badly, so reward models are trained on comparisons.
  • The scores are only meaningful relative to each other.
  • The reward model learns whatever pattern predicted the human's choice, including patterns you did not intend.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install torch

Runs on a CPU in a few seconds.

The reward-model loss, in full

For a preference pair, the loss is

python
loss = -F.logsigmoid(score(chosen) - score(rejected))

That single line is the whole method. It is logistic regression on the difference of two scores.

Training one, and then breaking it

reward_model.py
import torch
import torch.nn.functional as F

torch.manual_seed(0)

# Each answer is described by three features the labeller can react to.
FEATURES = ["correct", "polite", "long"]
TRUE_W = torch.tensor([3.0, 1.0, 0.0])       # labellers care about correctness, a little
                                             # about politeness, and nothing about length


def sample(n, length_correlates):
    correct = torch.randint(0, 2, (n, 1)).float()
    polite = torch.randint(0, 2, (n, 1)).float()
    if length_correlates:
        # in this dataset the correct answers happen to also be the long ones
        long = correct.clone()
    else:
        long = torch.randint(0, 2, (n, 1)).float()
    return torch.cat([correct, polite, long], 1)


def make_pairs(n, length_correlates):
    a, b = sample(n, length_correlates), sample(n, length_correlates)
    ra, rb = a @ TRUE_W, b @ TRUE_W
    # Bradley-Terry: P(a preferred) = sigmoid(r_a - r_b)
    p = torch.sigmoid(ra - rb)
    a_wins = (torch.rand(n) < p).float()
    chosen = torch.where(a_wins[:, None].bool(), a, b)
    rejected = torch.where(a_wins[:, None].bool(), b, a)
    return chosen, rejected


def train_rm(chosen, rejected, steps=400):
    w = torch.zeros(3, requires_grad=True)
    opt = torch.optim.Adam([w], lr=0.05)
    for _ in range(steps):
        margin = chosen @ w - rejected @ w
        loss = -F.logsigmoid(margin).mean()      # the reward-model loss, in full
        opt.zero_grad()
        loss.backward()
        opt.step()
    return w.detach(), loss.item()


for corr in (False, True):
    ch, rj = make_pairs(4000, corr)
    w, loss = train_rm(ch, rj)
    te_ch, te_rj = make_pairs(2000, corr)
    acc = ((te_ch @ w) > (te_rj @ w)).float().mean()
    tag = "length correlated with correctness" if corr else "length independent"
    print(f"\n--- training data: {tag} ---")
    print(f"  final loss {loss:.4f}, held-out pair accuracy {acc:.1%}")
    for name, learned, true in zip(FEATURES, w.tolist(), TRUE_W.tolist()):
        print(f"  weight on {name:<8} learned {learned:>6.2f}   true {true:>5.2f}")

    # now score four hand-made answers with the learned reward model
    probe = torch.tensor([[1., 1., 0.],   # correct, polite, short
                          [1., 0., 1.],   # correct, blunt,  long
                          [0., 1., 1.],   # WRONG,   polite, long
                          [0., 0., 0.]])  # wrong,   blunt,  short
    labels = ["correct+polite+short", "correct+blunt+long",
              "WRONG+polite+long", "wrong+blunt+short"]
    print("  reward scores:")
    for lab, r in zip(labels, (probe @ w).tolist()):
        print(f"    {lab:<22} {r:>6.2f}")
Output
--- training data: length independent ---
  final loss 0.4237, held-out pair accuracy 72.0%
  weight on correct  learned   3.06   true  3.00
  weight on polite   learned   0.93   true  1.00
  weight on long     learned   0.03   true  0.00
  reward scores:
    correct+polite+short     3.99
    correct+blunt+long       3.09
    WRONG+polite+long        0.95
    wrong+blunt+short        0.00

--- training data: length correlated with correctness ---
  final loss 0.4222, held-out pair accuracy 68.9%
  weight on correct  learned   1.49   true  3.00
  weight on polite   learned   1.03   true  1.00
  weight on long     learned   1.49   true  0.00
  reward scores:
    correct+polite+short     2.52
    correct+blunt+long       2.97
    WRONG+polite+long        2.52
    wrong+blunt+short        0.00

Written against PyTorch 2.5.1, CPU. torch.manual_seed(0) fixes the sampled preferences, so this reproduces on this build; a different build may shift the last decimal.

This output is the whole lesson

With clean data, the reward model recovered the truth. Learned weights 3.06, 0.93, 0.03 against true 3.0, 1.0, 0.0. It was never told the true weights, only which of two answers a person picked.

Held-out accuracy was 72%, not 99%, and that is correct. The labels themselves are noisy — Bradley-Terry says a human picks the better answer with probability sigmoid(margin), not certainty. A reward model scoring 95% on real preference data is memorising annotator quirks, not learning quality. A ceiling near 70–75% is normal on human preference data and is not a bug.

With confounded data, the model split the credit. correct fell from 3.06 to 1.49 and long rose from 0.03 to 1.49. Nothing in the loss can separate two features that always move together.

Look at the last block of scores. WRONG+polite+long scores 2.52, exactly equal to correct+polite+short. The reward model now believes a wrong long answer is as good as a right short one. Optimise a chat model against this and you get reward hacking, on purpose, from a reward model with a perfectly respectable loss curve.

Both configurations had almost identical loss: 0.4237 and 0.4222. The loss cannot tell you which reward model is broken. Only probing it with cases you constructed can.

A real reward model

In practice the linear w becomes a language model with the language-modelling head replaced by a single scalar output, read from the final token's hidden state:

python
from transformers import AutoModelForSequenceClassification
model = AutoModelForSequenceClassification.from_pretrained(
    "Qwen/Qwen3-0.6B", num_labels=1)      # one output = a scalar reward

No output block — this downloads a model and needs a GPU for anything real. What matters is num_labels=1. The loss is exactly the -logsigmoid(chosen - rejected) line above, applied to those scalars. TRL's RewardTrainer wraps it.

The reward model is usually initialised from the SFT model, so it starts already understanding the task's language. Some recipes freeze most layers and train only the head plus the top few blocks.

Common mistakes

Reading the scores as absolute. A reward of 4.1 means nothing on its own. The loss is invariant to adding a constant to every score, so the zero point is arbitrary and drifts between training runs. Compare scores only within one reward model, on one prompt.

Comparing across prompts. Reward models are not calibrated across different questions. A score of 3 on a hard question may be better than a 6 on an easy one.

Training on pairs where both answers are equal. Ties carry no gradient signal and add noise. Drop them, or use a margin-aware loss.

Never probing the model. Build a small set of adversarial pairs — same content, one much longer; same content, one with more bullet points; correct-but-terse against wrong-but-fluent. Score them before you trust the reward model with a training run.

Reusing a stale reward model. As the policy improves, its outputs leave the distribution the reward model was trained on, and the scores degrade. This is why RLHF pipelines refresh preference data between rounds.

Try it yourself

Set TRUE_W = torch.tensor([3.0, 1.0, -1.0]), so labellers actively dislike long answers, and keep length_correlates=True. Predict the learned weights before running it. The answer explains why "the labellers said they wanted short answers" does not save you from length bias.

What to learn next

Researcher — Mathematics and papers.

Bradley–Terry

The standard reward model assumes preferences follow the Bradley–Terry model (Bradley and Terry, 1952). For prompt $x$ and completions $y_1, y_2$:

$$ p(y_1 \succ y_2 \mid x) = \frac{\exp r^(x,y_1)}{\exp r^(x,y_1) + \exp r^(x,y_2)} = \sigma!\left(r^(x,y_1) - r^*(x,y_2)\right) $$

$r^*$ is the latent reward and $\sigma$ the logistic function. Fitting $r_\phi$ by maximum likelihood over a preference dataset $\mathcal{D} = {(x, y_w, y_l)}$ gives

$$ \mathcal{L}R(\phi) = -\mathbb{E}{(x,y_w,y_l)\sim\mathcal{D}}\left[\log \sigma!\left(r_\phi(x,y_w) - r_\phi(x,y_l)\right)\right] $$

Two structural properties follow immediately and both matter downstream.

Shift invariance. $r_\phi$ and $r_\phi + c(x)$ give identical loss for any function of the prompt alone. The reward scale has no absolute meaning, and per-prompt offsets are unidentifiable. This is why PPO-style RLHF normalises advantages per prompt and why GRPO subtracts a group mean.

Transitivity is assumed, not observed. Bradley–Terry cannot represent intransitive human preferences ($A \succ B \succ C \succ A$), which real annotator populations exhibit. Munos et al., 2024 (Nash Learning from Human Feedback) and the SPPO line drop the BT assumption and solve for a Nash equilibrium of the preference game instead.

Where the signal ceiling comes from

Inter-annotator agreement on open-ended preference data is typically 60–75%. InstructGPT (Ouyang et al., 2022) reported around 73% agreement between held-out human labellers, and their reward model's held-out accuracy was in the same range. A reward model cannot exceed the agreement rate of the labellers it was trained on, so accuracy well above that band is evidence of leakage or annotator-identity memorisation rather than quality.

RewardBench (Lambert et al., 2024) is the standard public benchmark. It reveals wide variance across categories — chat, reasoning, safety — and a persistent weakness on subtly-wrong reasoning, which is precisely the case where the reward signal is most needed.

Length bias, quantified

The confounding demonstrated in the code is documented at scale.

  • Singhal et al., 2023 (A Long Way to Go: Investigating Length Correlations in RLHF) found reward-model score correlates with response length at $r > 0.7$ on standard preference datasets, and that a large fraction of RLHF's apparent improvement is reproducible by optimising length alone.
  • Dubois et al., 2024 (Length-Controlled AlpacaEval) built a regression-based debiasing of the evaluation itself, and showed length control substantially reorders the leaderboard.
  • Park et al., 2024 (Disentangling Length from Quality in Direct Preference Optimization) show the same pathology arises in DPO without any explicit reward model.

Mitigations that are actually used: length-normalising the reward, adding a length penalty, balancing lengths within preference pairs at collection time, and reporting length-controlled win rates. None of them fully removes the effect.

Reward over-optimisation

Gao et al., 2023 (Scaling Laws for Reward Model Overoptimization) is the key empirical result. Using a large "gold" reward model to label data for a smaller proxy reward model, they measured true reward as a function of KL distance $d = \sqrt{D_{\mathrm{KL}}(\pi | \pi_{\text{ref}})}$ from the reference policy, finding:

$$ R(d) = d\,(\alpha - \beta d) $$

for best-of-$n$ sampling, and a logarithmic analogue for RL. True reward rises, peaks, then falls as the policy exploits the proxy. The peak location scales with reward-model size and data; the shape does not go away. This is Goodhart's law with a fitted curve, and it is why the KL penalty exists.

Variants beyond a single scalar

ApproachIdeaTrade-off
Ensemble of reward modelsaverage or take a conservative quantilereduces over-optimisation; $k\times$ cost
Multi-objective / multi-headseparate scores for helpfulness, safety, verbosityinterpretable; needs per-attribute labels
Process reward model (PRM)score each reasoning step, not the final answerfar better on maths; step labels are expensive
Generative reward model / LLM judgean LLM outputs a judgement in textflexible; inherits the judge's biases
Rule-based verifierno learned model at allexact where applicable — see RLVR

Lightman et al., 2023 (Let's Verify Step by Step) is the reference for process supervision: a PRM trained on 800k step-level labels solved 78% of a MATH test subset by reranking, substantially beating an outcome-supervised model. The cost is the label collection.

Rule-based verifiers are the most consequential recent shift. Where the answer can be checked mechanically — a unit test, a numeric match — the reward is exact, unhackable in the Bradley–Terry sense, and free. That is the foundation of the 2024–2025 reasoning models.

Papers

What to learn next