Reward model
In one sentence A reward model is a network trained on human preference data to score how good an answer is, standing in for human judgement at scale.
Updated
A reward model is a model trained to predict which answers humans would prefer, so it can hand out scores when no human is present.
Training an LLM with RLHF needs millions of judgements — far more than any team of humans can provide live. So the humans' taste is bottled. Raters compare pairs of answers for a while; a model is trained on those comparisons until it can predict the raters' choices; then this stand-in judges the millions of answers the LLM produces during training. It is the difference between the master chef tasting every plate and the master chef training a palate — then letting that palate run the kitchen's quality control.
Mechanically, a reward model is usually the LLM's own architecture with the text-prediction head swapped for a single-number output: text in, score out. It trains on (prompt, better answer, worse answer) triples, learning to score the better one higher.
The concept's importance goes beyond the training pipeline — it is where alignment's hardest problem becomes concrete. The reward model is an imperfect bottling of human intent, and the LLM optimises the bottle, not the intent. Gaps between them get exploited: confident tone, flattering agreement, and padded length all score well without being better — the phenomenon called reward hacking, and a root cause of sycophancy. Related uses of the same idea: LLM-as-judge evaluation, and best-of-n sampling, where the reward model picks the best of several generated drafts.
Where to go next
- Full lesson: RLHF
- Related terms: rlhf, dpo, alignment, llm-as-judge