AI glossary

RLHF (reinforcement learning from human feedback)

In one sentence RLHF tunes a language model using human preferences between answers, which is how raw text predictors became helpful assistants.

By Updated

RLHF trains a language model to produce answers humans prefer, by learning from comparisons between candidate answers rather than from correct-answer examples.

Some skills cannot be taught from an answer key, because no single right answer exists. You teach a new cook "which of these two dals tastes better?" — and from many such comparisons, taste emerges. "Helpful, honest, harmless" is a taste, not a formula, and RLHF is how that taste gets into a model.

The pipeline has three stages. First, instruction-tuning teaches a pretrained model the assistant format. Second, humans are shown pairs of model answers and pick the better one; a reward-model is trained on these picks, learning to score any answer the way the raters would. Third, reinforcement-learning (usually the PPO algorithm) tunes the LLM to generate answers the reward model scores highly — with a leash that punishes drifting too far from the original model, so it does not degenerate into reward-hacking gibberish.

pretrain → instruction-tune → collect human preferences
        → train reward model → RL against the reward model

This recipe, published with InstructGPT in 2022, produced ChatGPT and defines the modern assistant. Its known failure mode matters daily: models learn to please raters, and confident, agreeable, lengthy answers rate well — one root of sycophancy and polished hallucination. Simpler successors like DPO now compete with the full RL machinery.

Where to go next