Direct preference optimisation
DPO trains a model straight from human comparisons, skipping the reward model and the reinforcement learning loop entirely.
- 15 min read
- 3 reading levels
- Updated
Read these first
On this page 8
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
DPO trains the model straight from "this answer was better than that one". No reward model, no reinforcement learning.
The analogy you have already lived
Sit in an optician's chair. The lens wheel clicks around and you are asked, "Better with one, or with two?" You answer. Click. "One, or two?" You answer again.
The optician never asks you to name your prescription. Your comparisons are the adjustment. Twenty clicks later the lenses are right.
That is DPO. The old method wrote down a prescription first — the reward model. Then it fitted lenses to that written number. DPO adjusts straight from your answers.
Why it exists
The older method has three moving parts. Collect comparisons. Train a reward model on them. Then run reinforcement learning against that reward model.
Every one of those parts can go wrong.
Reinforcement learning is fiddly. It has many knobs, it is unstable, and small mistakes produce nonsense.
Four models sit in memory at once. The model being trained, the frozen reference, the reward model, and a value model. On a small GPU, that alone rules it out.
Errors compound. The reward model is an imperfect copy of human taste, and the reinforcement learning stage then over-optimises that imperfect copy.
Somebody noticed the maths could be rearranged. Write down what the reinforcement learning stage is trying to achieve. Solve it on paper. The reward model cancels out.
What is left is an ordinary training loss you can run like any other.
How it works
Take a question with two answers, one preferred by a human.
question: "Explain rainbows to a child."
chosen : "Sunlight bends through raindrops and splits into colours."
rejected: "Rainbows are an optical phenomenon involving refraction."At every step, DPO checks four numbers.
how likely the CHOSEN answer is, under the model being trained
how likely the CHOSEN answer is, under the frozen starting model
how likely the REJECTED answer is, under the model being trained
how likely the REJECTED answer is, under the frozen starting modelThen it pushes the chosen answer to become relatively more likely than the rejected one. Relative, that is, to where the frozen copy started.
The frozen copy is still there. It is doing the same leash job as before. What has disappeared is the reward model and the whole reinforcement learning machinery.
The catch nobody puts on the slide
Read this twice, because it is genuinely surprising.
DPO pushes the gap between chosen and rejected. It does not promise the chosen answer becomes more likely on its own.
In practice, both often become less likely, with the rejected one falling faster. The gap grows exactly as designed. Probability drains out of both, and goes somewhere you did not choose.
The developer section below shows this happening in a printout you can run yourself.
Where you have already seen this
- An optician's "better with one, or two?"
- A tailor adjusting from "this one fits better" rather than from measurements.
- Any tool with a side-by-side A/B choice that quietly tunes itself.
Remember this
- DPO learns straight from comparisons — no reward model, no reinforcement learning.
- It still needs a frozen copy of the starting model.
- It grows the gap between chosen and rejected, which is not the same as improving the chosen answer.
What to learn next
- Building preference data — the pairs this method eats.
- GRPO — the online method that replaced PPO for reasoning.
- Reward models — the component DPO removes, and why it still matters.
Developer — Code and libraries.
Setup
pip install torchRuns on a CPU in a few seconds.
The DPO loss, complete
margin = beta * ((logp_chosen - ref_logp_chosen) - (logp_rejected - ref_logp_rejected))
loss = -F.logsigmoid(margin).mean()Compare that with the reward-model loss, which was -logsigmoid(r_chosen - r_rejected). It is the same loss. The reward has been replaced by beta * log(pi / pi_ref) — the model's own log-probability ratio against the frozen reference.
Watching DPO recover a reward model it never had
import torch
import torch.nn.functional as F
torch.manual_seed(0)
ANSWERS = ["A: correct+clear", "B: correct+terse", "C: partly right",
"D: wrong+fluent", "E: wrong+rambling", "F: refuses"]
REF_LOGITS = torch.tensor([0.4, 0.2, 0.6, 0.5, 0.1, 0.3]) # SFT policy, nearly uniform
TRUE_R = torch.tensor([2.0, 1.4, 0.6, -0.5, -1.2, 0.0]) # what humans actually think
BETA = 0.5
ref_logp = REF_LOGITS.log_softmax(0)
def sample_pairs(n):
"""Humans compare two sampled answers; Bradley-Terry decides the winner."""
i = torch.multinomial(ref_logp.exp(), n, replacement=True)
j = torch.multinomial(ref_logp.exp(), n, replacement=True)
keep = i != j
i, j = i[keep], j[keep]
i_wins = torch.rand(len(i)) < torch.sigmoid(TRUE_R[i] - TRUE_R[j])
return torch.where(i_wins, i, j), torch.where(i_wins, j, i)
win, lose = sample_pairs(6000)
print(f"{len(win)} preference pairs, no reward model anywhere in this script")
logits = REF_LOGITS.clone().requires_grad_(True)
opt = torch.optim.Adam([logits], lr=0.05)
print(f"\n{'step':>5} {'DPO loss':>9} {'log pi(chosen)':>15} {'log pi(rejected)':>17} {'margin':>8}")
for step in range(601):
logp = logits.log_softmax(0)
pi_w, pi_l = logp[win], logp[lose]
ref_w, ref_l = ref_logp[win], ref_logp[lose]
# the DPO loss, complete
margin = BETA * ((pi_w - ref_w) - (pi_l - ref_l))
loss = -F.logsigmoid(margin).mean()
if step % 150 == 0:
print(f"{step:>5} {loss.item():>9.4f} {pi_w.mean().item():>15.4f} "
f"{pi_l.mean().item():>17.4f} {margin.mean().item():>8.4f}")
opt.zero_grad()
loss.backward()
opt.step()
learned = logits.log_softmax(0).exp().detach()
# the closed-form optimum of the KL-regularised objective, for comparison
analytic = (ref_logp + TRUE_R / BETA).softmax(0)
print(f"\n{'answer':<18} {'true reward':>12} {'ref prob':>9} {'DPO prob':>9} {'closed form':>12}")
for a, r, p0, p1, p2 in zip(ANSWERS, TRUE_R, ref_logp.exp(), learned, analytic):
print(f"{a:<18} {r:>12.1f} {p0:>9.3f} {p1:>9.3f} {p2:>12.3f}")
print(f"\nimplicit reward recovered by DPO, beta * log(pi/pi_ref):")
implicit = BETA * (logits.log_softmax(0).detach() - ref_logp)
implicit = implicit - implicit.mean() + TRUE_R.mean() # rewards are shift-invariant
for a, r, ir in zip(ANSWERS, TRUE_R, implicit):
print(f" {a:<18} true {r:>5.1f} recovered {ir:>6.2f}")4973 preference pairs, no reward model anywhere in this script
step DPO loss log pi(chosen) log pi(rejected) margin
0 0.6931 -1.7691 -1.7944 0.0000
150 0.4649 -2.6598 -4.8383 1.0766
300 0.4648 -2.7282 -4.9666 1.1065
450 0.4648 -2.7283 -4.9668 1.1066
600 0.4648 -2.7283 -4.9668 1.1066
answer true reward ref prob DPO prob closed form
A: correct+clear 2.0 0.173 0.829 0.743
B: correct+terse 1.4 0.141 0.121 0.183
C: partly right 0.6 0.211 0.037 0.055
D: wrong+fluent -0.5 0.191 0.003 0.006
E: wrong+rambling -1.2 0.128 0.001 0.001
F: refuses 0.0 0.156 0.010 0.012
implicit reward recovered by DPO, beta * log(pi/pi_ref):
A: correct+clear true 2.0 recovered 2.23
B: correct+terse true 1.4 recovered 1.37
C: partly right true 0.6 recovered 0.58
D: wrong+fluent true -0.5 recovered -0.65
E: wrong+rambling true -1.2 recovered -1.30
F: refuses true 0.0 recovered 0.06Written against PyTorch 2.5.1, CPU, seeded — reproducible on this build. The pair count varies with the seed because self-comparisons are dropped.
Four things in that output, and one of them is a warning
The implicit reward matched the truth. True 2.0, 1.4, 0.6, -0.5, -1.2, 0.0; recovered 2.23, 1.37, 0.58, -0.65, -1.30, 0.06. No reward model was ever trained. The reward is read off the trained policy as beta * log(pi/pi_ref) — this is not a metaphor, it is what the derivation says.
The mean was subtracted before printing, because reward is only defined up to a per-prompt constant. That is the same shift-invariance seen in the reward-model lesson.
The learned policy approximates the closed-form answer. DPO reached 0.829 on answer A where the KL-regularised optimum is 0.743. Close, not identical — with finite noisy comparisons and a policy free to over-concentrate, DPO overshoots. That overshoot is characteristic, not a bug in this script.
Loss flattened at 0.4648, not near zero. As with the reward model, the labels are noisy by construction. A DPO run whose loss approaches zero is memorising the pairs.
log pi(chosen) went DOWN, from −1.7691 to −2.7283. The chosen answers became about 2.6 times less likely, while the margin grew from 0 to 1.11. This is the documented behaviour of DPO, not an artefact of this toy: probability drained out of most answers and piled onto one. If your goal was "make good answers more likely", the objective you wrote does not say that.
Doing it for real
TRL (v1.12.0) reduces the whole thing to:
from trl import DPOTrainer
from datasets import load_dataset
trainer = DPOTrainer(
model="Qwen/Qwen3-0.6B",
train_dataset=load_dataset("trl-lib/ultrafeedback_binarized", split="train"),
)
trainer.train()No output block — this needs a GPU and downloads several gigabytes; its logs are hardware-dependent.
The dataset format is three fields:
{"prompt": "The sky is", "chosen": " blue.", "rejected": " green."}Defaults worth knowing, all on DPOConfig:
| Setting | Default | Note |
|---|---|---|
beta | 0.1 | higher means stay closer to the reference |
loss_type | ["sigmoid"] | the original DPO loss |
learning_rate | 1e-6 | far lower than SFT's 2e-5 |
ref_model | None | defaults to the initial policy, frozen |
precompute_ref_log_probs | False | set True to drop the reference model from memory |
max_length | 1024 | truncation |
The metrics to watch during a DPO run are rewards/accuracies (fraction of pairs where the implicit chosen reward beats the rejected one) and rewards/margins. Also watch logps/chosen. If it falls steeply, the run is displacing likelihood rather than improving answers.
Common mistakes
Skipping SFT. DPO assumes the reference already produces reasonable text in the right format. Running DPO on a base model produces a well-ordered mess.
Using SFT's learning rate. 2e-5 will destroy a DPO run. Start at 1e-6 to 5e-6 for full fine-tuning, higher for adapters.
Preference pairs from a different model. DPO is an offline method, so the pairs come from wherever you collected them. If they came from a much stronger model, they are far off your policy's distribution, and the gradient signal is weak where it matters. Generating pairs from the model you are training — on-policy data — reliably works better.
Ignoring length. DPO shows the same length bias as reward models. If your chosen answers are systematically longer, you are training a length preference. loss_type="sigmoid_norm" (SimPO's length normalisation) is one mitigation.
Reading beta as PPO's beta. They play the same conceptual role and have different scales. DPO's 0.1 is not PPO's 0.1.
Not freezing the reference. Same trap as PPO. Verify ref_model.training is False and that no parameter requires grad.
Try it yourself
Change BETA to 2.0 and re-run. The learned distribution should stay much closer to the reference, and the recovered rewards should shrink toward zero. Work out why the ordering survives even though the magnitudes change.
What to learn next
- Building preference data — the pairs this method eats.
- GRPO — the online method that replaced PPO for reasoning.
- Reward models — the component DPO removes, and why it still matters.
Researcher — Mathematics and papers.
The derivation, in three steps
Step 1. The RLHF objective has the closed-form solution derived in the KL penalty lesson:
$$ \pi^*(y\mid x) = \frac{1}{Z(x)}\pi_{\text{ref}}(y\mid x)\exp!\left(\tfrac{1}{\beta}r(x,y)\right) $$
Step 2. Rearrange for the reward:
$$ r(x,y) = \beta \log \frac{\pi^*(y \mid x)}{\pi_{\text{ref}}(y\mid x)} + \beta \log Z(x) $$
Step 3. Substitute into the Bradley–Terry likelihood. The preference probability depends only on the difference of two rewards at the same prompt, and $\beta \log Z(x)$ is identical for both — so it cancels:
$$ p(y_w \succ y_l \mid x) = \sigma!\left(\beta\log\frac{\pi^(y_w|x)}{\pi_{\text{ref}}(y_w|x)} - \beta\log\frac{\pi^(y_l|x)}{\pi_{\text{ref}}(y_l|x)}\right) $$
Maximum likelihood on this gives the DPO loss:
$$ \mathcal{L}{\text{DPO}} = -\mathbb{E}{(x,y_w,y_l)}\left[\log\sigma!\left(\beta\log\frac{\pi_\theta(y_w|x)}{\pi_{\text{ref}}(y_w|x)} - \beta\log\frac{\pi_\theta(y_l|x)}{\pi_{\text{ref}}(y_l|x)}\right)\right] $$
The intractable partition function $Z(x)$ — the reason RLHF needed RL — disappears because it is prompt-dependent and reward-difference-invariant. That single cancellation is the entire contribution of the paper (Rafailov et al., 2023).
The gradient, and what it says about the toy result
$$ \nabla_\theta \mathcal{L}{\text{DPO}} = -\beta\,\mathbb{E}\Big[\underbrace{\sigma!\left(\hat r\theta(x,y_l) - \hat r_\theta(x,y_w)\right)}{\text{weight: high when the model is wrong}}\big(\nabla\theta\log\pi_\theta(y_w|x) - \nabla_\theta\log\pi_\theta(y_l|x)\big)\Big] $$
with $\hat r_\theta = \beta\log\frac{\pi_\theta}{\pi_{\text{ref}}}$. Two consequences.
Automatic example weighting. Pairs the model already ranks correctly contribute almost nothing. This is what makes DPO stable without a value network.
Only the difference of gradients is controlled. Nothing in the expression forces $\log \pi_\theta(y_w)$ upward in absolute terms. The toy run's fall from −1.77 to −2.73 is the objective behaving exactly as written.
Likelihood displacement
That last effect has a name and a literature.
Pal et al., 2024 (Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive) show that when $y_w$ and $y_l$ have small edit distance, DPO's gradient reduces $\log\pi_\theta(y_w)$, and propose DPOP — adding a term penalising the chosen log-probability dropping below the reference.
Razin et al., 2024 (Unintentional Unalignment: Likelihood Displacement in Direct Preference Optimization) go further and characterise where the displaced mass goes. It flows to responses with similar embeddings to the rejected one, which can include responses semantically opposite to the chosen one. Their headline case: training a model to refuse unsafe requests via DPO increased its unsafe-compliance rate from 4.6% to 94.4%. They introduce a centered hidden embedding similarity (CHES) score to identify and filter the pairs responsible.
This is the strongest practical reason to log logps/chosen throughout a DPO run.
The variant zoo, and what each fixes
| Loss | Change | Fixes |
|---|---|---|
| IPO (Azar et al., 2024) | replace $\log\sigma$ with a squared loss to a target margin | over-fitting to deterministic preferences |
| cDPO / robust | label smoothing on the preference | annotator noise |
| KTO (Ethayarajh et al., 2024) | needs only "good"/"bad" labels, not pairs | pairwise data is expensive |
| ORPO (Hong et al., 2024) | odds-ratio penalty added to the SFT loss; no reference model | two-stage pipeline; memory |
| SimPO (Meng et al., 2024) | length-normalised average log-prob as the implicit reward, plus a margin; no reference | length bias; reference cost |
| DPOP / DPO-Positive | penalise chosen log-prob falling | likelihood displacement |
| MPO | weighted combination of several of the above | task-specific trade-offs |
TRL exposes most of these as loss_type on a single DPOConfig, and supports weighted combinations via loss_type=[...] with loss_weights=[...].
Does DPO match PPO?
The honest answer is: usually, and not always, and the gap is about data rather than about the loss.
- Ivison et al., 2024 (Unpacking DPO and PPO) ran a controlled comparison and found PPO ahead on reasoning-heavy tasks, attributing most of the gap to PPO's online sampling.
- Xu et al., 2024 (Is DPO Superior to PPO for LLM Alignment?) reach the same conclusion more sharply, arguing DPO's offline nature leaves it exposed to distribution shift, and that PPO's advantage is largest exactly where the SFT policy is weakest.
- Tajwar et al., 2024 (Preference Fine-Tuning of LLMs Should Leverage Suboptimal, On-Policy Data) isolate the mechanism: on-policy sampling and negative gradients are what matter, not which loss you write down.
The practical synthesis, and what most open recipes now do: iterative DPO. Sample from the current policy, label the pairs (by humans, a reward model, or an LLM judge), run DPO, repeat. This recovers most of PPO's benefit while keeping DPO's simplicity, and it is what Llama 3's post-training and Tülu 3 both use.
Papers
- Rafailov et al., Direct Preference Optimization: Your Language Model is Secretly a Reward Model, NeurIPS 2023 — arxiv.org/abs/2305.18290
- Azar et al., A General Theoretical Paradigm to Understand Learning from Human Preferences (IPO), AISTATS 2024 — arxiv.org/abs/2310.12036
- Ethayarajh et al., KTO: Model Alignment as Prospect Theoretic Optimization, ICML 2024 — arxiv.org/abs/2402.01306
- Hong et al., ORPO: Monolithic Preference Optimization without Reference Model, EMNLP 2024 — arxiv.org/abs/2403.07691
- Pal et al., Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive, 2024 — arxiv.org/abs/2402.13228
- Meng et al., SimPO, NeurIPS 2024 — arxiv.org/abs/2405.14734
- Tajwar et al., Preference Fine-Tuning of LLMs Should Leverage Suboptimal, On-Policy Data, ICML 2024 — arxiv.org/abs/2404.14367
- Ivison et al., Unpacking DPO and PPO, NeurIPS 2024 — arxiv.org/abs/2406.09279
- Xu et al., Is DPO Superior to PPO for LLM Alignment?, ICML 2024 — arxiv.org/abs/2404.10719
- Razin et al., Unintentional Unalignment: Likelihood Displacement in DPO, ICLR 2025 — arxiv.org/abs/2410.08847
What to learn next
- Building preference data — the pairs this method eats.
- GRPO — the online method that replaced PPO for reasoning.
- Reward models — the component DPO removes, and why it still matters.