Post-training and Alignment

Direct preference optimisation

DPO trains a model straight from human comparisons, skipping the reward model and the reinforcement learning loop entirely.

On this page 8
  1. The short answer
  2. The analogy you have already lived
  3. Why it exists
  4. How it works
  5. The catch nobody puts on the slide
  6. Where you have already seen this
  7. Remember this
  8. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

DPO trains the model straight from "this answer was better than that one". No reward model, no reinforcement learning.

The analogy you have already lived

Sit in an optician's chair. The lens wheel clicks around and you are asked, "Better with one, or with two?" You answer. Click. "One, or two?" You answer again.

The optician never asks you to name your prescription. Your comparisons are the adjustment. Twenty clicks later the lenses are right.

That is DPO. The old method wrote down a prescription first — the reward model. Then it fitted lenses to that written number. DPO adjusts straight from your answers.

Why it exists

The older method has three moving parts. Collect comparisons. Train a reward model on them. Then run reinforcement learning against that reward model.

Every one of those parts can go wrong.

Reinforcement learning is fiddly. It has many knobs, it is unstable, and small mistakes produce nonsense.

Four models sit in memory at once. The model being trained, the frozen reference, the reward model, and a value model. On a small GPU, that alone rules it out.

Errors compound. The reward model is an imperfect copy of human taste, and the reinforcement learning stage then over-optimises that imperfect copy.

Somebody noticed the maths could be rearranged. Write down what the reinforcement learning stage is trying to achieve. Solve it on paper. The reward model cancels out.

What is left is an ordinary training loss you can run like any other.

How it works

Take a question with two answers, one preferred by a human.

   question:  "Explain rainbows to a child."
   chosen  :  "Sunlight bends through raindrops and splits into colours."
   rejected:  "Rainbows are an optical phenomenon involving refraction."

At every step, DPO checks four numbers.

   how likely the CHOSEN answer is, under the model being trained
   how likely the CHOSEN answer is, under the frozen starting model
   how likely the REJECTED answer is, under the model being trained
   how likely the REJECTED answer is, under the frozen starting model

Then it pushes the chosen answer to become relatively more likely than the rejected one. Relative, that is, to where the frozen copy started.

The frozen copy is still there. It is doing the same leash job as before. What has disappeared is the reward model and the whole reinforcement learning machinery.

The catch nobody puts on the slide

Read this twice, because it is genuinely surprising.

DPO pushes the gap between chosen and rejected. It does not promise the chosen answer becomes more likely on its own.

In practice, both often become less likely, with the rejected one falling faster. The gap grows exactly as designed. Probability drains out of both, and goes somewhere you did not choose.

The developer section below shows this happening in a printout you can run yourself.

Where you have already seen this

  • An optician's "better with one, or two?"
  • A tailor adjusting from "this one fits better" rather than from measurements.
  • Any tool with a side-by-side A/B choice that quietly tunes itself.

Remember this

  • DPO learns straight from comparisons — no reward model, no reinforcement learning.
  • It still needs a frozen copy of the starting model.
  • It grows the gap between chosen and rejected, which is not the same as improving the chosen answer.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install torch

Runs on a CPU in a few seconds.

The DPO loss, complete

python
margin = beta * ((logp_chosen - ref_logp_chosen) - (logp_rejected - ref_logp_rejected))
loss = -F.logsigmoid(margin).mean()

Compare that with the reward-model loss, which was -logsigmoid(r_chosen - r_rejected). It is the same loss. The reward has been replaced by beta * log(pi / pi_ref) — the model's own log-probability ratio against the frozen reference.

Watching DPO recover a reward model it never had

dpo_from_scratch.py
import torch
import torch.nn.functional as F

torch.manual_seed(0)

ANSWERS = ["A: correct+clear", "B: correct+terse", "C: partly right",
           "D: wrong+fluent", "E: wrong+rambling", "F: refuses"]
REF_LOGITS = torch.tensor([0.4, 0.2, 0.6, 0.5, 0.1, 0.3])   # SFT policy, nearly uniform
TRUE_R = torch.tensor([2.0, 1.4, 0.6, -0.5, -1.2, 0.0])     # what humans actually think
BETA = 0.5

ref_logp = REF_LOGITS.log_softmax(0)


def sample_pairs(n):
    """Humans compare two sampled answers; Bradley-Terry decides the winner."""
    i = torch.multinomial(ref_logp.exp(), n, replacement=True)
    j = torch.multinomial(ref_logp.exp(), n, replacement=True)
    keep = i != j
    i, j = i[keep], j[keep]
    i_wins = torch.rand(len(i)) < torch.sigmoid(TRUE_R[i] - TRUE_R[j])
    return torch.where(i_wins, i, j), torch.where(i_wins, j, i)


win, lose = sample_pairs(6000)
print(f"{len(win)} preference pairs, no reward model anywhere in this script")

logits = REF_LOGITS.clone().requires_grad_(True)
opt = torch.optim.Adam([logits], lr=0.05)

print(f"\n{'step':>5} {'DPO loss':>9} {'log pi(chosen)':>15} {'log pi(rejected)':>17} {'margin':>8}")
for step in range(601):
    logp = logits.log_softmax(0)
    pi_w, pi_l = logp[win], logp[lose]
    ref_w, ref_l = ref_logp[win], ref_logp[lose]
    # the DPO loss, complete
    margin = BETA * ((pi_w - ref_w) - (pi_l - ref_l))
    loss = -F.logsigmoid(margin).mean()
    if step % 150 == 0:
        print(f"{step:>5} {loss.item():>9.4f} {pi_w.mean().item():>15.4f} "
              f"{pi_l.mean().item():>17.4f} {margin.mean().item():>8.4f}")
    opt.zero_grad()
    loss.backward()
    opt.step()

learned = logits.log_softmax(0).exp().detach()
# the closed-form optimum of the KL-regularised objective, for comparison
analytic = (ref_logp + TRUE_R / BETA).softmax(0)

print(f"\n{'answer':<18} {'true reward':>12} {'ref prob':>9} {'DPO prob':>9} {'closed form':>12}")
for a, r, p0, p1, p2 in zip(ANSWERS, TRUE_R, ref_logp.exp(), learned, analytic):
    print(f"{a:<18} {r:>12.1f} {p0:>9.3f} {p1:>9.3f} {p2:>12.3f}")

print(f"\nimplicit reward recovered by DPO, beta * log(pi/pi_ref):")
implicit = BETA * (logits.log_softmax(0).detach() - ref_logp)
implicit = implicit - implicit.mean() + TRUE_R.mean()      # rewards are shift-invariant
for a, r, ir in zip(ANSWERS, TRUE_R, implicit):
    print(f"  {a:<18} true {r:>5.1f}   recovered {ir:>6.2f}")
Output
4973 preference pairs, no reward model anywhere in this script

 step  DPO loss  log pi(chosen)  log pi(rejected)   margin
    0    0.6931         -1.7691           -1.7944   0.0000
  150    0.4649         -2.6598           -4.8383   1.0766
  300    0.4648         -2.7282           -4.9666   1.1065
  450    0.4648         -2.7283           -4.9668   1.1066
  600    0.4648         -2.7283           -4.9668   1.1066

answer              true reward  ref prob  DPO prob  closed form
A: correct+clear            2.0     0.173     0.829        0.743
B: correct+terse            1.4     0.141     0.121        0.183
C: partly right             0.6     0.211     0.037        0.055
D: wrong+fluent            -0.5     0.191     0.003        0.006
E: wrong+rambling          -1.2     0.128     0.001        0.001
F: refuses                  0.0     0.156     0.010        0.012

implicit reward recovered by DPO, beta * log(pi/pi_ref):
  A: correct+clear   true   2.0   recovered   2.23
  B: correct+terse   true   1.4   recovered   1.37
  C: partly right    true   0.6   recovered   0.58
  D: wrong+fluent    true  -0.5   recovered  -0.65
  E: wrong+rambling  true  -1.2   recovered  -1.30
  F: refuses         true   0.0   recovered   0.06

Written against PyTorch 2.5.1, CPU, seeded — reproducible on this build. The pair count varies with the seed because self-comparisons are dropped.

Four things in that output, and one of them is a warning

The implicit reward matched the truth. True 2.0, 1.4, 0.6, -0.5, -1.2, 0.0; recovered 2.23, 1.37, 0.58, -0.65, -1.30, 0.06. No reward model was ever trained. The reward is read off the trained policy as beta * log(pi/pi_ref) — this is not a metaphor, it is what the derivation says.

The mean was subtracted before printing, because reward is only defined up to a per-prompt constant. That is the same shift-invariance seen in the reward-model lesson.

The learned policy approximates the closed-form answer. DPO reached 0.829 on answer A where the KL-regularised optimum is 0.743. Close, not identical — with finite noisy comparisons and a policy free to over-concentrate, DPO overshoots. That overshoot is characteristic, not a bug in this script.

Loss flattened at 0.4648, not near zero. As with the reward model, the labels are noisy by construction. A DPO run whose loss approaches zero is memorising the pairs.

log pi(chosen) went DOWN, from −1.7691 to −2.7283. The chosen answers became about 2.6 times less likely, while the margin grew from 0 to 1.11. This is the documented behaviour of DPO, not an artefact of this toy: probability drained out of most answers and piled onto one. If your goal was "make good answers more likely", the objective you wrote does not say that.

Doing it for real

TRL (v1.12.0) reduces the whole thing to:

python
from trl import DPOTrainer
from datasets import load_dataset

trainer = DPOTrainer(
    model="Qwen/Qwen3-0.6B",
    train_dataset=load_dataset("trl-lib/ultrafeedback_binarized", split="train"),
)
trainer.train()

No output block — this needs a GPU and downloads several gigabytes; its logs are hardware-dependent.

The dataset format is three fields:

python
{"prompt": "The sky is", "chosen": " blue.", "rejected": " green."}

Defaults worth knowing, all on DPOConfig:

SettingDefaultNote
beta0.1higher means stay closer to the reference
loss_type["sigmoid"]the original DPO loss
learning_rate1e-6far lower than SFT's 2e-5
ref_modelNonedefaults to the initial policy, frozen
precompute_ref_log_probsFalseset True to drop the reference model from memory
max_length1024truncation

The metrics to watch during a DPO run are rewards/accuracies (fraction of pairs where the implicit chosen reward beats the rejected one) and rewards/margins. Also watch logps/chosen. If it falls steeply, the run is displacing likelihood rather than improving answers.

Common mistakes

Skipping SFT. DPO assumes the reference already produces reasonable text in the right format. Running DPO on a base model produces a well-ordered mess.

Using SFT's learning rate. 2e-5 will destroy a DPO run. Start at 1e-6 to 5e-6 for full fine-tuning, higher for adapters.

Preference pairs from a different model. DPO is an offline method, so the pairs come from wherever you collected them. If they came from a much stronger model, they are far off your policy's distribution, and the gradient signal is weak where it matters. Generating pairs from the model you are training — on-policy data — reliably works better.

Ignoring length. DPO shows the same length bias as reward models. If your chosen answers are systematically longer, you are training a length preference. loss_type="sigmoid_norm" (SimPO's length normalisation) is one mitigation.

Reading beta as PPO's beta. They play the same conceptual role and have different scales. DPO's 0.1 is not PPO's 0.1.

Not freezing the reference. Same trap as PPO. Verify ref_model.training is False and that no parameter requires grad.

Try it yourself

Change BETA to 2.0 and re-run. The learned distribution should stay much closer to the reference, and the recovered rewards should shrink toward zero. Work out why the ordering survives even though the magnitudes change.

What to learn next

Researcher — Mathematics and papers.

The derivation, in three steps

Step 1. The RLHF objective has the closed-form solution derived in the KL penalty lesson:

$$ \pi^*(y\mid x) = \frac{1}{Z(x)}\pi_{\text{ref}}(y\mid x)\exp!\left(\tfrac{1}{\beta}r(x,y)\right) $$

Step 2. Rearrange for the reward:

$$ r(x,y) = \beta \log \frac{\pi^*(y \mid x)}{\pi_{\text{ref}}(y\mid x)} + \beta \log Z(x) $$

Step 3. Substitute into the Bradley–Terry likelihood. The preference probability depends only on the difference of two rewards at the same prompt, and $\beta \log Z(x)$ is identical for both — so it cancels:

$$ p(y_w \succ y_l \mid x) = \sigma!\left(\beta\log\frac{\pi^(y_w|x)}{\pi_{\text{ref}}(y_w|x)} - \beta\log\frac{\pi^(y_l|x)}{\pi_{\text{ref}}(y_l|x)}\right) $$

Maximum likelihood on this gives the DPO loss:

$$ \mathcal{L}{\text{DPO}} = -\mathbb{E}{(x,y_w,y_l)}\left[\log\sigma!\left(\beta\log\frac{\pi_\theta(y_w|x)}{\pi_{\text{ref}}(y_w|x)} - \beta\log\frac{\pi_\theta(y_l|x)}{\pi_{\text{ref}}(y_l|x)}\right)\right] $$

The intractable partition function $Z(x)$ — the reason RLHF needed RL — disappears because it is prompt-dependent and reward-difference-invariant. That single cancellation is the entire contribution of the paper (Rafailov et al., 2023).

The gradient, and what it says about the toy result

$$ \nabla_\theta \mathcal{L}{\text{DPO}} = -\beta\,\mathbb{E}\Big[\underbrace{\sigma!\left(\hat r\theta(x,y_l) - \hat r_\theta(x,y_w)\right)}{\text{weight: high when the model is wrong}}\big(\nabla\theta\log\pi_\theta(y_w|x) - \nabla_\theta\log\pi_\theta(y_l|x)\big)\Big] $$

with $\hat r_\theta = \beta\log\frac{\pi_\theta}{\pi_{\text{ref}}}$. Two consequences.

Automatic example weighting. Pairs the model already ranks correctly contribute almost nothing. This is what makes DPO stable without a value network.

Only the difference of gradients is controlled. Nothing in the expression forces $\log \pi_\theta(y_w)$ upward in absolute terms. The toy run's fall from −1.77 to −2.73 is the objective behaving exactly as written.

Likelihood displacement

That last effect has a name and a literature.

Pal et al., 2024 (Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive) show that when $y_w$ and $y_l$ have small edit distance, DPO's gradient reduces $\log\pi_\theta(y_w)$, and propose DPOP — adding a term penalising the chosen log-probability dropping below the reference.

Razin et al., 2024 (Unintentional Unalignment: Likelihood Displacement in Direct Preference Optimization) go further and characterise where the displaced mass goes. It flows to responses with similar embeddings to the rejected one, which can include responses semantically opposite to the chosen one. Their headline case: training a model to refuse unsafe requests via DPO increased its unsafe-compliance rate from 4.6% to 94.4%. They introduce a centered hidden embedding similarity (CHES) score to identify and filter the pairs responsible.

This is the strongest practical reason to log logps/chosen throughout a DPO run.

The variant zoo, and what each fixes

LossChangeFixes
IPO (Azar et al., 2024)replace $\log\sigma$ with a squared loss to a target marginover-fitting to deterministic preferences
cDPO / robustlabel smoothing on the preferenceannotator noise
KTO (Ethayarajh et al., 2024)needs only "good"/"bad" labels, not pairspairwise data is expensive
ORPO (Hong et al., 2024)odds-ratio penalty added to the SFT loss; no reference modeltwo-stage pipeline; memory
SimPO (Meng et al., 2024)length-normalised average log-prob as the implicit reward, plus a margin; no referencelength bias; reference cost
DPOP / DPO-Positivepenalise chosen log-prob fallinglikelihood displacement
MPOweighted combination of several of the abovetask-specific trade-offs

TRL exposes most of these as loss_type on a single DPOConfig, and supports weighted combinations via loss_type=[...] with loss_weights=[...].

Does DPO match PPO?

The honest answer is: usually, and not always, and the gap is about data rather than about the loss.

  • Ivison et al., 2024 (Unpacking DPO and PPO) ran a controlled comparison and found PPO ahead on reasoning-heavy tasks, attributing most of the gap to PPO's online sampling.
  • Xu et al., 2024 (Is DPO Superior to PPO for LLM Alignment?) reach the same conclusion more sharply, arguing DPO's offline nature leaves it exposed to distribution shift, and that PPO's advantage is largest exactly where the SFT policy is weakest.
  • Tajwar et al., 2024 (Preference Fine-Tuning of LLMs Should Leverage Suboptimal, On-Policy Data) isolate the mechanism: on-policy sampling and negative gradients are what matter, not which loss you write down.

The practical synthesis, and what most open recipes now do: iterative DPO. Sample from the current policy, label the pairs (by humans, a reward model, or an LLM judge), run DPO, repeat. This recovers most of PPO's benefit while keeping DPO's simplicity, and it is what Llama 3's post-training and Tülu 3 both use.

Papers

What to learn next