Post-training and Alignment

Building preference data

Preference data is a pile of "this answer beat that one" judgements, and almost every alignment failure can be traced back to how it was collected.

On this page 10
  1. The short answer
  2. The analogy you have already lived
  3. Why it exists
  4. How it is made
  5. Who does the labelling
  6. The trap that catches everyone
  7. What is honestly hard here
  8. Where you have already seen this
  9. Remember this
  10. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

Preference data is a list of pairs of answers, each with a note saying which one a person liked better.

The analogy you have already lived

Walk past the sampling counter in a big supermarket. Two paper cups, two versions of the same drink, and someone asking which you prefer.

Notice what they do not ask. Nobody asks you to rate the drink out of ten. Nobody asks you to explain your answer. One question, one finger point, five seconds.

That is how preference data is collected. Simple question, enormous number of repetitions.

Why it exists

Every alignment method you will meet eats this same food. Reward models eat it. DPO eats it. So does almost everything else.

That makes the data the most important artefact in the whole pipeline, and it is usually the least examined.

Model architecture gets papers. Data collection gets a spreadsheet nobody reads.

How it is made

There are three ways to fill the two cups, and they behave very differently.

Ask two different models. Cheap, and lopsided. If one model is much better, the labeller is really labelling "which model", not "which answer".

Ask one model twice. Sample the same model with some randomness, twice. Both answers come from the model you are training, which is exactly what you want. This is called on-policy data.

Take answers from somewhere else. Public datasets, another company's model. Easy to obtain, and often about a model very unlike yours.

Who does the labelling

Two options, both flawed.

Humans. Slow and expensive. They disagree with each other constantly. Two careful people agree about two thirds of the time on open-ended answers. That is not carelessness. Many pairs genuinely have no better answer.

Another model. Fast and cheap. Also biased in ways you can measure and cannot remove. It prefers longer answers, its own style, and its own outputs.

The trap that catches everyone

Your labellers pick the longer answer far more often than they should.

They are not being lazy. Longer answers usually are more thorough. But the model you train from that data cannot tell "thorough" from "long", so it learns "long".

There is a fix, and it hurts. Keep only pairs where the two answers have similar length. It works, and it throws away most of your data.

What is honestly hard here

Averaging over several labellers feels safe. It is not always safe.

If two of your three labellers share the same bias, the majority vote makes that bias stronger, not weaker. The developer section measures this happening.

Averaging removes random mistakes. It cannot remove a bias that everybody shares.

Where you have already seen this

  • A supermarket taste test with two paper cups.
  • Two photos side by side, choosing which to keep.
  • Being asked "before or after?" while someone adjusts something.

Remember this

  • Preference data is pairs plus one bit: which one won.
  • Answers sampled from the model you are training work better than borrowed ones.
  • Labellers prefer long answers, and averaging several labellers does not fix a shared bias.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install torch

Runs on a CPU in under a second.

Simulating an annotation panel

Three annotators judge the same 4,000 pairs. One is careful, one is hurried and noisy, one skim-reads and reacts to length. There is a ground truth, which real projects never have.

annotators.py
import torch

torch.manual_seed(7)
N = 4000

# Each candidate answer has a quality and a length. Quality is what we want to
# capture. Length is a surface feature three annotators react to differently.
qual_a, qual_b = torch.randn(N), torch.randn(N)
len_a, len_b = torch.randn(N), torch.randn(N)

ANNOTATORS = {                     # (weight on quality, weight on length)
    "careful":   (2.0, 0.0),
    "hurried":   (0.6, 0.0),       # noisier, but unbiased
    "skim-reads": (1.0, 1.5),      # prefers the longer answer
}


def label(wq, wl):
    margin = wq * (qual_a - qual_b) + wl * (len_a - len_b)
    return (torch.rand(N) < torch.sigmoid(margin)).long()   # 1 = A preferred


votes = {name: label(*w) for name, w in ANNOTATORS.items()}
truth = (qual_a > qual_b).long()


def agreement(x, y):
    return (x == y).float().mean().item()


def kappa(x, y):
    po = agreement(x, y)
    px = x.float().mean().item()
    py = y.float().mean().item()
    pe = px * py + (1 - px) * (1 - py)
    return (po - pe) / (1 - pe)


names = list(votes)
print("pairwise agreement between annotators (and Cohen's kappa)")
for i in range(len(names)):
    for j in range(i + 1, len(names)):
        a, b = votes[names[i]], votes[names[j]]
        print(f"  {names[i]:<11} vs {names[j]:<11} agree {agreement(a, b):.1%}  "
              f"kappa {kappa(a, b):.3f}")

print("\nagreement with the ground truth 'higher quality wins'")
for n in names:
    print(f"  {n:<11} {agreement(votes[n], truth):.1%}")

majority = (sum(votes.values()) >= 2).long()
print(f"  {'majority-of-3':<11} {agreement(majority, truth):.1%}")

print("\nhow often does each label set pick the LONGER answer?")
longer_a = (len_a > len_b).long()
for n in names:
    print(f"  {n:<11} {agreement(votes[n], longer_a):.1%}")
print(f"  {'majority-of-3':<11} {agreement(majority, longer_a):.1%}")
print(f"  {'ground truth':<11} {agreement(truth, longer_a):.1%}")

# Filtering to pairs where the two answers are close in length removes the
# confound at the cost of throwing data away.
close = (len_a - len_b).abs() < 0.3
print(f"\nkeeping only the {close.float().mean():.0%} of pairs with similar lengths:")
print(f"  skim-reads agreement with truth  {agreement(votes['skim-reads'][close], truth[close]):.1%}"
      f"  (was {agreement(votes['skim-reads'], truth):.1%})")
print(f"  skim-reads picks the longer answer {agreement(votes['skim-reads'][close], longer_a[close]):.1%}"
      f"  (was {agreement(votes['skim-reads'], longer_a):.1%})")
Output
pairwise agreement between annotators (and Cohen's kappa)
  careful     vs hurried     agree 64.7%  kappa 0.295
  careful     vs skim-reads  agree 64.0%  kappa 0.279
  hurried     vs skim-reads  agree 57.8%  kappa 0.157

agreement with the ground truth 'higher quality wins'
  careful     83.2%
  hurried     66.5%
  skim-reads  65.8%
  majority-of-3 79.5%

how often does each label set pick the LONGER answer?
  careful     52.0%
  hurried     50.5%
  skim-reads  74.4%
  majority-of-3 60.9%
  ground truth 50.9%

keeping only the 16% of pairs with similar lengths:
  skim-reads agreement with truth  70.2%  (was 65.8%)
  skim-reads picks the longer answer 53.1%  (was 74.4%)

Written against PyTorch 2.5.1, CPU, manual_seed(7) — reproducible on this build.

Four results, and two of them are uncomfortable

Agreement between annotators is 58–65%, with kappa between 0.16 and 0.30. Those are not broken annotators. They are in the range reported for real open-ended preference annotation. Kappa below 0.4 is conventionally called "fair", and this is the material every alignment method is built on.

The majority vote was WORSE than the best single annotator. 79.5% against the careful annotator's 83.2%. Two weak annotators outvoted a strong one. Majority voting assumes your labellers are roughly equally competent, and when they are not, it actively destroys signal. Identify and weight your good annotators; do not average them away.

The majority vote imported the length bias. The truth picks the longer answer 50.9% of the time. The skim-reader picks it 74.4% of the time. The majority-of-three picks it 60.9% — more than half the bias survived. Averaging removes independent noise. It cannot remove a bias that one annotator holds strongly.

Length-balancing worked, and cost 84% of the data. Restricting to similar-length pairs cut the skim-reader's length preference from 74.4% to 53.1%, and raised its agreement with truth from 65.8% to 70.2%. That is the trade in its most honest form: a much cleaner signal from a much smaller dataset.

Generating the pairs

The standard on-policy recipe, in outline:

python
# for each prompt, sample k completions from the CURRENT policy
completions = model.generate(prompt_ids, num_return_sequences=4,
                             do_sample=True, temperature=1.0, top_p=0.95)
# score them (reward model, LLM judge, or a verifier), then pair best with worst

No output block — the text depends on the model, the sampling seed and the library version, and inventing a sample would teach you to expect something that will not happen.

Three details that decide whether this works:

  • Temperature around 1.0. Too low and all four completions are near-identical, so the pair carries no signal. Too high and you are labelling noise.
  • Best against worst, not best against second-best. A wider margin gives a cleaner gradient.
  • Discard ties. A pair the judge scores equally contributes nothing and adds noise.

Public datasets, and what each is actually good for

DatasetPairsJudgeWatch out for
Anthropic HH-RLHF~170khumansold models; helpful and harmless splits differ sharply
UltraFeedback~64kGPT-4 ratingsjudge bias; strong length correlation
Nectar~183kGPT-4 rankings7-way rankings, needs conversion to pairs
SHP~385kReddit vote countspopularity is not quality
Tülu 3 preference mixcuratedmixeddocumented and decontaminated

These are all off-policy for your model. They are the right way to start and the wrong way to finish. Iterative on-policy collection is what closes the last of the gap.

Common mistakes

Never measuring annotator agreement. If you do not know your kappa, you do not know your ceiling. Have several annotators label the same 200 pairs before scaling up.

Averaging annotators without checking them. As measured above, this can be worse than picking your best annotator.

Not stratifying by prompt type. A dataset that is 80% creative writing will produce a reward model that is confident about prose and useless about code.

Not decontaminating. If prompts overlap your evaluation sets, your win rates are fiction. See benchmark contamination.

Using an LLM judge without calibrating it. Take 200 pairs, have humans label them, and measure your judge against those humans. Report that number alongside every result the judge produced.

Ignoring position bias in LLM judges. Judges systematically prefer the first option shown. Present each pair in both orders and keep only the pairs where the judge agrees with itself.

Try it yourself

Add a fourth annotator with weights (1.0, 1.5) — a second skim-reader. Recompute the majority vote. Now two of four share the length bias, and the majority's length preference should rise further. That is the failure mode in one edit.

What to learn next

Researcher — Mathematics and papers.

Where the noise ceiling comes from

Under Bradley–Terry, an annotator with quality weight $w$ picks correctly with probability $\sigma(w\,\Delta q)$, where $\Delta q$ is the true quality gap. Integrating over the distribution of gaps gives an expected accuracy strictly below 1 for any finite $w$. Even a perfect annotator is stochastic on near-ties, because near-ties genuinely have no better answer.

Reported human agreement on open-ended preference tasks:

  • InstructGPT (Ouyang et al., 2022): ~73% agreement between held-out labellers on their comparison task.
  • Anthropic HH (Bai et al., 2022): comparable, with the helpfulness split noticeably noisier than harmlessness.
  • MT-Bench / Chatbot Arena (Zheng et al., 2023): human–human agreement around 81% on their curated pairs, with GPT-4 reaching similar agreement with humans as humans do with each other.

This bounds every downstream artefact. A reward model reported at 90% on the same distribution is fitting annotator identity, not quality.

Majority voting is not variance reduction when biases are shared

Write annotator $i$'s decision as $\mathrm{sign}(w_i \Delta q + b_i \Delta \ell + \varepsilon_i)$, with $\Delta \ell$ the length gap. Majority voting averages the independent $\varepsilon_i$ — variance falls as $1/n$ — but it does not touch the systematic $b_i \Delta \ell$ term. If $\mathbb{E}[b_i] \neq 0$ across your panel, the aggregate keeps the bias while looking more reliable.

Two better estimators, both standard in the crowdsourcing literature:

  • Dawid–Skene (1979): an EM procedure jointly estimating per-annotator confusion matrices and latent true labels. It downweights weak annotators instead of averaging them.
  • Bradley–Terry with annotator effects: fit a per-annotator scale $w_i$ and bias term, and report the debiased latent reward.

Both need each item labelled by several annotators, which is what makes a small overlap sample non-optional.

LLM judges

Zheng et al., 2023 (Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena) documented three biases that all replicate:

  1. Position bias. GPT-4 as judge changes its answer when the two responses are swapped, on a substantial fraction of pairs. Mitigation: evaluate both orders, keep only self-consistent judgements.
  2. Verbosity bias. Longer answers win more, controlling for content.
  3. Self-enhancement bias. A judge prefers text from its own model family.

Dubois et al., 2024 (Length-Controlled AlpacaEval) fit a generalised linear model with a length term and report the length-controlled win rate. It raised the correlation with Chatbot Arena from 0.94 to 0.98 and substantially reordered the leaderboard. If you report win rates, report the length-controlled version too.

On-policy versus off-policy

Tajwar et al., 2024 (Preference Fine-Tuning of LLMs Should Leverage Suboptimal, On-Policy Data) is the sharpest empirical statement of the case. Their finding: methods using on-policy sampling and a negative gradient (pushing down bad responses) outperform purely offline methods, and the advantage is largest when the target distribution is far from the reference. Contrastive offline methods on off-policy data underperform in exactly that regime.

The consequence for data collection: pairs generated by a much stronger model are further from your policy, not closer to what you need. They teach a direction your policy cannot take in one step.

Iterative / online DPO is the practical answer, and both Llama 3 and Tülu 3 use a version of it: sample from the current policy, label, train, resample. Each round is on-policy for the policy that generated it.

Length balancing, and what it costs

Filtering to $|\Delta\ell| < \tau$ is unbiased with respect to length by construction, and it changes the prompt distribution too — short-answer prompts survive filtering more often. Alternatives that keep more data:

  • Reward regression debiasing. Fit reward on quality features and length jointly, then zero the length coefficient at inference (Chen et al., 2024, ODIN uses a two-head reward model with an explicit length head, discarded at RL time).
  • Length-normalised objectives. SimPO's average-log-probability reward, or loss_type="sigmoid_norm" in TRL.
  • Explicit length penalty in the reward.

None fully removes the effect, because length genuinely correlates with quality in the data-generating process. The goal is to remove the spurious portion, and that requires a causal assumption you cannot verify from observational preference data alone. See confounding for why this is hard in general.

Constitutional AI and RLAIF

Bai et al., 2022 (Constitutional AI) replaced human harmlessness labels with model-generated critiques and revisions guided by a written list of principles, then trained a preference model on AI-generated comparisons. Lee et al., 2024 (RLAIF vs RLHF) found AI feedback matching human feedback on summarisation and dialogue, with the caveat that the AI labeller inherits its own training's biases.

The practical position most teams settle on: AI feedback for scale and coverage, a human-labelled calibration set of a few hundred pairs to measure the AI labeller against, and human labels for anything safety-critical.

Papers

What to learn next