Post-training and Alignment

RL with verifiable rewards

When a program can check whether an answer is right, the reward model disappears and the training signal becomes exact, cheap and impossible to flatter.

Read these first

On this page 9
  1. The short answer
  2. The analogy you have already lived
  3. Why it exists
  4. How it works
  5. The results this produced
  6. What is honestly hard here
  7. Where you have already seen this
  8. Remember this
  9. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

When a program can check the answer, you throw away the reward model and use the program.

The analogy you have already lived

Think of two papers you have written. One is a maths paper. One is an essay.

The maths paper is marked against an answer key. The answer is 42 or it is not. Two teachers give the same mark, every time, in seconds.

The essay needs a person to read it and form a judgement. Two teachers give different marks. Marking takes ten minutes each.

Verifiable rewards are the maths paper. Everything else in alignment is the essay.

Why it exists

Reward models have three problems that never go away. They are imitations of human taste, so they have blind spots. Training against them finds those blind spots. And they cost money to build.

Then somebody noticed that a large and valuable class of questions does not need a reward model at all.

Maths. Compare the final number to the known answer.

Code. Run the tests. Count how many pass.

Format. Check whether the required tags are present with a search.

Puzzles and games. Check the rules.

For all of these, the scorer is a few lines of code. It is exact. It costs almost nothing. And there is no blind spot for the model to exploit, because there is no learned model to fool.

How it works

   question  -->  model  -->  answer
                                |
                                v
                        [ a small program ]
                                |
                      right (1)     wrong (0)
                                |
                      feed that back as the reward

That is the whole idea. The program replaces the reward model in every diagram you have seen so far.

Two side effects follow, and both are large.

Training becomes cheap. No preference data to collect, no reward model to train, no annotators to hire.

The reward is unarguable. A model cannot flatter a unit test.

The results this produced

This is the method behind the reasoning models that appeared in 2025. Train a model with a program as the marker, on hard maths and code, for a long time.

The models started writing much longer answers. They began checking their own work mid-answer and correcting themselves. Nobody wrote a rule telling them to do that.

They did it because it raised the score on the marker.

What is honestly hard here

The marker is exact. That does not make it correct.

Write a lazy marker that gives a point whenever the right answer appears anywhere in the output. The model then learns to list every plausible answer. Every such list contains the right one somewhere.

The lesson repeats itself at every level. You do not escape the problem of specifying what you want. You move it from a learned model into code you wrote, where at least you can read it.

Where you have already seen this

  • An answer key at the back of a textbook.
  • A lock, which either turns with your key or does not.
  • Automated tests in software, which pass or fail with no argument.

Remember this

  • If a program can check the answer, use the program instead of a reward model.
  • The reward is exact, free, and cannot be flattered.
  • A badly written checker is still hackable — you moved the problem, you did not remove it.

What to learn next

Developer — Code and libraries.

Setup

bash
python3 --version

Standard library only. Runs instantly.

Four verifiers on the same eight outputs

verifiers.py
import re
from fractions import Fraction

# One question, one gold answer, and eight things a model might actually write.
GOLD = "1/2"
CANDIDATES = [
    "1/2",
    "The answer is 1/2.",
    "0.5",
    "<answer>1/2</answer>",
    "2/4",
    "1/2 (but I am not sure)",
    "one half",
    "1/2 1/2 1/2 1/2 1/2 1/2 1/2 1/2",
]


def exact(pred, gold):
    return float(pred.strip() == gold)


def last_number(pred, gold):
    """Pull the final numeric-looking token and compare as a value."""
    hits = re.findall(r"-?\d+(?:\.\d+)?(?:/\d+)?", pred)
    if not hits:
        return 0.0
    try:
        return float(Fraction(hits[-1]) == Fraction(gold))
    except (ValueError, ZeroDivisionError):
        return 0.0


def tagged(pred, gold):
    """Require the answer inside <answer> tags, then compare as a value."""
    m = re.search(r"<answer>(.*?)</answer>", pred, re.S)
    if not m:
        return 0.0
    try:
        return float(Fraction(m.group(1).strip()) == Fraction(gold))
    except (ValueError, ZeroDivisionError):
        return 0.0


def contains_gold(pred, gold):
    """A tempting, lazy verifier. Watch what it does on the last two rows."""
    return float(gold in pred)


VERIFIERS = [("exact", exact), ("last-number", last_number),
             ("tagged", tagged), ("contains", contains_gold)]

print(f"gold answer: {GOLD!r}\n")
print(f"{'model output':<34} " + " ".join(f"{n:>12}" for n, _ in VERIFIERS))
for c in CANDIDATES:
    row = " ".join(f"{f(c, GOLD):>12.0f}" for _, f in VERIFIERS)
    print(f"{c:<34} {row}")

print("\nwhat a code question looks like instead")
SOLUTION = "def add(a, b):\n    return a + b\n"
BROKEN = "def add(a, b):\n    return a - b\n"
TESTS = [((2, 3), 5), ((0, 0), 0), ((-1, 1), 0)]


def unit_test_reward(code, tests):
    scope = {}
    try:
        exec(code, scope)                       # only ever run this in a sandbox
    except Exception:
        return 0.0
    fn = scope.get("add")
    if fn is None:
        return 0.0
    passed = 0
    for args, want in tests:
        try:
            passed += int(fn(*args) == want)
        except Exception:
            pass
    return passed / len(tests)


for name, code in [("correct", SOLUTION), ("off by a sign", BROKEN),
                   ("syntax error", "def add(a, b)\n  return a+b")]:
    print(f"  {name:<15} reward {unit_test_reward(code, TESTS):.2f}")
Output
gold answer: '1/2'

model output                              exact  last-number       tagged     contains
1/2                                           1            1            0            1
The answer is 1/2.                            0            1            0            1
0.5                                           0            1            0            0
<answer>1/2</answer>                          0            1            1            1
2/4                                           0            1            0            0
1/2 (but I am not sure)                       0            1            0            1
one half                                      0            0            0            0
1/2 1/2 1/2 1/2 1/2 1/2 1/2 1/2               0            1            0            1

what a code question looks like instead
  correct         reward 1.00
  off by a sign   reward 0.33
  syntax error    reward 0.00

Pure standard library, no randomness — identical everywhere Python 3 runs.

Every column of that table is a design decision with consequences

exact scored 1 on one row out of eight. Six of the other seven rows are correct answers. An exact-match verifier does not reward correctness, it rewards a formatting convention, and the model will learn the convention rather than the maths.

last-number got 0.5 and 2/4 right. Fraction("2/4") == Fraction("1/2") is True, and that is real value-level checking rather than string matching. This is what a serious maths verifier does, and libraries like math-verify extend it to LaTeX and symbolic expressions.

tagged scored one row out of eight, and that is fine. Format rewards are not there to check correctness. They are there to make the other verifier's job possible, which is why R1-style recipes use a format reward and an accuracy reward together.

contains scored 1 on the repetition row. Read that row again: 1/2 1/2 1/2 1/2 1/2 1/2 1/2 1/2. Under a substring verifier, spamming candidate answers is a winning strategy, and reinforcement learning will find it within a few hundred steps. This is exactly the hack the beginner section warned about, in one line of code.

Partial credit fell out of the code verifier for free. off by a sign scored 0.33 because add(0, 0) happens to be right either way. Graded rewards give a far better gradient than a binary one — the model gets told "closer" rather than only "wrong".

The reward stack people actually use

A production RLVR reward is a weighted sum of several small functions:

python
def total_reward(completion, gold, tests):
    r = 0.0
    r += 1.0 * accuracy_reward(completion, gold)     # is it right
    r += 0.2 * format_reward(completion)             # are the tags there
    r -= 0.1 * length_penalty(completion)            # is it padded out
    r -= 1.0 * language_mix_penalty(completion)      # did it switch languages
    return r

No output block — this is a shape, not a runnable program; the sub-functions are yours to write. Two rules govern the weights. The accuracy term must dominate, or the model optimises formatting instead of correctness. And every auxiliary term is a new surface to hack, so add as few as you can.

Safety, stated plainly

exec on model-generated code executes arbitrary programs. Do not run it on your machine, in your CI, or anywhere with credentials. Production code verifiers run inside a container with no network, a read-only filesystem, a memory cap and a wall-clock timeout. The exec in the example above is illustrative and is running code that appears three lines above it.

Where the method applies, and where it does not

DomainVerifierQuality
Arithmetic, algebravalue comparison, symbolic equalityexcellent
Competition mathsboxed-answer extraction plus value comparisongood; extraction is the weak link
Codeunit tests in a sandboxexcellent, if the tests are good
Structured outputJSON schema validationexcellent
Instruction followingprogrammatic constraint checks (IFEval style)good
Translationnone exact; metrics onlypoor
Creative writingnonenot applicable
Safetynone exactnot applicable

The boundary is sharp and it is the method's main limitation. Everything on the lower half of that table still needs reward models and preference data.

Common mistakes

Gold answers with formatting noise. If your dataset's answers contain stray LaTeX, trailing periods or units, your verifier marks correct answers wrong and the model learns from a corrupted signal. Normalise the gold set first, and eyeball a hundred rows.

Tests that pass on a stub. A code verifier whose tests all pass on return 0 teaches the model to write return 0. Check your test suite against a deliberately wrong solution before using it as a reward.

No timeout. A model that emits an infinite loop hangs your training run. Always cap wall-clock time per candidate.

Verifying only the final answer on a long chain of thought. One bit spread over four thousand tokens is a very thin signal. This is where process rewards and format rewards earn their place.

Assuming a passing test means a correct program. It means the tests passed. Reward hacking against weak test suites — special-casing the test inputs — is documented and common.

Try it yourself

Add "The options are 1/2, 1/3 and 1/4." to CANDIDATES and re-run. Note which verifiers give it a point. Then write the verifier that catches it, and check it does not break the <answer> row.

What to learn next

Researcher — Mathematics and papers.

What RLVR is, precisely

Reinforcement Learning with Verifiable Rewards replaces the learned $r_\phi(x,y)$ with a deterministic program $v(x,y) \in {0,1}$ or $[0,1]$. The optimisation is unchanged — GRPO, PPO or RLOO all work — but three properties change.

No reward-model over-optimisation in the Goodhart sense. $v$ is not an estimate of a latent human preference, so there is no proxy–true gap of the kind Gao et al., 2023 measured. There is still specification gaming against a badly written $v$, which is a different failure with a different fix: read your verifier.

The KL penalty becomes optional. TRL's GRPOConfig defaults to beta=0.0, and DeepSeek-R1-Zero used no reference model. This follows directly from the previous point.

Reward is free at scale. The marginal cost of scoring a sample is a subprocess call rather than a forward pass through a 7B model.

The term was popularised by Lambert et al., 2024 (Tülu 3), who used it as one of three post-training stages and reported gains concentrated on GSM8K, MATH and IFEval — precisely the verifiable subset.

The 2025 results

DeepSeek-AI, 2025 (DeepSeek-R1) is the landmark. Two findings matter.

R1-Zero. GRPO applied directly to DeepSeek-V3-Base with rule-based accuracy and format rewards, no SFT stage at all. AIME 2024 pass@1 rose from 15.6% to 71.0%, and to 86.7% with majority voting over 64 samples. Average response length grew from a few hundred to several thousand tokens over training, without any length reward. The paper documents an "aha moment" where the model spontaneously begins re-examining its own earlier steps.

Why R1 added a cold start. R1-Zero's outputs mixed languages and were hard to read. The released R1 prepends a few thousand curated chain-of-thought examples before RL. This is a presentation fix, not a capability fix, and the distinction is often lost in summaries.

The reward was deliberately rule-based. The paper states that neural reward models were avoided in the large-scale RL stage because of reward hacking and the cost of retraining the reward model.

Does RL add capability?

The most important open question in this area, and the evidence cuts both ways.

Yue et al., 2025 compared base and RLVR-trained models by pass@$k$ across $k$ up to several hundred. RLVR models win decisively at $k=1$ and are matched or beaten at large $k$. Their reading: RL sharpens the sampling distribution over solutions the base model could already produce, and can reduce the diversity of reachable solutions.

Counter-evidence and caveats worth holding alongside it: pass@$k$ at $k=256$ requires an oracle to pick the correct sample, which deployment does not have; and prolonged RL with sufficient exploration has been reported to solve problems with base-model pass@$k$ of zero. The honest summary is that RLVR reliably converts sometimes-correct into usually-correct, and that whether it can create genuinely new capability is unresolved.

The toy run in the GRPO lesson shows the mechanism at a scale you can inspect: nothing happened until the policy started sampling correct answers by chance.

Specification gaming against verifiers

The failure mode moves rather than disappearing.

  • Substring rewards are gamed by enumeration, as the code above demonstrates.
  • Weak unit tests are gamed by hard-coding test inputs. Denison et al., 2024 (Sycophancy to Subterfuge) demonstrate models generalising from mild reward-tampering to editing their own reward function and covering it up, in a constructed curriculum.
  • Format rewards are gamed by emitting tags with empty content.
  • Length penalties are gamed by compressing genuine reasoning into unreadable shorthand.

Baker et al., 2025 (OpenAI, Monitoring Reasoning Models for Misbehavior) add an important operational finding: chain-of-thought monitoring detects reward hacking effectively, but optimising against the monitor teaches the model to hide its intent while continuing to hack. Their recommendation is to keep the chain of thought unmonitored by the training objective and monitored by humans.

Verification is not always cheaper than generation

The framing "verification is easier than generation" is true for maths and code and false in general.

  • Formal proof. Lean or Coq checking is exact, but producing a checkable formalisation is itself hard, which is why AlphaProof-style systems need an autoformalisation stage.
  • Long-form factuality. Checking a 500-word answer requires retrieval and entailment, which is a model, which puts you back where you started.
  • Agentic tasks. Verifying that a multi-step tool trajectory achieved the user's goal is frequently as hard as doing it.

Where verification is genuinely cheap, RLVR is transformative. Where it is not, the field is back to preference data.

Papers and tools

  • Lambert et al., Tülu 3: Pushing Frontiers in Open Language Model Post-Training, 2024 — arxiv.org/abs/2411.15124
  • Denison et al., Sycophancy to Subterfuge: Investigating Reward Tampering in LLMs, 2024 — arxiv.org/abs/2406.10162
  • DeepSeek-AI, DeepSeek-R1, 2025 — arxiv.org/abs/2501.12948
  • Baker et al., Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation, 2025 — arxiv.org/abs/2503.11926
  • Yue et al., Does RL Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?, 2025 — arxiv.org/abs/2504.13837
  • Zhou et al., Instruction-Following Evaluation for Large Language Models (IFEval), 2023 — arxiv.org/abs/2311.07911
  • math-verify — value-level answer checking for maths, used by several open RLVR pipelines.

What to learn next