Post-training and Alignment

GRPO

GRPO trains a model by sampling several answers to the same question and rewarding the ones that beat the group's own average, with no value network at all.

On this page 8
  1. The short answer
  2. The analogy you have already lived
  3. Why it exists
  4. How it works
  5. The property that surprises people
  6. Where you have already seen this
  7. Remember this
  8. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

GRPO asks the model the same question several times and scores every answer. It rewards the ones that beat the group's own average.

The analogy you have already lived

Think of relative grading in a classroom. The teacher does not decide in advance that 80 marks is a distinction. She looks at what the whole class scored, then grades against that average.

If the paper was brutal and the class average was 40, your 55 is excellent. If the paper was easy and the average was 85, your 55 is poor.

GRPO grades the model the same way. Ask one question eight times, score the eight answers, and compare each one to the average of those eight.

Why it exists

The older method, PPO, needed a second network. Its whole job was to guess how good a situation was likely to be. That guess is the baseline you compare an answer against.

That second network — the value network — is roughly the same size as the model you are training. It doubles the memory. It has its own training problems. It is another thing to get wrong.

Somebody asked a good question. If you are going to sample eight answers anyway, why not use their average as the baseline?

No second network. No extra memory. The baseline comes free, from samples you already had.

How it works

   question: "What is 17 x 24?"
        |
        +--> answer 1  --> checked --> right  (1)
        +--> answer 2  --> checked --> wrong  (0)
        +--> answer 3  --> checked --> right  (1)
        +--> answer 4  --> checked --> wrong  (0)
        +--> answer 5  --> checked --> wrong  (0)
        +--> answer 6  --> checked --> right  (1)
        +--> answer 7  --> checked --> wrong  (0)
        +--> answer 8  --> checked --> wrong  (0)
                                        |
                       group average is 3 out of 8
                                        |
        answers that were right: above average -> make more likely
        answers that were wrong: below average -> make less likely

Notice who is doing the scoring. For maths and code, the scorer is a program, not a model. Run the code. Check the number. That answer cannot be argued with, and it cannot be gamed the way a learned scorer can.

The property that surprises people

Look at what happens when the whole group is right, or the whole group is wrong.

All eight right? Every answer equals the average. Nothing is above or below. No learning happens from that question.

All eight wrong? The same thing. No signal at all.

Only questions the model gets partly right teach it anything. Questions that are far too easy or far too hard are wasted compute.

That is not a flaw to be fixed. It is a description of what reinforcement learning is. The model can only be pushed toward things it already sometimes does.

Where you have already seen this

  • Relative grading against a class average.
  • Several people trying the same shot, and comparing each attempt to the group.
  • A cooking contest where the judge tastes all the entries before deciding.

Remember this

  • Sample several answers to one question, and use their average as the baseline.
  • That removes the second network PPO needed, and halves the memory.
  • A question where every sample agrees teaches nothing.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install torch

Runs on a CPU in about forty seconds.

GRPO in one screen

The task: given two digits, output (a + b) mod 10. The reward is a Python comparison, not a learned model. The model is never shown a correct answer during the GRPO phase — only told whether its own sample was right.

grpo_from_scratch.py
import torch
import torch.nn as nn
import torch.nn.functional as F

# Task: given two digits, answer (a + b) mod 10. The reward is a program, not a
# model: 1.0 for the right digit, 0.0 otherwise. The model is never shown a
# correct answer during GRPO - only told whether its own sample was right.
V, SEP, GROUP, EPS_LOW, EPS_HIGH, MU = 13, 10, 16, 0.2, 0.28, 2


class Policy(nn.Module):
    def __init__(self):
        super().__init__()
        self.emb = nn.Embedding(V, 48)
        self.pos = nn.Embedding(4, 48)
        layer = nn.TransformerEncoderLayer(48, 4, 96, batch_first=True, dropout=0.0)
        self.blocks = nn.TransformerEncoder(layer, 2)
        self.head = nn.Linear(48, V)

    def forward(self, x):
        h = self.emb(x) + self.pos(torch.arange(x.shape[1]))
        return self.head(self.blocks(h))[:, -1]      # logits for the next token


def prompts(n, gen):
    ab = torch.randint(0, 10, (n, 2), generator=gen)
    return torch.cat([ab, torch.full((n, 1), SEP)], 1), (ab.sum(1) % 10)


def accuracy(model, gen):
    x, y = prompts(1000, gen)
    with torch.no_grad():
        return (model(x).argmax(-1) == y).float().mean().item()


def sft(model, steps, gen, lr=3e-3):
    """A short supervised warm-up on a handful of solved examples."""
    opt = torch.optim.AdamW(model.parameters(), lr=lr)
    x, y = prompts(160, gen)                       # only 160 worked examples
    for _ in range(steps):
        loss = F.cross_entropy(model(x), y)
        opt.zero_grad()
        loss.backward()
        opt.step()


def grpo(model, steps, gen, evalgen, tag):
    opt = torch.optim.AdamW(model.parameters(), lr=5e-4)
    print(f"\n--- GRPO {tag} ---")
    print(f"{'step':>5} {'sample reward':>14} {'greedy acc':>11} {'clipped':>8} {'dead groups':>12}")
    for step in range(1, 401):
        x, y = prompts(16, gen)
        xg, yg = x.repeat_interleave(GROUP, 0), y.repeat_interleave(GROUP, 0)

        with torch.no_grad():                                  # roll out the old policy
            logits_old = model(xg)
            sampled = torch.multinomial(logits_old.softmax(-1), 1, generator=gen).squeeze(1)
            logp_old = logits_old.log_softmax(-1).gather(1, sampled[:, None]).squeeze(1)

        reward = (sampled == yg).float()                       # the verifier

        r = reward.view(-1, GROUP)
        # the GRPO advantage: no value network, only the group's own mean and spread
        adv = ((r - r.mean(1, keepdim=True)) / (r.std(1, keepdim=True) + 1e-4)).reshape(-1)
        dead = (r.std(1) == 0).float().mean().item()           # groups with no signal at all

        clipped_frac = 0.0
        for _ in range(MU):
            logp = model(xg).log_softmax(-1).gather(1, sampled[:, None]).squeeze(1)
            ratio = (logp - logp_old).exp()
            unclipped, clipped = ratio * adv, ratio.clamp(1 - EPS_LOW, 1 + EPS_HIGH) * adv
            loss = -torch.min(unclipped, clipped).mean()
            clipped_frac = (unclipped != clipped).float().mean().item()
            opt.zero_grad()
            loss.backward()
            opt.step()

        if step % 100 == 0 or step == 1:
            print(f"{step:>5} {reward.mean().item():>14.3f} {accuracy(model, evalgen):>11.3f} "
                  f"{clipped_frac:>8.2f} {dead:>12.2f}")


print(f"group size {GROUP}, clip range [{1 - EPS_LOW}, {1 + EPS_HIGH}], {MU} inner updates")
print("random guessing scores 0.100\n")

torch.manual_seed(0)
cold = Policy()
print(f"cold model, greedy accuracy before GRPO: "
      f"{accuracy(cold, torch.Generator().manual_seed(99)):.3f}")
grpo(cold, 400, torch.Generator().manual_seed(1),
     torch.Generator().manual_seed(99), "from a RANDOM model")

torch.manual_seed(0)
warm = Policy()
sft(warm, 25, torch.Generator().manual_seed(5))
print(f"\nwarm model, greedy accuracy after a short SFT: "
      f"{accuracy(warm, torch.Generator().manual_seed(99)):.3f}")
grpo(warm, 400, torch.Generator().manual_seed(1),
     torch.Generator().manual_seed(99), "from an SFT-warmed model")
Output
group size 16, clip range [0.8, 1.28], 2 inner updates
random guessing scores 0.100

cold model, greedy accuracy before GRPO: 0.091

--- GRPO from a RANDOM model ---
 step  sample reward  greedy acc  clipped  dead groups
    1          0.117       0.106     0.00         0.31
  100          0.086       0.235     0.00         0.31
  200          0.574       0.712     0.06         0.50
  300          0.664       0.847     0.01         0.69
  400          0.812       0.923     0.00         1.00

warm model, greedy accuracy after a short SFT: 0.469

--- GRPO from an SFT-warmed model ---
 step  sample reward  greedy acc  clipped  dead groups
    1          0.375       0.451     0.04         0.19
  100          0.359       0.659     0.00         0.69
  200          0.730       0.775     0.00         0.75
  300          0.727       0.843     0.01         0.75
  400          0.859       0.825     0.02         0.88

Written against PyTorch 2.5.1, CPU, all generators seeded — byte-identical across repeated runs on this build. Different hardware may shift the trajectory slightly; the pattern is what reproduces.

Every column in that table is worth reading

The cold model went from 9.1% to 92.3% with no labelled answer ever shown to it. All it received was a sampled == yg comparison. This is reinforcement learning from a verifiable reward in miniature, and it is the shape of what produced the 2025 reasoning models.

Nothing happened for the first hundred steps. Accuracy crawled 0.106 → 0.235. Then it took off. That flat-then-sharp shape is characteristic of RL: the model must stumble onto the behaviour before the reward can reinforce it.

The dead groups column rises to 1.00 as the model improves. By step 400, every group of 16 samples is unanimous, so every advantage is zero and every gradient is zero. The algorithm has run out of signal precisely because it succeeded. This is not a bug in the toy; it is the motivation for DAPO's dynamic sampling, described below.

The warm start converged faster early and no better late. 0.659 against 0.235 at step 100; 0.825 against 0.923 at step 400. SFT bought early progress and no ceiling. Real pipelines use a warm start for a different reason: on real reasoning tasks the answer space is astronomically large, and a random model's success rate is not 10% but effectively zero — so there is nothing to reinforce.

clipped stayed near zero. With only two inner updates the policy barely moves per batch, so the ratio stays inside the clip range. Raise MU and this column rises, which is exactly what the clip is there to bound.

The real thing

TRL (v1.12.0):

python
from datasets import load_dataset
from trl import GRPOTrainer
from trl.rewards import accuracy_reward

dataset = load_dataset("trl-lib/DeepMath-103K", split="train")

trainer = GRPOTrainer(
    model="Qwen/Qwen2.5-0.5B-Instruct",
    reward_funcs=accuracy_reward,
    train_dataset=dataset,
)
trainer.train()

No output block — this needs a GPU and downloads gigabytes, and its logs are hardware-dependent.

reward_funcs accepts your own Python functions. That is the point of the method: the reward is code you wrote.

python
def format_reward(completions, **kwargs):
    """1.0 if the answer is wrapped in the tags we asked for."""
    return [1.0 if "<answer>" in c and "</answer>" in c else 0.0 for c in completions]

Defaults in TRL v1.12.0 that differ from the original paper, and matter:

SettingDefaultWhy
beta0.0KL penalty off; no reference model is loaded
loss_type"dapo"token-level normalisation, not the paper's per-sequence
scale_rewards"group"divide by the group's standard deviation
num_generations8the group size
epsilon / epsilon_high0.2 / Noneset epsilon_high higher for DAPO's "clip-higher"
max_completion_length512

beta=0.0 is the biggest departure. With a program as the reward there is no proxy to over-optimise in the Goodhart sense, so the KL leash that RLHF needs is optional here.

Common mistakes

Too small a group. With num_generations=2 almost every group is unanimous and almost every batch is wasted. 8 to 16 is the usual range, and larger is better on hard problems.

A reward function that crashes. Your reward is arbitrary Python running on model output. Wrap it in try/except and return 0.0 on failure, or one malformed generation kills the run.

Rewarding only the final answer on long outputs. A single scalar spread over 4,000 tokens is a very weak signal. Format rewards and step-level rewards fill the gap.

Forgetting that generation dominates the cost. Each step generates group_size × batch_size completions. That is where the wall-clock time goes, which is why production GRPO uses vLLM for the rollout phase.

Ignoring the dead-group fraction. Log it. When it approaches 1.0 your run has stopped learning, whatever the reward curve says.

Length blowing up. GRPO on reasoning tasks reliably makes outputs longer. Some of that is real reasoning; some is padding. Cap it and shape the reward for overlong outputs.

Try it yourself

Set GROUP = 4 and re-run the cold model. Watch dead groups start much higher and learning stall. That single number explains most of the difference between a GRPO run that works and one that does not.

What to learn next

Researcher — Mathematics and papers.

The objective

GRPO (Shao et al., 2024, DeepSeekMath) samples a group ${o_1,\dots,o_G}$ from $\pi_{\theta_{\text{old}}}(\cdot\mid q)$ and optimises

$$ \mathcal{J}{\text{GRPO}}(\theta) = \mathbb{E}\left[\frac{1}{G}\sum{i=1}^{G}\frac{1}{|o_i|}\sum_{t=1}^{|o_i|} \min!\Big(\rho_{i,t}\hat A_{i,t},\ \mathrm{clip}(\rho_{i,t}, 1-\epsilon, 1+\epsilon)\hat A_{i,t}\Big) - \beta\, \mathbb{D}{\mathrm{KL}}!\left[\pi\theta | \pi_{\text{ref}}\right]\right] $$

with importance ratio $\rho_{i,t} = \dfrac{\pi_\theta(o_{i,t}\mid q, o_{i,<t})}{\pi_{\theta_{\text{old}}}(o_{i,t}\mid q, o_{i,<t})}$ and the group-relative advantage

$$ \hat A_{i,t} = \frac{r_i - \mathrm{mean}({r_1,\dots,r_G})}{\mathrm{std}({r_1,\dots,r_G})} $$

Three structural points.

The advantage is constant across $t$. Every token in a completion receives the same advantage. There is no per-token credit assignment, and no GAE. This is what removes the value network.

The baseline is a Monte Carlo estimate of $V(q)$. PPO learns $V$; GRPO estimates it from $G$ samples at each step. The variance of that estimate is $O(1/G)$, which is why group size is the algorithm's most important hyperparameter.

The KL term is in the loss, not the reward. GRPO uses the k3 estimator directly in the objective, unlike PPO-style RLHF which folds a per-token KL into the reward signal.

Where GRPO is biased, and the corrections

Liu et al., 2025 (Understanding R1-Zero-Like Training: A Critical Perspective, "Dr. GRPO") identify two biases in the formula above.

  1. Length bias, from the $1/|o_i|$ term. Dividing by response length means each token of a long wrong answer is penalised less than each token of a short wrong answer. The optimisation therefore prefers longer wrong answers — which explains part of the response-length growth attributed to "more reasoning".
  2. Difficulty bias, from dividing by $\mathrm{std}(r)$. Questions where the group's rewards happen to have low variance get their advantages inflated, over-weighting questions that are nearly always right or nearly always wrong.

Dr. GRPO removes both: drop the $1/|o_i|$ factor (use a constant $L$ instead) and drop the standard-deviation division. TRL exposes this as loss_type="dr_grpo" and scale_rewards=False.

DAPO (Yu et al., 2025, Decoupled Clip and Dynamic sAmpling Policy Optimization) makes four changes, all of which are now common practice:

  • Clip-higher. Decouple $\epsilon_{\text{low}}$ and $\epsilon_{\text{high}}$, raising the upper bound. The symmetric clip caps how much probability a low-probability token can gain, which drives entropy collapse; a larger upper bound preserves exploration.
  • Dynamic sampling. Oversample and discard groups where accuracy is exactly 0 or exactly 1. This directly targets the dead groups column in the code above — those groups contribute nothing but still consume batch capacity and inflate gradient variance.
  • Token-level policy-gradient loss. Average over all tokens in the batch rather than averaging per-sequence means, so long sequences carry proportionate weight.
  • Overlong reward shaping. A soft length penalty starting at one threshold and strong enough past a second threshold to cancel a correct answer's reward.

GSPO (Zheng et al., 2025, Qwen) argues the importance ratio should be defined at the sequence level, since the reward is a sequence-level quantity, and reports better stability for mixture-of-experts models where token-level ratios are especially noisy.

What RL on verifiable rewards actually adds

This is contested and worth stating carefully.

DeepSeek-R1 (DeepSeek-AI, 2025) reported R1-Zero — GRPO applied directly to a base model with rule-based rewards, no SFT — reaching 71.0% on AIME 2024 (86.7% with majority voting), up from 15.6%. It also reported the emergence of longer chains of thought and self-correction as training proceeded. The released R1 added a small cold-start SFT stage before RL, because R1-Zero's outputs suffered from poor readability and language mixing.

The critical question is whether RL teaches new capability or sharpens existing capability.

  • Yue et al., 2025 (Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?) compare pass@$k$ at large $k$ and find RL-trained models do not exceed their base models — RL concentrates probability on solutions the base model could already sample, improving pass@1 while sometimes reducing pass@256.
  • The counter-position is that pass@$k$ at large $k$ is not the quantity anyone deploys, and that reliable pass@1 is the capability.

Both readings are consistent with the toy result above: nothing was learned in the first 100 steps because nothing was being sampled to reinforce.

Cost profile

ComponentPPOGRPO
Models in memorypolicy, ref, reward, valuepolicy (+ ref if $\beta>0$)
Forward passes per steprollout + policy + ref + value + rewardrollout + policy ($\times\mu$)
Dominant costgenerationgeneration
Extra hyperparametersGAE $\lambda$, value LR, value clipgroup size

GRPO's real-world speed advantage is smaller than the memory saving suggests, because generation dominates the step time in both. Its decisive advantage is that it removes an entire model and its hyperparameters from a notoriously fragile pipeline.

Papers

What to learn next