PPO
PPO is the default reinforcement learning algorithm today — an actor-critic method with one small rule that stops the policy changing too much in a single update.
- 15 min read
- 3 reading levels
- Updated
Read these first
On this page 9
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
PPO is an actor-critic method with a rule: never change your behaviour by more than a little in one go.
That one rule is why it is the algorithm most people reach for first.
The analogy you have already lived
Think about a shower with a badly labelled hot tap.
The water is cold. You turn the tap a long way. Ten seconds later it is scalding, so you yank it back. Now it is freezing. You are stuck oscillating, and you never find the setting you wanted.
Now do it properly. Turn the tap a small amount. Wait. Feel it. Turn it a small amount again. You arrive at the right temperature in under a minute, and you never get burned.
PPO is the second approach. Small, deliberate adjustments, each one checked before the next.
The problem it fixes
The actor-critic lesson ended on a warning. Two learners depending on each other can drift apart, and a single overlarge update can wreck a policy that was working.
The wreckage has a specific shape. The agent collects some experience. It computes an update. The update is too big. The new policy now behaves nothing like the one that gathered the data. The data no longer describes it. So the next update is advice about a policy that has ceased to exist.
One bad step and the agent is worse than it was an hour ago, with no route back.
This is the single most common way reinforcement learning fails, and before PPO the fix was mathematically heavy and slow to run.
What PPO does instead
Play for a while and record what happened. Then improve the policy — but refuse to let the chance of any action move too far from what it was when the data was gathered.
chance of pushing left, when the data was collected: 0.40
PPO will let it move to anywhere in: 0.32 to 0.48
| |
-20% +20%
an update that wants 0.90? the extra benefit is ignored.The refusal is the whole innovation. If an update wants to move a probability well beyond the allowed band, PPO stops counting the benefit of going further. There is nothing to gain from the extra distance, so the optimiser does not go there.
The band is usually twenty percent either way. That number is a choice, not a law, and it is the main knob people turn.
The bonus this buys you
The policy is not allowed to run away. So PPO can safely do something the earlier methods could not: use the same batch of experience several times.
REINFORCE had to throw each episode away after one gradient step. PPO takes several passes over the same data before collecting more. Same experience, several times the learning.
That is why PPO reaches a good score in far fewer environment steps than plain policy gradients. It is a large part of why it became the default.
What is honestly hard here
PPO is famous for being simple and it is famous among practitioners for something else: the details matter enormously.
A careful study went through public PPO implementations. It found dozens of small choices that papers rarely mention. How observations are scaled. How the learning rate decays. How the starting weights are set. Whether advantages are normalised. Several of those contribute more to the final score than the clipping rule the paper is named after.
This is not a criticism of PPO. It is a warning about reading any reinforcement learning paper. The algorithm in the equations is often not the algorithm that produced the numbers.
If you are implementing it yourself, start from a known-good reference implementation and change one thing at a time.
Where you have already seen it
- The assistant you have used. PPO was the algorithm in the original RLHF recipe for tuning language models to human preferences.
- Robot control, in both simulation and hardware.
- Game-playing agents trained at very large scale.
- The default in most reinforcement learning libraries, because it works acceptably on a wide range of problems without much tuning.
Remember this
- PPO limits how far the policy may move in one update. That is the whole idea.
- The limit lets it reuse the same experience several times, which is where the efficiency comes from.
- It is the sensible default, and its implementation details matter more than its equation.
What to learn next
- RLHF — PPO doing the job it is most widely used for today.
- Reward shaping — because a stable algorithm on a wrong reward is still wrong.
- Actor-critic methods — the foundation PPO is built on.
Developer — Code and libraries.
Setup
pip install gymnasium torch numpyCPU only, and this is the fastest script in the section — about eleven seconds to reach the maximum score on CartPole.
PPO, complete, in sixty lines
import gymnasium as gym
import numpy as np
import torch
import torch.nn as nn
torch.manual_seed(0); np.random.seed(0)
torch.set_num_threads(1)
env = gym.make("CartPole-v1")
actor = nn.Sequential(nn.Linear(4, 64), nn.Tanh(), nn.Linear(64, 2))
critic = nn.Sequential(nn.Linear(4, 64), nn.Tanh(), nn.Linear(64, 1))
opt = torch.optim.Adam(list(actor.parameters()) + list(critic.parameters()), lr=3e-3)
ROLLOUT, EPOCHS, MINIBATCH = 1024, 4, 128
GAMMA, LAM, CLIP = 0.99, 0.95, 0.2
state, _ = env.reset(seed=0)
episode_return, finished = 0.0, []
for update in range(40):
S, A, LOGP, R, D, V = [], [], [], [], [], []
for _ in range(ROLLOUT):
st = torch.tensor(state)
with torch.no_grad():
dist = torch.distributions.Categorical(logits=actor(st))
action = dist.sample()
S.append(state); A.append(int(action))
LOGP.append(dist.log_prob(action).item()) # the old policy's opinion, frozen
V.append(critic(st).item())
state, reward, terminated, truncated, _ = env.step(int(action))
R.append(reward); D.append(float(terminated))
episode_return += reward
if terminated or truncated:
finished.append(episode_return)
episode_return = 0.0
state, _ = env.reset()
with torch.no_grad():
last_value = critic(torch.tensor(state)).item()
# generalised advantage estimation: a smoothed "was that better than expected?"
adv, gae = np.zeros(ROLLOUT, dtype=np.float32), 0.0
for t in reversed(range(ROLLOUT)):
next_v = last_value if t == ROLLOUT - 1 else V[t + 1]
delta = R[t] + GAMMA * next_v * (1 - D[t]) - V[t]
gae = delta + GAMMA * LAM * (1 - D[t]) * gae
adv[t] = gae
ret = adv + np.array(V, dtype=np.float32)
S = torch.tensor(np.array(S)); A = torch.tensor(A)
LOGP = torch.tensor(LOGP); ADV = torch.tensor(adv); RET = torch.tensor(ret)
ADV = (ADV - ADV.mean()) / (ADV.std() + 1e-8)
for _ in range(EPOCHS):
for i in torch.randperm(ROLLOUT).split(MINIBATCH):
dist = torch.distributions.Categorical(logits=actor(S[i]))
ratio = torch.exp(dist.log_prob(A[i]) - LOGP[i]) # how much the policy moved
unclipped = ratio * ADV[i]
clipped = torch.clamp(ratio, 1 - CLIP, 1 + CLIP) * ADV[i]
actor_loss = -torch.min(unclipped, clipped).mean() # the clip is the whole idea
critic_loss = nn.functional.mse_loss(critic(S[i]).squeeze(1), RET[i])
loss = actor_loss + 0.5 * critic_loss - 0.01 * dist.entropy().mean()
opt.zero_grad(); loss.backward(); opt.step()
if (update + 1) % 5 == 0 and finished:
print(f"update {update + 1:2d} {(update + 1) * ROLLOUT:6d} steps "
f"average episode return: {np.mean(finished[-20:]):6.1f}")
env.close()update 5 5120 steps average episode return: 74.2 update 10 10240 steps average episode return: 162.2 update 15 15360 steps average episode return: 201.2 update 20 20480 steps average episode return: 242.7 update 25 25600 steps average episode return: 383.8 update 30 30720 steps average episode return: 459.6 update 35 35840 steps average episode return: 486.6 update 40 40960 steps average episode return: 500.0
Exact numbers will vary — the mid-rollout env.reset() is unseeded, and floating-point order differs between builds. What should not vary is the shape: a steady climb, no collapse.
Compare that with the DQN lesson, where the score reached 350 and then fell back to 72 on a longer run. That contrast is the reason PPO became the default.
The three lines that are PPO
ratio = torch.exp(dist.log_prob(A[i]) - LOGP[i])
clipped = torch.clamp(ratio, 1 - CLIP, 1 + CLIP) * ADV[i]
actor_loss = -torch.min(ratio * ADV[i], clipped).mean()ratio is the new policy's probability of the action divided by the old one's. Subtracting logs and exponentiating is the numerically stable way to write that division. A ratio of 1.0 means the policy has not moved on this action.
torch.clamp(ratio, 0.8, 1.2) caps the ratio inside the band.
torch.min(unclipped, clipped) is the subtle part, and it is worth being precise about. Taking the minimum of the two means the objective is a pessimistic estimate.
- Advantage positive, ratio already above 1.2: the clipped term is smaller, so it is chosen, and its gradient is zero. No reward for pushing further.
- Advantage negative, ratio below 0.8: again the clipped term is chosen, and again the gradient is zero. No reward for pushing further down.
- Advantage negative and ratio very large: here
minpicks the unclipped term, which is a large negative number with a live gradient. An action that turned out badly and has become much more likely gets pulled back hard, with no cap.
That asymmetry is deliberate. The clip removes the incentive to over-improve, and it does not remove the ability to undo a mistake.
The rest of the machinery
GAE, the backwards loop. delta is the one-step surprise: reward plus the critic's view of the next state, minus its view of this one. The accumulator gae blends these across time steps with weight GAMMA * LAM. With LAM = 1.0 you get the full Monte-Carlo advantage; with LAM = 0.0 you get the one-step version. 0.95 is the standard compromise.
(1 - D[t]) appears twice in the GAE loop, and both are necessary. A terminated state has no future, so neither the bootstrap nor the accumulated advantage may cross that boundary.
ADV = (ADV - ADV.mean()) / (ADV.std() + 1e-8) — advantage normalisation. Not in the original paper, present in essentially every implementation, and it matters more than most of the hyperparameters.
EPOCHS = 4 and minibatches — four passes over the same 1024 steps. This is the data reuse the clip makes safe. Raise it to 20 and the ratio drifts far outside the band on the later passes, where the gradient is zero, so you burn compute for nothing and risk instability.
nn.Tanh() rather than ReLU — the convention in PPO implementations for continuous control, and it does measurably matter. Tanh's bounded output keeps activations in a stable range across the many small updates.
Common mistakes
Recomputing LOGP from the current policy. It must be captured under torch.no_grad() at collection time and frozen. Recompute it inside the epoch loop and ratio is always exactly 1.0, the clip never activates, and you have written plain policy gradient with extra steps.
Too many epochs. More is not better. Past a handful of passes the data no longer describes the current policy, and the clip protects you by producing zero gradient rather than by giving useful learning.
Bootstrapping through a truncation. D[t] here stores terminated only. CartPole's 500-step limit is a truncation, and the pole is still up. Storing terminated or truncated teaches the critic that a perfect episode ends in worthlessness.
Assuming the clip guarantees a small policy change. It does not. The clip bounds the ratio per action, not the KL divergence overall, and a large enough learning rate can still move the policy a long way. Production implementations monitor the approximate KL each update and stop early if it exceeds a threshold.
Sharing a network body without care. With a shared trunk, the value loss coefficient becomes critical. Separate networks, as here, are more forgiving on small problems.
Try it yourself
Set CLIP = 100.0, effectively disabling the clip, and run again. On some seeds it still works, and on others it collapses partway through — which is exactly the point, since without the clip you are running unconstrained policy gradient with four epochs of reuse. Then set CLIP = 0.02 and watch it learn correctly but slowly, because every update is now tiny. The default of 0.2 sits between those two failure modes.
What to learn next
- RLHF — PPO doing the job it is most widely used for today.
- Reward shaping — because a stable algorithm on a wrong reward is still wrong.
- Actor-critic methods — the foundation PPO is built on.
Researcher — Mathematics and papers.
From TRPO to PPO
TRPO (Schulman et al., 2015) maximises a surrogate objective subject to a hard KL constraint:
$$ \max_\theta \; \mathbb{E}t!\left[ \frac{\pi\theta(a_t \mid s_t)}{\pi_{\theta_{\text{old}}}(a_t \mid s_t)} \hat{A}_t \right] \quad \text{subject to} \quad \mathbb{E}t!\left[ \mathrm{KL}!\left( \pi{\theta_{\text{old}}}(\cdot \mid s_t) \,|\, \pi_\theta(\cdot \mid s_t) \right) \right] \le \delta $$
The motivation is a monotonic improvement guarantee (Kakade and Langford, 2002): the true return satisfies $J(\pi') \ge L_{\pi}(\pi') - C \cdot \max_s \mathrm{KL}(\pi | \pi')$ with $C = \frac{4\epsilon\gamma}{(1-\gamma)^2}$ and $\epsilon = \max_{s,a}|A^\pi(s,a)|$. Optimising the right-hand side never decreases $J$. In practice the theoretical $C$ forces steps so small that TRPO uses $\delta$ as a tuned hyperparameter instead, discarding the guarantee it was derived from.
TRPO solves the constrained problem with a conjugate-gradient step against a Fisher-vector product, plus a backtracking line search. It works, and it is awkward: second-order, hard to combine with parameter sharing, and unpleasant with recurrent networks.
The clipped surrogate
PPO (Schulman et al., 2017) replaces the constraint with a clip. With $r_t(\theta) = \frac{\pi_\theta(a_t\mid s_t)}{\pi_{\theta_{\text{old}}}(a_t\mid s_t)}$:
$$ L^{\text{CLIP}}(\theta) = \mathbb{E}_t\left[ \min\Big( r_t(\theta)\hat{A}_t, \; \mathrm{clip}\big(r_t(\theta), 1-\epsilon, 1+\epsilon\big)\hat{A}_t \Big) \right] $$
The $\min$ makes $L^{\text{CLIP}}$ a lower bound on the unclipped surrogate. Gradient behaviour, stated exactly:
| Case | $\min$ selects | Gradient |
|---|---|---|
| $\hat{A}_t > 0$, $r_t < 1+\epsilon$ | unclipped | active |
| $\hat{A}_t > 0$, $r_t \ge 1+\epsilon$ | clipped | zero |
| $\hat{A}_t < 0$, $r_t > 1-\epsilon$ | unclipped | active |
| $\hat{A}_t < 0$, $r_t \le 1-\epsilon$ | clipped | zero |
The full objective adds a value loss and an entropy bonus: $L = L^{\text{CLIP}} - c_1 (V_\phi(s_t) - \hat{R}_t)^2 + c_2 \mathcal{H}\pi_\theta$, with $c_1 \approx 0.5$ and $c_2 \approx 0.01$.
PPO has no monotonic improvement guarantee. The clip bounds the per-sample likelihood ratio, not the expected KL, and Engstrom et al. (2020) exhibit runs where the mean KL substantially exceeds what the clip suggests. It is a heuristic that works well, which is a different claim from a theorem.
The implementation details
Engstrom et al. (2020), Implementation Matters in Deep Policy Gradients, and Andrychowicz et al. (2020), What Matters in On-Policy Reinforcement Learning?, both ran large ablations. Their combined findings are uncomfortable and worth taking seriously.
The choices that mattered most, several of which appear nowhere in the PPO paper:
- Observation normalisation with a running mean and variance. Large effect on continuous control.
- Reward scaling by a running estimate of the return's standard deviation.
- Advantage normalisation per minibatch.
- Orthogonal initialisation with a policy output layer scaled by $0.01$, so the initial policy is near-uniform.
- Learning rate annealing to zero over training.
- Gradient clipping at global norm 0.5.
- Value function loss clipping — contested; several studies find it neutral or harmful.
- Adam epsilon at $10^{-5}$ rather than the default $10^{-8}$.
Engstrom et al.'s sharpest result: TRPO with PPO's code-level optimisations matches PPO. Much of the reported gap between the two algorithms came from the implementation, not the objective.
The practical consequence is that a from-scratch PPO that underperforms a library's is usually missing details, not misunderstanding the maths. Huang et al. (2022), The 37 Implementation Details of Proximal Policy Optimization, is the reference to work through.
Variants
- PPO-Penalty. Adds $-\beta \,\mathrm{KL}$ to the objective with $\beta$ adapted to hit a target KL. Used in the original RLHF work, where an explicit KL budget against the pretrained model is the point rather than an implementation detail.
- Early stopping on KL. Abandon remaining epochs once the approximate KL exceeds roughly $1.5\times$ the target. Cheap and effective.
- Phasic Policy Gradient (Cobbe et al., 2021). Separates policy and value optimisation into distinct phases, letting the value function take many more passes than the policy safely.
- DAPO, GRPO and relatives (2024-2025). Group-relative variants that drop the learned critic and estimate the advantage from a group of sampled responses to the same prompt. Now common in LLM reasoning training, where a value network over token sequences is expensive and poorly conditioned.
References
- Schulman et al. (2015), Trust Region Policy Optimization — arxiv.org/abs/1502.05477.
- Schulman et al. (2017), Proximal Policy Optimization Algorithms — arxiv.org/abs/1707.06347.
- Engstrom et al. (2020), Implementation Matters in Deep Policy Gradients — arxiv.org/abs/2005.12729.
- Andrychowicz et al. (2020), What Matters in On-Policy Reinforcement Learning? — arxiv.org/abs/2006.05990.
- Huang et al. (2022), The 37 Implementation Details of Proximal Policy Optimization.
- Cobbe et al. (2021), Phasic Policy Gradient — arxiv.org/abs/2009.04416.
What to learn next
- RLHF — PPO doing the job it is most widely used for today.
- Reward shaping — because a stable algorithm on a wrong reward is still wrong.
- Actor-critic methods — the foundation PPO is built on.