Policy gradients
Instead of scoring every action and picking the best, a policy gradient method adjusts the behaviour itself — doing more of whatever led to a good outcome.
- 14 min read
- 3 reading levels
- Updated
Read these first
On this page 10
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
A policy gradient method changes the behaviour directly: whatever you did before something good happened, do more of that.
There is no table of scores. There is only a habit, being adjusted.
The analogy you have already lived
Think about throwing a crumpled paper ball into a bin across the room.
You do not compute a score for every possible arm angle and release point, then pick the best one. Nobody has ever done that. You throw. It misses left. Your next throw drifts a little right, without you deciding anything in words.
Your throwing habit shifts a small amount in the direction that worked.
That is a policy gradient. The habit is the policy. The small shift in a good direction is the gradient.
How this differs from everything so far
Every method so far has done the same thing: score the options, then pick the highest score. Q-learning and DQN both work that way.
Policy gradient methods skip the scoring completely.
value methods: situation -> a score for each action -> pick the biggest
policy methods: situation -> "do this, with this much chance" -> do itThe output is not a score. It is a set of chances. In CartPole it might be "push left, 70 percent of the time; push right, 30 percent".
Why anyone would want this
Three real reasons, and each one is a problem the value methods cannot fix.
A steering wheel has no biggest option. DQN picks the highest-scoring action out of a short list. A steering angle is a smooth range, and you cannot check every value in a range. A policy can output "turn about eight degrees left, give or take two" without listing anything.
Sometimes being unpredictable is correct. In rock-paper-scissors, any fixed choice loses to a human within three rounds. The best play is genuinely random. A value method picks the maximum, which is always the same choice. A policy can be random on purpose.
Sometimes a small change in scores flips the whole behaviour. With a value method, a tiny wobble in two nearly-tied scores swaps which action wins, and behaviour jumps. Policies change smoothly, because the chances move a little rather than the choice flipping.
How the adjustment works
Play a whole episode. Do not learn anything yet — watch.
At the end, look at how it went overall. Then walk back through every action you took.
the episode went well -> every action taken becomes a bit more likely
the episode went badly -> every action taken becomes a bit less likelyThat is genuinely the whole algorithm. It is called REINFORCE, published by Ronald Williams in 1992.
The objection you are about to raise is the right one. A good episode probably contained some bad moves, and they get encouraged too. That is true, and it is exactly why this method is slow and noisy. It relies on averaging: over thousands of episodes, a genuinely bad move will appear in enough failures that it gets pushed down overall.
One repair that changes everything
Compare against average, not against zero.
Suppose every episode scores between 400 and 500. A 410 is a poor result. It is still a positive number, so plain REINFORCE encourages everything you did in it.
The fix: subtract the typical score first. Now a 410 comes out negative and gets discouraged, while a 490 comes out positive. The thing you subtract is called a baseline.
This one change is the difference between a method that barely works and one that works. The next lesson is about learning a smart baseline instead of using a simple average.
The price you pay
You must throw your experience away after using it once. These methods learn about the habit you currently have. Change the habit, and everything you recorded describes a habit you no longer have.
That is called being on-policy, and it is expensive. DQN could replay a moment from an hour ago a hundred times. REINFORCE uses each episode once and deletes it.
You are trading sample efficiency for stability and for the ability to handle smooth actions. Whether that is a good trade depends on whether your episodes are cheap. In a simulator, they are. On a real robot, they are not.
Where you have already seen this
- Robot arms learning to grasp, where joint angles are smooth ranges.
- Data centre and building control, where the setting is a dial rather than a switch.
- Language models tuned with human feedback. The model's choice of next word is already a set of chances over words — it is a policy by construction. That is why RLHF uses policy methods and not DQN.
Remember this
- A policy outputs chances over actions, not scores. You act by sampling from it.
- Good episode: make everything you did more likely. Bad episode: less likely.
- Always subtract a baseline, or the method is far weaker than it needs to be.
What to learn next
- Actor-critic methods — learning the baseline instead of guessing it.
- PPO — how to reuse data without the policy exploding.
- Gradient descent — the optimiser underneath all of this.
Developer — Code and libraries.
Setup
pip install gymnasium torch numpyCPU is enough. This takes about forty-five seconds.
REINFORCE with a baseline, complete
import gymnasium as gym
import numpy as np
import torch
import torch.nn as nn
torch.manual_seed(0); np.random.seed(0)
torch.set_num_threads(1)
env = gym.make("CartPole-v1")
# The policy outputs a probability for each action, not a score for each action.
policy = nn.Sequential(nn.Linear(4, 64), nn.ReLU(), nn.Linear(64, 2))
opt = torch.optim.Adam(policy.parameters(), lr=0.01)
GAMMA = 0.99
scores = []
for episode in range(600):
state, _ = env.reset(seed=episode)
log_probs, rewards, done = [], [], False
while not done: # play one full episode first
logits = policy(torch.tensor(state))
dist = torch.distributions.Categorical(logits=logits)
action = dist.sample() # sample, do not take the max
log_probs.append(dist.log_prob(action))
state, reward, terminated, truncated, _ = env.step(int(action))
rewards.append(reward)
done = terminated or truncated
# return from each step onwards: what actually happened after that move
returns, running = [], 0.0
for r in reversed(rewards):
running = r + GAMMA * running
returns.insert(0, running)
returns = torch.tensor(returns)
returns = (returns - returns.mean()) / (returns.std() + 1e-8) # baseline: better than average?
loss = -(torch.stack(log_probs) * returns).sum() # push up the log-probability of good moves
opt.zero_grad(); loss.backward(); opt.step()
scores.append(sum(rewards))
if (episode + 1) % 100 == 0:
print(f"episode {episode + 1:3d} average of last 100: {np.mean(scores[-100:]):6.1f}")
env.close()episode 100 average of last 100: 67.7 episode 200 average of last 100: 204.3 episode 300 average of last 100: 189.7 episode 400 average of last 100: 481.8 episode 500 average of last 100: 500.0 episode 600 average of last 100: 498.8
500 is the maximum CartPole allows. This reached it and stayed there — noticeably steadier than the DQN in the previous lesson, which oscillated between 300 and 500 on the same task.
Exact numbers will vary across PyTorch versions and machines even with the seed fixed. The pattern to expect is under 100 early, a dip somewhere in the middle, and close to 500 by the end.
Line by line, and the one line everybody gets wrong
dist.sample() — not argmax. The randomness is the exploration. There is no epsilon anywhere in this script, and none is needed: a policy that is uncertain explores by construction, and as it becomes confident the sampling narrows on its own. Replace this with argmax and learning stops dead, because the policy can no longer discover anything it does not already prefer.
loss = -(log_probs * returns).sum() — the line to stare at.
This is not a loss in the supervised sense. There is no correct label anywhere. It is an expression whose gradient happens to equal the policy gradient. log_prob of the action taken, multiplied by how well the episode went, negated because optimisers descend and we want to ascend.
Read what the gradient does. A positive return makes the optimiser increase the log-probability of that action in that state — do more of it. A negative return decreases it. That is the whole method in one line.
(returns - returns.mean()) / (returns.std() + 1e-8) — the baseline, plus a scale normalisation. Delete this line and re-run. Every CartPole reward is +1, so every return is positive, so every action gets encouraged, including the one that dropped the pole. Learning becomes drastically slower and much noisier. This is the single highest-value line in the script.
returns.insert(0, running) — walking backwards accumulates the discounted return from each step onwards in one pass. Building it forwards costs a nested loop and is a common source of off-by-one bugs.
opt.step() once per episode — one gradient step per episode. Compare with DQN, which took a step on every environment step. That is where the sample inefficiency lives.
Why the variance is so high
Consider what the return actually measures. The return after step 3 includes everything that happened at steps 4 through 200 — almost all of which had nothing to do with the action at step 3.
So the learning signal for one action is contaminated by hundreds of unrelated decisions. The estimate is unbiased, meaning it is correct on average, and its variance is enormous. Reducing that variance without introducing bias is the whole research programme that produced actor-critic and PPO.
Common mistakes
Using the total episode return for every action. A common shortcut is loss = -(log_probs * total_return).sum(). It works, and it is worse. An action cannot influence rewards that came before it, so including them adds variance and no information. Use the return from that step onwards, as this script does.
Forgetting the minus sign. Without it you are minimising the return. Your agent will get impressively good at dropping the pole immediately.
Reusing an episode for a second gradient step. These methods are on-policy. Once the parameters move, the recorded log-probabilities describe a policy that no longer exists. Reusing data needs importance weighting, which is what PPO does carefully.
Normalising returns inside a batch of one very short episode. With three steps, returns.std() is noise. On short episodes prefer subtracting a running mean over episodes instead of per-episode normalisation.
A learning rate that is too high. At lr=0.1 this collapses to a policy that always pushes the same way, and it never recovers. A policy that has become deterministic has no gradient left to explore with. This is the failure mode PPO was designed to prevent.
Try it yourself
Delete the returns-normalisation line and run again. Watch how much slower it is. Then put it back and change GAMMA to 0.9: the effective horizon drops to about ten steps, so the agent stops caring about balancing more than ten steps ahead, and its score plateaus far below 500. Both experiments take a minute and teach more than reading about variance does.
What to learn next
- Actor-critic methods — learning the baseline instead of guessing it.
- PPO — how to reuse data without the policy exploding.
- Gradient descent — the optimiser underneath all of this.
Researcher — Mathematics and papers.
The policy gradient theorem
Parameterise $\pi_\theta(a \mid s)$ and maximise $J(\theta) = \mathbb{E}{\tau \sim \pi\theta}[G(\tau)]$. The difficulty is that the trajectory distribution depends on $\theta$, so the gradient cannot pass through the expectation directly.
The log-derivative trick, using $\nabla_\theta p_\theta(x) = p_\theta(x) \nabla_\theta \log p_\theta(x)$:
$$ \nabla_\theta J(\theta) = \mathbb{E}{\tau \sim \pi\theta} \left[ \sum_{t=0}^{T} \nabla_\theta \log \pi_\theta(a_t \mid s_t) \, G_t \right] $$
Because $\log p_\theta(\tau) = \log p(s_0) + \sum_t \left[\log \pi_\theta(a_t \mid s_t) + \log P(s_{t+1} \mid s_t, a_t)\right]$ and the environment terms carry no $\theta$, they vanish under the gradient. The transition dynamics disappear entirely. This is why policy gradients are model-free without needing any approximation.
Sutton et al. (2000) give the general form: $$ \nabla_\theta J(\theta) = \mathbb{E}{s \sim d^{\pi}, a \sim \pi\theta}\left[ \nabla_\theta \log \pi_\theta(a \mid s) \, Q^{\pi_\theta}(s,a) \right] $$ with $d^\pi$ the discounted state visitation distribution. Symbols: $G_t$ the return from $t$, $Q^{\pi}$ the action-value under the current policy, $\tau$ a trajectory.
Baselines are free
For any function $b(s)$ that does not depend on $a$: $$ \mathbb{E}{a \sim \pi\theta}\left[ \nabla_\theta \log \pi_\theta(a \mid s) \, b(s) \right] = b(s) \nabla_\theta \sum_a \pi_\theta(a \mid s) = b(s) \nabla_\theta 1 = 0 $$
So subtracting $b(s)$ from the return leaves the gradient unbiased while changing its variance. The variance-minimising baseline is $b^*(s) = \frac{\mathbb{E}[(\nabla \log \pi)^2 G]}{\mathbb{E}[(\nabla \log \pi)^2]}$, a gradient-magnitude-weighted average of returns. In practice $b(s) = V^\pi(s)$ is used, which is close enough and gives $A^\pi(s,a) = Q^\pi(s,a) - V^\pi(s)$, the advantage. That substitution is the step into actor-critic.
Note that the per-episode standardisation used in the code is not a state-dependent baseline — it is a batch normalisation of returns. It reduces scale sensitivity and it does introduce a small bias. It is nonetheless the near-universal practical choice.
Variance, quantified
The REINFORCE estimator's variance grows roughly as $O(T^2)$ in the horizon under mild assumptions, because $G_t$ sums $T-t$ random rewards and appears in $T$ terms. Three standard reductions:
- Reward-to-go. Replace $G(\tau)$ with $G_t$, dropping rewards that precede the action. Unbiased by causality.
- Baseline subtraction. Unbiased, as shown above.
- Bootstrapping. Replace $G_t$ with $r_t + \gamma V(s_{t+1})$. This does introduce bias and cuts variance sharply. GAE (Schulman et al., 2016) parameterises the whole spectrum with $\lambda$.
The on-policy constraint
The expectation is over $\tau \sim \pi_\theta$ for the current $\theta$. After one gradient step the data is off-policy and the estimator is biased. Correcting with importance sampling gives $$ \nabla_\theta J = \mathbb{E}{\tau \sim \pi{\theta_{\text{old}}}}\left[ \frac{\pi_\theta(a\mid s)}{\pi_{\theta_{\text{old}}}(a \mid s)} \nabla_\theta \log \pi_\theta(a\mid s) A(s,a) \right] $$ but the importance ratio has unbounded variance over long trajectories. Constraining or clipping that ratio is precisely what TRPO and PPO do, and it is what makes limited data reuse safe.
Natural policy gradients
The plain gradient is steepest ascent in Euclidean parameter space, which is the wrong geometry — equal parameter changes produce wildly unequal changes in the policy distribution. The natural gradient (Amari, 1998; Kakade, 2002) preconditions by the Fisher information matrix: $$ \tilde{\nabla}\theta J = F(\theta)^{-1} \nabla\theta J, \qquad F(\theta) = \mathbb{E}\left[ \nabla_\theta \log \pi_\theta \, \nabla_\theta \log \pi_\theta^\top \right] $$ This is steepest ascent under the KL divergence between policies, which is invariant to reparameterisation. Computing $F^{-1}$ is infeasible for large networks; TRPO approximates it with conjugate gradients, and PPO abandons it for a much cheaper clipping heuristic that works nearly as well.
Continuous actions
For $\mathcal{A} \subseteq \mathbb{R}^d$, output a distribution — typically a diagonal Gaussian $\pi_\theta(a\mid s) = \mathcal{N}(\mu_\theta(s), \operatorname{diag}(\sigma_\theta^2))$. The log-probability is differentiable and the same estimator applies unchanged. A state-independent learned $\log\sigma$ is standard and usually more stable than a state-dependent one. Bounded action spaces need care: naive clipping biases the gradient, and a tanh-squashed Gaussian with the corresponding log-det-Jacobian correction (as in SAC) is the cleaner treatment.
References
- Williams (1992), Simple statistical gradient-following algorithms for connectionist reinforcement learning — REINFORCE.
- Sutton, McAllester, Singh and Mansour (2000), Policy Gradient Methods for Reinforcement Learning with Function Approximation.
- Kakade (2002), A Natural Policy Gradient.
- Schulman et al. (2016), High-Dimensional Continuous Control Using Generalized Advantage Estimation — arxiv.org/abs/1506.02438.
- Greensmith, Bartlett and Baxter (2004), Variance Reduction Techniques for Gradient Estimates in Reinforcement Learning.
What to learn next
- Actor-critic methods — learning the baseline instead of guessing it.
- PPO — how to reuse data without the policy exploding.
- Gradient descent — the optimiser underneath all of this.