Actor-critic methods
Two networks working together — one decides what to do, the other judges how well things are going — which cuts the noise that makes plain policy gradients so slow.
- 14 min read
- 3 reading levels
- Updated
Read these first
On this page 10
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
One network chooses the actions. A second network judges how well the situation is going. The first learns from the second's opinion instead of waiting for the end.
The analogy you have already lived
Imagine learning to cook a new dish with a friend who keeps tasting.
Without the friend, you find out how it went only when everyone eats, an hour later. You then have to guess which of your forty decisions was the problem. Too much salt? Wrong heat? Wrong moment for the tomatoes?
With the friend tasting as you go, you get an opinion after every step. "Better than it was." "Worse than it was." You can correct the thing you did five seconds ago.
The cook is the actor, the part that picks the moves. The tasting friend is the critic, the part that judges how things are going.
The problem this fixes
The previous lesson had an honest weakness. REINFORCE plays a whole episode, sees the total, and then encourages or discourages every single action in it based on that one number.
A great episode probably contained three bad moves. They get encouraged. A terrible episode probably contained a brilliant save. It gets discouraged.
Over thousands of episodes it averages out. But "thousands of episodes" is the price, and the learning curve is jumpy the whole way.
The critic replaces the end-of-episode verdict with an immediate opinion.
What the critic actually says
The critic answers one question: "from here, how well is this likely to go?"
Not "what should I do". Only "how good is this situation".
Then the actor gets a far sharper signal:
what the critic expected from this situation: 50
what actually happened after the action: 62
-------------------------------------------------
the action did better than expected by: +12 -> do more of itThat difference has a name: the advantage. How much better this action turned out than the situation deserved.
Notice how much cleaner that is. It does not matter whether 62 is a big number or a small one. What matters is whether it beat what was expected from that spot.
Why "better than expected" is the right question
Suppose an agent is in a hopeless position. Every action leads to a score of about 10, and the best possible is 11.
REINFORCE sees a return of 11, a positive number, and encourages what was done.
Actor-critic sees that the critic expected 10, the outcome was 11, and the advantage is +1. Small, honest, correctly proportioned.
Now put the agent in a wonderful position where everything gives about 500. It takes a poor action and gets 480. REINFORCE sees a large positive number and encourages a bad move. Actor-critic sees an advantage of minus 20 and pushes it down.
Comparing against expectation, not against zero, is the whole idea.
The catch nobody mentions first
At the start of training, the critic is useless. It has seen nothing and its opinions are noise.
So the actor is learning from a judge that is itself learning. Two learners, each depending on the other, each moving under the other's feet.
This is genuinely hard, and it is worth saying plainly. When actor-critic training fails, it is usually this. The critic falls behind, so the actor gets bad advice, so it visits strange situations, which the critic has never seen, so its advice gets worse. The whole thing spirals.
That fragility is the reason PPO exists. It limits how far the actor may move in one step. That limit is what keeps the pair together.
The shape of it
the situation
/ \
/ \
[ actor ] [ critic ]
| |
chances over actions one number:
| "how good is
pick one it here?"
| |
v v
do it ----> compare what happened
with what was expected
|
v
actor: do more of what beat expectation
critic: get better at expectingWhere you have already seen it
Almost everywhere modern reinforcement learning is used, because almost every modern method is an actor-critic method.
- Robot control in simulation and on real hardware.
- Game-playing agents in complex real-time games.
- Language model tuning. The version of RLHF used on the assistants you have talked to is an actor-critic method with an extra safety strap.
Remember this
- The actor picks actions. The critic judges how good the situation is.
- The actor learns from the advantage: how much better the outcome was than expected.
- Two learners depending on each other is powerful and fragile, in that order.
What to learn next
- PPO — the actor-critic method that fixed the instability well enough to be a default.
- Policy gradients — the estimator the critic is reducing the variance of.
- RLHF — actor-critic applied to language models.
Developer — Code and libraries.
Setup
pip install gymnasium torch numpyCPU only. This runs in about twenty-five seconds — roughly half the time REINFORCE took to reach the same score.
Actor and critic, both from scratch
import gymnasium as gym
import numpy as np
import torch
import torch.nn as nn
torch.manual_seed(0); np.random.seed(0)
torch.set_num_threads(1)
env = gym.make("CartPole-v1")
actor = nn.Sequential(nn.Linear(4, 64), nn.ReLU(), nn.Linear(64, 2)) # what to do
critic = nn.Sequential(nn.Linear(4, 64), nn.ReLU(), nn.Linear(64, 1)) # how good this state is
opt = torch.optim.Adam(list(actor.parameters()) + list(critic.parameters()), lr=0.01)
GAMMA = 0.99
scores = []
for episode in range(600):
state, _ = env.reset(seed=episode)
states, actions, rewards, done = [], [], [], False
while not done:
logits = actor(torch.tensor(state))
dist = torch.distributions.Categorical(logits=logits)
action = dist.sample()
states.append(state); actions.append(int(action))
state, reward, terminated, truncated, _ = env.step(int(action))
rewards.append(reward)
done = terminated or truncated
returns, running = [], 0.0
for r in reversed(rewards):
running = r + GAMMA * running
returns.insert(0, running)
s = torch.tensor(np.array(states))
a = torch.tensor(actions)
g = torch.tensor(returns, dtype=torch.float32)
values = critic(s).squeeze(1)
advantage = g - values # was the outcome better than the critic expected?
dist = torch.distributions.Categorical(logits=actor(s))
actor_loss = -(dist.log_prob(a) * advantage.detach()).mean()
critic_loss = nn.functional.mse_loss(values, g)
entropy = dist.entropy().mean() # a nudge to keep some randomness alive
loss = actor_loss + 0.5 * critic_loss - 0.01 * entropy
opt.zero_grad(); loss.backward(); opt.step()
scores.append(sum(rewards))
if (episode + 1) % 100 == 0:
print(f"episode {episode + 1:3d} avg of last 100: {np.mean(scores[-100:]):6.1f}"
f" critic's guess at the start state: {critic(torch.tensor(states[0])).item():6.1f}")
env.close()episode 100 avg of last 100: 15.3 critic's guess at the start state: 15.0 episode 200 avg of last 100: 271.5 critic's guess at the start state: 73.3 episode 300 avg of last 100: 353.5 critic's guess at the start state: 67.2 episode 400 avg of last 100: 236.8 critic's guess at the start state: 84.9 episode 500 avg of last 100: 443.5 critic's guess at the start state: 92.7 episode 600 avg of last 100: 500.0 critic's guess at the start state: 104.6
Exact numbers vary by PyTorch build and machine even with seeds fixed. The trend is what matters.
Watch the critic learn its own job
The second column is the interesting one. The critic's estimate of the starting position climbs: 15, then 73, then 92, then 105.
Early on it says 15 because the pole really does fall after about 15 steps. It is not wrong; the agent is bad. As the actor improves, the same starting position becomes genuinely more valuable, and the critic tracks it.
There is a ceiling. With GAMMA = 0.99 and a 500-step limit, the largest possible discounted return is about 99.3. The critic settles a little above that at 104.6, which is honest approximation error rather than a bug — the critic is a small network fitting noisy targets, and it overshoots slightly. If you saw 400 there, that would be a bug.
Notice also the dip at episode 400 (236.8) before the recovery. Actor-critic on CartPole is more stable than DQN and it is not monotone.
Line by line
advantage = g - values — the whole method. g is what actually happened. values is what the critic expected. The difference is the advantage.
advantage.detach() — this is the line people get wrong, and it is silent when wrong. Without detach, the actor's loss sends gradients into the critic, and the optimiser discovers it can reduce the actor loss by making the critic's predictions worse. The critic is corrupted, the advantage becomes meaningless, and training limps along without any error message. Detach the advantage before it touches the actor.
One optimiser over both parameter lists — convenient here. In larger setups the critic usually wants a higher learning rate than the actor, so two optimisers is the more common arrangement.
0.5 * critic_loss — the critic's loss is on the scale of returns (up to about 100) while the actor's is on the scale of log-probabilities (around 1). Without a weight, the critic's gradients drown the actor's. That 0.5 is a real tuning knob, not decoration.
- 0.01 * entropy — the entropy bonus. Entropy measures how undecided the policy is. Subtracting it from the loss rewards staying undecided. Without it, a policy that gets an early lead can collapse to always-push-left, and a deterministic policy has nothing left to explore with. Set the coefficient to 0.0 and you will see this happen on some seeds.
nn.functional.mse_loss(values, g) — the critic is doing plain supervised regression. Its inputs are states, its targets are the observed returns. Nothing about it is special; the difficulty is that its targets keep changing as the actor improves.
The version used in practice
This script uses full-episode returns for the critic's target, which makes it a Monte-Carlo actor-critic. Real implementations bootstrap instead:
# one-step target: reward now, plus the critic's opinion of where we landed
target = reward + GAMMA * critic(next_state) * (1 - terminated)That reduces variance sharply and adds bias. Generalised advantage estimation (GAE) blends the two with a parameter lam: lam = 1.0 gives the Monte-Carlo version above, lam = 0.0 gives the one-step version, and lam = 0.95 is the near-universal default. You will meet GAE fully in the PPO lesson.
Common mistakes
Forgetting .detach() on the advantage. Described above. This is the most common actor-critic bug and it does not raise anything.
Sharing a network body between actor and critic without balancing the losses. A shared trunk saves compute and couples the two objectives. If the critic's gradients dominate, the shared features become good for value prediction and useless for choosing actions.
No entropy bonus. Premature collapse to a deterministic policy. Symptom: the score climbs, then freezes exactly, and the policy's entropy is near zero.
Using an advantage of zero mean but wild scale. Normalising advantages per batch ((adv - adv.mean()) / (adv.std() + 1e-8)) is standard and makes the actor's learning rate far less sensitive.
Bootstrapping through a time-limit truncation. When CartPole stops at 500 steps the pole is still up. Treating that as a terminal state teaches the critic that surviving 500 steps is worth nothing after step 500.
Try it yourself
Set the entropy coefficient to 0.0 and run five seeds. On at least one, the policy will collapse early and the score will freeze low. Then set it to 0.1 and watch the opposite: the agent stays deliberately random and never commits, so the score plateaus in the middle. The right value is a balance, and finding it by hand is the fastest way to understand what entropy is buying you.
What to learn next
- PPO — the actor-critic method that fixed the instability well enough to be a default.
- Policy gradients — the estimator the critic is reducing the variance of.
- RLHF — actor-critic applied to language models.
Researcher — Mathematics and papers.
The construction
Start from the policy gradient with a state-dependent baseline:
$$ \nabla_\theta J(\theta) = \mathbb{E}{s \sim d^\pi, a \sim \pi\theta}\left[ \nabla_\theta \log \pi_\theta(a \mid s) \, A^{\pi}(s,a) \right], \qquad A^\pi(s,a) = Q^\pi(s,a) - V^\pi(s) $$
$A^\pi$ is the advantage function. Since $\mathbb{E}_{a\sim\pi}[A^\pi(s,a)] = 0$, it has lower magnitude and lower variance than $Q^\pi$ while giving the same expected gradient.
Neither $Q^\pi$ nor $V^\pi$ is known, so the critic $V_\phi$ estimates $V^\pi$ and the advantage is estimated from data. Common estimators, in increasing bias and decreasing variance:
$$ \hat{A}_t^{(\infty)} = G_t - V_\phi(s_t) \qquad \hat{A}t^{(n)} = \sum{k=0}^{n-1}\gamma^k r_{t+k+1} + \gamma^n V_\phi(s_{t+n}) - V_\phi(s_t) \qquad \hat{A}_t^{(1)} = \delta_t = r_{t+1} + \gamma V_\phi(s_{t+1}) - V_\phi(s_t) $$
Generalised advantage estimation
Schulman et al. (2016) take an exponentially weighted average over all $n$:
$$ \hat{A}t^{\text{GAE}(\gamma,\lambda)} = \sum{l=0}^{\infty} (\gamma\lambda)^l \delta_{t+l}, \qquad \delta_t = r_{t+1} + \gamma V_\phi(s_{t+1}) - V_\phi(s_t) $$
$\lambda = 1$ recovers the Monte-Carlo advantage (unbiased, high variance); $\lambda = 0$ gives the one-step TD advantage (biased by the critic's error, low variance). $\lambda \in [0.92, 0.98]$ is standard. Computed backwards in one pass:
gae = 0.0
for t in reversed(range(T)):
delta = r[t] + gamma * V[t+1] * (1 - done[t]) - V[t]
gae = delta + gamma * lam * (1 - done[t]) * gae
adv[t] = gaeNote that the bias introduced by $\lambda < 1$ is proportional to the critic's error, so GAE and critic quality are not independent knobs.
Compatible function approximation
Sutton et al. (2000) prove that substituting an approximation $f_w(s,a)$ for $Q^\pi$ leaves the gradient exact if two conditions hold:
- $\nabla_w f_w(s,a) = \nabla_\theta \log \pi_\theta(a \mid s)$ — the critic's features are the policy's score function.
- $w$ minimises $\mathbb{E}[(Q^\pi(s,a) - f_w(s,a))^2]$.
This is the compatible function approximation theorem, and it is why actor-critic is more than a heuristic. In practice condition (1) is violated by every deep implementation, so the gradient is biased. The theorem tells you what you are giving up, which is more useful than a guarantee nobody satisfies.
The algorithm family
| Method | Distinguishing feature |
|---|---|
| A3C (Mnih et al., 2016) | Asynchronous parallel workers, lock-free Hogwild updates |
| A2C | Synchronous batched A3C; equal or better, and simpler |
| TRPO (Schulman et al., 2015) | Hard KL trust region via conjugate gradient |
| PPO (Schulman et al., 2017) | Clipped surrogate, first-order, data reuse |
| DDPG / TD3 | Deterministic actor, off-policy, continuous actions |
| SAC (Haarnoja et al., 2018) | Maximum-entropy objective, off-policy, twin critics |
A3C's asynchrony was originally credited with decorrelating updates. Later analysis showed the synchronous A2C matches or beats it, and the parallelism, not the asynchrony, was the useful part.
SAC deserves separate attention. It maximises $\mathbb{E}[\sum_t \gamma^t (r_t + \alpha \mathcal{H}(\pi(\cdot \mid s_t)))]$ — return plus policy entropy — making the entropy bonus part of the objective rather than a regulariser bolted on. With automatic tuning of $\alpha$ against a target entropy, it is the most reliable off-policy continuous-control method available and typically the right default for robotics.
Why the coupled system is unstable
Three interacting failure modes, worth naming precisely:
- Non-stationary regression. $V_\phi$ chases $V^{\pi_\theta}$ while $\theta$ moves. The critic's target distribution shifts with every actor update.
- Distribution shift. The critic is accurate only where the current policy visits. A large actor update moves the visitation distribution into regions where the critic extrapolates, and its errors then drive the actor further out.
- Error amplification. A biased advantage produces a biased gradient, which changes the policy, which changes the states the critic sees. Nothing in the loop damps this.
Every practical stabiliser attacks one of these: trust regions and clipping bound (2); target networks slow (1); advantage normalisation bounds the scale of (3).
References
- Barto, Sutton and Anderson (1983), Neuronlike adaptive elements that can solve difficult learning control problems — the original actor-critic, on the cart-pole task used above.
- Konda and Tsitsiklis (2000), Actor-Critic Algorithms — the two-timescale convergence analysis.
- Mnih et al. (2016), Asynchronous Methods for Deep Reinforcement Learning — arxiv.org/abs/1602.01783.
- Schulman et al. (2016), High-Dimensional Continuous Control Using Generalized Advantage Estimation — arxiv.org/abs/1506.02438.
- Haarnoja et al. (2018), Soft Actor-Critic — arxiv.org/abs/1801.01290.
What to learn next
- PPO — the actor-critic method that fixed the instability well enough to be a default.
- Policy gradients — the estimator the critic is reducing the variance of.
- RLHF — actor-critic applied to language models.