Reinforcement Learning

Actor-critic methods

Two networks working together — one decides what to do, the other judges how well things are going — which cuts the noise that makes plain policy gradients so slow.

On this page 10
  1. The short answer
  2. The analogy you have already lived
  3. The problem this fixes
  4. What the critic actually says
  5. Why "better than expected" is the right question
  6. The catch nobody mentions first
  7. The shape of it
  8. Where you have already seen it
  9. Remember this
  10. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

One network chooses the actions. A second network judges how well the situation is going. The first learns from the second's opinion instead of waiting for the end.

The analogy you have already lived

Imagine learning to cook a new dish with a friend who keeps tasting.

Without the friend, you find out how it went only when everyone eats, an hour later. You then have to guess which of your forty decisions was the problem. Too much salt? Wrong heat? Wrong moment for the tomatoes?

With the friend tasting as you go, you get an opinion after every step. "Better than it was." "Worse than it was." You can correct the thing you did five seconds ago.

The cook is the actor, the part that picks the moves. The tasting friend is the critic, the part that judges how things are going.

The problem this fixes

The previous lesson had an honest weakness. REINFORCE plays a whole episode, sees the total, and then encourages or discourages every single action in it based on that one number.

A great episode probably contained three bad moves. They get encouraged. A terrible episode probably contained a brilliant save. It gets discouraged.

Over thousands of episodes it averages out. But "thousands of episodes" is the price, and the learning curve is jumpy the whole way.

The critic replaces the end-of-episode verdict with an immediate opinion.

What the critic actually says

The critic answers one question: "from here, how well is this likely to go?"

Not "what should I do". Only "how good is this situation".

Then the actor gets a far sharper signal:

   what the critic expected from this situation:   50
   what actually happened after the action:        62
   -------------------------------------------------
   the action did better than expected by:        +12  ->  do more of it

That difference has a name: the advantage. How much better this action turned out than the situation deserved.

Notice how much cleaner that is. It does not matter whether 62 is a big number or a small one. What matters is whether it beat what was expected from that spot.

Why "better than expected" is the right question

Suppose an agent is in a hopeless position. Every action leads to a score of about 10, and the best possible is 11.

REINFORCE sees a return of 11, a positive number, and encourages what was done.

Actor-critic sees that the critic expected 10, the outcome was 11, and the advantage is +1. Small, honest, correctly proportioned.

Now put the agent in a wonderful position where everything gives about 500. It takes a poor action and gets 480. REINFORCE sees a large positive number and encourages a bad move. Actor-critic sees an advantage of minus 20 and pushes it down.

Comparing against expectation, not against zero, is the whole idea.

The catch nobody mentions first

At the start of training, the critic is useless. It has seen nothing and its opinions are noise.

So the actor is learning from a judge that is itself learning. Two learners, each depending on the other, each moving under the other's feet.

This is genuinely hard, and it is worth saying plainly. When actor-critic training fails, it is usually this. The critic falls behind, so the actor gets bad advice, so it visits strange situations, which the critic has never seen, so its advice gets worse. The whole thing spirals.

That fragility is the reason PPO exists. It limits how far the actor may move in one step. That limit is what keeps the pair together.

The shape of it

                  the situation
                   /         \
                  /           \
            [ actor ]      [ critic ]
                |               |
       chances over actions   one number:
                |             "how good is
             pick one          it here?"
                |               |
                v               v
            do it  ---->  compare what happened
                          with what was expected
                                |
                                v
                   actor: do more of what beat expectation
                   critic: get better at expecting

Where you have already seen it

Almost everywhere modern reinforcement learning is used, because almost every modern method is an actor-critic method.

  • Robot control in simulation and on real hardware.
  • Game-playing agents in complex real-time games.
  • Language model tuning. The version of RLHF used on the assistants you have talked to is an actor-critic method with an extra safety strap.

Remember this

  • The actor picks actions. The critic judges how good the situation is.
  • The actor learns from the advantage: how much better the outcome was than expected.
  • Two learners depending on each other is powerful and fragile, in that order.

What to learn next

  • PPO — the actor-critic method that fixed the instability well enough to be a default.
  • Policy gradients — the estimator the critic is reducing the variance of.
  • RLHF — actor-critic applied to language models.

Developer — Code and libraries.

Setup

bash
pip install gymnasium torch numpy

CPU only. This runs in about twenty-five seconds — roughly half the time REINFORCE took to reach the same score.

Actor and critic, both from scratch

actor_critic.py
import gymnasium as gym
import numpy as np
import torch
import torch.nn as nn

torch.manual_seed(0); np.random.seed(0)
torch.set_num_threads(1)

env = gym.make("CartPole-v1")

actor  = nn.Sequential(nn.Linear(4, 64), nn.ReLU(), nn.Linear(64, 2))   # what to do
critic = nn.Sequential(nn.Linear(4, 64), nn.ReLU(), nn.Linear(64, 1))   # how good this state is
opt = torch.optim.Adam(list(actor.parameters()) + list(critic.parameters()), lr=0.01)
GAMMA = 0.99
scores = []

for episode in range(600):
    state, _ = env.reset(seed=episode)
    states, actions, rewards, done = [], [], [], False

    while not done:
        logits = actor(torch.tensor(state))
        dist = torch.distributions.Categorical(logits=logits)
        action = dist.sample()
        states.append(state); actions.append(int(action))
        state, reward, terminated, truncated, _ = env.step(int(action))
        rewards.append(reward)
        done = terminated or truncated

    returns, running = [], 0.0
    for r in reversed(rewards):
        running = r + GAMMA * running
        returns.insert(0, running)

    s = torch.tensor(np.array(states))
    a = torch.tensor(actions)
    g = torch.tensor(returns, dtype=torch.float32)

    values = critic(s).squeeze(1)
    advantage = g - values                     # was the outcome better than the critic expected?

    dist = torch.distributions.Categorical(logits=actor(s))
    actor_loss  = -(dist.log_prob(a) * advantage.detach()).mean()
    critic_loss = nn.functional.mse_loss(values, g)
    entropy     = dist.entropy().mean()        # a nudge to keep some randomness alive

    loss = actor_loss + 0.5 * critic_loss - 0.01 * entropy
    opt.zero_grad(); loss.backward(); opt.step()

    scores.append(sum(rewards))
    if (episode + 1) % 100 == 0:
        print(f"episode {episode + 1:3d}  avg of last 100: {np.mean(scores[-100:]):6.1f}"
              f"   critic's guess at the start state: {critic(torch.tensor(states[0])).item():6.1f}")

env.close()
Output
episode 100  avg of last 100:   15.3   critic's guess at the start state:   15.0
episode 200  avg of last 100:  271.5   critic's guess at the start state:   73.3
episode 300  avg of last 100:  353.5   critic's guess at the start state:   67.2
episode 400  avg of last 100:  236.8   critic's guess at the start state:   84.9
episode 500  avg of last 100:  443.5   critic's guess at the start state:   92.7
episode 600  avg of last 100:  500.0   critic's guess at the start state:  104.6

Exact numbers vary by PyTorch build and machine even with seeds fixed. The trend is what matters.

Watch the critic learn its own job

The second column is the interesting one. The critic's estimate of the starting position climbs: 15, then 73, then 92, then 105.

Early on it says 15 because the pole really does fall after about 15 steps. It is not wrong; the agent is bad. As the actor improves, the same starting position becomes genuinely more valuable, and the critic tracks it.

There is a ceiling. With GAMMA = 0.99 and a 500-step limit, the largest possible discounted return is about 99.3. The critic settles a little above that at 104.6, which is honest approximation error rather than a bug — the critic is a small network fitting noisy targets, and it overshoots slightly. If you saw 400 there, that would be a bug.

Notice also the dip at episode 400 (236.8) before the recovery. Actor-critic on CartPole is more stable than DQN and it is not monotone.

Line by line

advantage = g - values — the whole method. g is what actually happened. values is what the critic expected. The difference is the advantage.

advantage.detach() — this is the line people get wrong, and it is silent when wrong. Without detach, the actor's loss sends gradients into the critic, and the optimiser discovers it can reduce the actor loss by making the critic's predictions worse. The critic is corrupted, the advantage becomes meaningless, and training limps along without any error message. Detach the advantage before it touches the actor.

One optimiser over both parameter lists — convenient here. In larger setups the critic usually wants a higher learning rate than the actor, so two optimisers is the more common arrangement.

0.5 * critic_loss — the critic's loss is on the scale of returns (up to about 100) while the actor's is on the scale of log-probabilities (around 1). Without a weight, the critic's gradients drown the actor's. That 0.5 is a real tuning knob, not decoration.

- 0.01 * entropy — the entropy bonus. Entropy measures how undecided the policy is. Subtracting it from the loss rewards staying undecided. Without it, a policy that gets an early lead can collapse to always-push-left, and a deterministic policy has nothing left to explore with. Set the coefficient to 0.0 and you will see this happen on some seeds.

nn.functional.mse_loss(values, g) — the critic is doing plain supervised regression. Its inputs are states, its targets are the observed returns. Nothing about it is special; the difficulty is that its targets keep changing as the actor improves.

The version used in practice

This script uses full-episode returns for the critic's target, which makes it a Monte-Carlo actor-critic. Real implementations bootstrap instead:

python
# one-step target: reward now, plus the critic's opinion of where we landed
target = reward + GAMMA * critic(next_state) * (1 - terminated)

That reduces variance sharply and adds bias. Generalised advantage estimation (GAE) blends the two with a parameter lam: lam = 1.0 gives the Monte-Carlo version above, lam = 0.0 gives the one-step version, and lam = 0.95 is the near-universal default. You will meet GAE fully in the PPO lesson.

Common mistakes

Forgetting .detach() on the advantage. Described above. This is the most common actor-critic bug and it does not raise anything.

Sharing a network body between actor and critic without balancing the losses. A shared trunk saves compute and couples the two objectives. If the critic's gradients dominate, the shared features become good for value prediction and useless for choosing actions.

No entropy bonus. Premature collapse to a deterministic policy. Symptom: the score climbs, then freezes exactly, and the policy's entropy is near zero.

Using an advantage of zero mean but wild scale. Normalising advantages per batch ((adv - adv.mean()) / (adv.std() + 1e-8)) is standard and makes the actor's learning rate far less sensitive.

Bootstrapping through a time-limit truncation. When CartPole stops at 500 steps the pole is still up. Treating that as a terminal state teaches the critic that surviving 500 steps is worth nothing after step 500.

Try it yourself

Set the entropy coefficient to 0.0 and run five seeds. On at least one, the policy will collapse early and the score will freeze low. Then set it to 0.1 and watch the opposite: the agent stays deliberately random and never commits, so the score plateaus in the middle. The right value is a balance, and finding it by hand is the fastest way to understand what entropy is buying you.

What to learn next

  • PPO — the actor-critic method that fixed the instability well enough to be a default.
  • Policy gradients — the estimator the critic is reducing the variance of.
  • RLHF — actor-critic applied to language models.

Researcher — Mathematics and papers.

The construction

Start from the policy gradient with a state-dependent baseline:

$$ \nabla_\theta J(\theta) = \mathbb{E}{s \sim d^\pi, a \sim \pi\theta}\left[ \nabla_\theta \log \pi_\theta(a \mid s) \, A^{\pi}(s,a) \right], \qquad A^\pi(s,a) = Q^\pi(s,a) - V^\pi(s) $$

$A^\pi$ is the advantage function. Since $\mathbb{E}_{a\sim\pi}[A^\pi(s,a)] = 0$, it has lower magnitude and lower variance than $Q^\pi$ while giving the same expected gradient.

Neither $Q^\pi$ nor $V^\pi$ is known, so the critic $V_\phi$ estimates $V^\pi$ and the advantage is estimated from data. Common estimators, in increasing bias and decreasing variance:

$$ \hat{A}_t^{(\infty)} = G_t - V_\phi(s_t) \qquad \hat{A}t^{(n)} = \sum{k=0}^{n-1}\gamma^k r_{t+k+1} + \gamma^n V_\phi(s_{t+n}) - V_\phi(s_t) \qquad \hat{A}_t^{(1)} = \delta_t = r_{t+1} + \gamma V_\phi(s_{t+1}) - V_\phi(s_t) $$

Generalised advantage estimation

Schulman et al. (2016) take an exponentially weighted average over all $n$:

$$ \hat{A}t^{\text{GAE}(\gamma,\lambda)} = \sum{l=0}^{\infty} (\gamma\lambda)^l \delta_{t+l}, \qquad \delta_t = r_{t+1} + \gamma V_\phi(s_{t+1}) - V_\phi(s_t) $$

$\lambda = 1$ recovers the Monte-Carlo advantage (unbiased, high variance); $\lambda = 0$ gives the one-step TD advantage (biased by the critic's error, low variance). $\lambda \in [0.92, 0.98]$ is standard. Computed backwards in one pass:

python
gae = 0.0
for t in reversed(range(T)):
    delta = r[t] + gamma * V[t+1] * (1 - done[t]) - V[t]
    gae = delta + gamma * lam * (1 - done[t]) * gae
    adv[t] = gae

Note that the bias introduced by $\lambda < 1$ is proportional to the critic's error, so GAE and critic quality are not independent knobs.

Compatible function approximation

Sutton et al. (2000) prove that substituting an approximation $f_w(s,a)$ for $Q^\pi$ leaves the gradient exact if two conditions hold:

  1. $\nabla_w f_w(s,a) = \nabla_\theta \log \pi_\theta(a \mid s)$ — the critic's features are the policy's score function.
  2. $w$ minimises $\mathbb{E}[(Q^\pi(s,a) - f_w(s,a))^2]$.

This is the compatible function approximation theorem, and it is why actor-critic is more than a heuristic. In practice condition (1) is violated by every deep implementation, so the gradient is biased. The theorem tells you what you are giving up, which is more useful than a guarantee nobody satisfies.

The algorithm family

MethodDistinguishing feature
A3C (Mnih et al., 2016)Asynchronous parallel workers, lock-free Hogwild updates
A2CSynchronous batched A3C; equal or better, and simpler
TRPO (Schulman et al., 2015)Hard KL trust region via conjugate gradient
PPO (Schulman et al., 2017)Clipped surrogate, first-order, data reuse
DDPG / TD3Deterministic actor, off-policy, continuous actions
SAC (Haarnoja et al., 2018)Maximum-entropy objective, off-policy, twin critics

A3C's asynchrony was originally credited with decorrelating updates. Later analysis showed the synchronous A2C matches or beats it, and the parallelism, not the asynchrony, was the useful part.

SAC deserves separate attention. It maximises $\mathbb{E}[\sum_t \gamma^t (r_t + \alpha \mathcal{H}(\pi(\cdot \mid s_t)))]$ — return plus policy entropy — making the entropy bonus part of the objective rather than a regulariser bolted on. With automatic tuning of $\alpha$ against a target entropy, it is the most reliable off-policy continuous-control method available and typically the right default for robotics.

Why the coupled system is unstable

Three interacting failure modes, worth naming precisely:

  1. Non-stationary regression. $V_\phi$ chases $V^{\pi_\theta}$ while $\theta$ moves. The critic's target distribution shifts with every actor update.
  2. Distribution shift. The critic is accurate only where the current policy visits. A large actor update moves the visitation distribution into regions where the critic extrapolates, and its errors then drive the actor further out.
  3. Error amplification. A biased advantage produces a biased gradient, which changes the policy, which changes the states the critic sees. Nothing in the loop damps this.

Every practical stabiliser attacks one of these: trust regions and clipping bound (2); target networks slow (1); advantage normalisation bounds the scale of (3).

References

  • Barto, Sutton and Anderson (1983), Neuronlike adaptive elements that can solve difficult learning control problems — the original actor-critic, on the cart-pole task used above.
  • Konda and Tsitsiklis (2000), Actor-Critic Algorithms — the two-timescale convergence analysis.
  • Mnih et al. (2016), Asynchronous Methods for Deep Reinforcement Learning — arxiv.org/abs/1602.01783.
  • Schulman et al. (2016), High-Dimensional Continuous Control Using Generalized Advantage Estimation — arxiv.org/abs/1506.02438.
  • Haarnoja et al. (2018), Soft Actor-Critic — arxiv.org/abs/1801.01290.

What to learn next

  • PPO — the actor-critic method that fixed the instability well enough to be a default.
  • Policy gradients — the estimator the critic is reducing the variance of.
  • RLHF — actor-critic applied to language models.

What to learn next

These follow on from what you just read.

  • Reinforcement Learning

    PPO

    PPO is the default reinforcement learning algorithm today — an actor-critic method with one small rule that stops the policy changing too much in a single update.

  • Reinforcement Learning

    Reward shaping

    Adding small rewards along the way helps an agent that would otherwise never stumble onto its goal — and done carelessly it teaches the agent to cheat instead.

  • Reinforcement Learning

    RLHF — learning from human feedback

    RLHF turns "which of these two answers do you prefer?" into a reward the model can be trained against — it is how a raw language model became a usable assistant.