Reinforcement Learning

What is reinforcement learning?

Reinforcement learning is how a computer learns by trying things, seeing what happens, and doing more of whatever worked.

On this page 9
  1. The short answer
  2. The analogy you have already lived
  3. Why this had to be invented
  4. How it works
  5. Where you have already seen it
  6. Three words you will keep meeting
  7. What is honestly hard here
  8. Remember this
  9. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

Reinforcement learning is learning by trying things and living with the consequences.

Nobody hands the learner the right answer. It acts, something happens, and it adjusts.

The analogy you have already lived

Think about the day you learned to ride a bicycle. Nobody gave you a list of correct handlebar angles. You pushed off, wobbled, and scraped a knee. You pushed off again and stayed up for three seconds. Then five. Then you were riding.

Every wobble was information. Every metre you stayed upright told you to keep doing whatever you were doing.

That is reinforcement learning, and you already know how it feels from the inside.

A dog learning to give its paw works the same way. It tries something, a treat appears or it does not, and the behaviour that earned treats happens more often.

Why this had to be invented

Most machine learning needs a teacher with an answer key. You show a photo, and the label says "cat". You show a house, and the label says "eighty lakh rupees". The lesson on supervised learning covers that setup.

That falls apart for a whole class of problems.

Nobody knows the right answer for every moment. What is the correct steering angle when a bike is leaning eleven degrees left at nine kilometres an hour? No human can write that down for every situation.

The reward arrives much later than the decision. You lose a chess game on move sixty. The bad move might have been move twelve. No answer key tells you which one.

Your choices change what you see next. In photo labelling, the next photo does not depend on your answer. On a bicycle, turning left decides what happens next. The learner shapes its own future data.

Reinforcement learning was built for exactly these situations.

How it works

The whole field is one loop, repeated millions of times.

        +----------------------------------------------+
        |                                              |
        v                                              |
   [ the learner ] --- picks an action ---> [ the world ]
        ^                                              |
        |                                              |
        +--- new situation, plus a score for that ------+
             action (a treat, or nothing, or a scolding)

Round and round. Nobody says "that was the correct move". The world only says "here is what happened, and here is your score".

Over enough rounds, the learner works out which actions tend to lead to good scores. That is the entire idea.

This is the part that surprises people. The learner is never told the right action. It is only ever told how well things went. Working out which of its hundred past moves deserves the credit is the hard part, and it has a name: credit assignment.

Where you have already seen it

  • A robot vacuum that gets better at covering your flat without falling down the stairs.
  • Game AI that beat the world champion at Go, having learned by playing against itself.
  • Data centre cooling, where a system learns fan and pump settings that cut the electricity bill.
  • ChatGPT and similar assistants. Human testers said which of two replies they preferred, and the model was tuned to produce more of the preferred kind. That method is RLHF, and it gets its own lesson.

Three words you will keep meeting

  • The agent is the learner. The thing making choices.
  • The environment is everything else. The world the agent acts inside.
  • The reward is the score the environment hands back after each action.

The next lesson takes these three apart properly.

What is honestly hard here

Reinforcement learning is the least reliable branch of AI to actually get working. This is not a beginner's difficulty that fades with practice. Three real problems:

It needs a staggering number of tries. A child learns a bicycle in an afternoon. A reinforcement learning agent might need a million attempts. This is why most of it happens inside simulations, where a million attempts cost nothing.

It gets worse before it gets better, and sometimes it gets worse and stays there. Training curves in this field climb and then fall off a cliff for no visible reason. Everyone who works here has watched it happen.

It does exactly what you rewarded, not what you meant. Reward a boat-racing agent for collecting points, and it may find a lagoon full of pickups. It spins in circles hitting them forever, and never finishes the race. That really happened. The lesson on reward shaping is mostly about this problem.

If those three sound off-putting, that is the correct reaction. It is also why the field is interesting.

Remember this

  • Reinforcement learning learns from consequences, not from an answer key.
  • Everything is one loop: act, see what happened, take the score, adjust.
  • The hard parts are waiting for delayed rewards and writing a score that means what you intended.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install numpy

That is the only dependency for this lesson. It runs in under a second on any laptop, and no GPU is involved.

The smallest agent that genuinely learns

Six rooms in a row. The agent starts in room 0. A sweet shop is in room 5. It can step left or step right. It has no map, no rules, and no idea the sweet shop exists.

It keeps a scorecard: one number for every combination of "which room I am in" and "which way I stepped". Every number starts at zero, meaning "no opinion yet".

first_agent.py
import numpy as np

# A corridor of 6 rooms. The agent starts at room 0. The sweet shop is room 5.
N_ROOMS, GOAL = 6, 5
rng = np.random.default_rng(0)

# The scorecard: how good each action looks from each room. 0 = "no idea yet".
scores = np.zeros((N_ROOMS, 2))          # column 0 = go left, column 1 = go right

def step(room, action):
    """The world's rule. The agent never gets to read this function."""
    nxt = max(0, room - 1) if action == 0 else min(GOAL, room + 1)
    return nxt, (1.0 if nxt == GOAL else 0.0)

steps_taken = []
for episode in range(30):
    room, steps = 0, 0
    while room != GOAL and steps < 100:
        # Early on, wander. Later, trust the scorecard. This is the whole trick.
        explore = rng.random() < max(0.05, 1.0 - episode / 15)
        action = rng.integers(2) if explore else int(np.argmax(scores[room]))

        nxt, reward = step(room, action)
        best_next = 0.0 if nxt == GOAL else scores[nxt].max()
        # Nudge this room's score toward "reward now, plus what comes after".
        scores[room, action] += 0.5 * (reward + 0.9 * best_next - scores[room, action])

        room, steps = nxt, steps + 1
    steps_taken.append(steps)

print("steps per attempt, first ten :", steps_taken[:10])
print("steps per attempt, last  ten :", steps_taken[-10:])
print("fewest possible steps        :", GOAL)
print()
print("scorecard (rows = rooms, columns = left / right):")
print(np.round(scores, 2))
Output
steps per attempt, first ten : [26, 20, 15, 14, 29, 10, 8, 7, 7, 5]
steps per attempt, last  ten : [5, 5, 5, 7, 5, 5, 5, 5, 5, 5]
fewest possible steps        : 5

scorecard (rows = rooms, columns = left / right):
[[0.14 0.66]
 [0.07 0.73]
 [0.57 0.81]
 [0.6  0.9 ]
 [0.33 1.  ]
 [0.   0.  ]]

Twenty-six steps on the first attempt, five on the last. Five is the theoretical best. The agent found it without ever being told which direction the shop was in.

Read the scorecard — it tells the story

Look down the right-hand column: 0.66, 0.73, 0.81, 0.90, 1.00. The scores climb as you get nearer the shop. The agent built a gradient out of nothing but a single reward at the end.

Room 5 stays at zero for both actions. It is the shop. The episode ends there, so nothing is ever recorded from it.

The left column is not all zero either. Room 2 scores 0.57 for going left. Stepping left is not fatal, and the agent worked out that you can recover. It only prefers right.

Line by line, the parts that matter

explore = rng.random() < max(0.05, 1.0 - episode / 15) — on episode 0 this is true almost always, so the agent flails at random. By episode 15 it is true five percent of the time. That trade-off has a name and a whole lesson: exploration vs exploitation.

reward + 0.9 * best_next — the target is not only the reward now. It is the reward now plus a discounted look at how good the next room is. The 0.9 is a discount factor: a reward one step later counts for ninety percent of a reward right now. Without this term, only room 4 would ever learn anything, because only room 4 ever touches the reward.

scores[room, action] += 0.5 * (target - scores[room, action]) — move the old number halfway toward the new estimate rather than replacing it. The 0.5 is a learning rate, the same idea you met in gradient descent.

steps < 100 — a safety valve. An agent that never finds the shop would loop forever. Every serious reinforcement learning setup has one of these.

You have now seen Q-learning without the name attached. The Q-learning lesson makes it formal.

Common mistakes

Expecting it to work without exploration. Delete the explore line and force action = np.argmax(scores[room]). Every score is zero, argmax returns index 0, and the agent walks left into the wall forever. Zero reward, zero learning, and no error message. This is the most common way a first agent fails.

Reading the scorecard before it has data. After three episodes those numbers are noise. Reinforcement learning results mean nothing until the agent has visited each state many times.

Forgetting to treat the goal as an ending. Drop the 0.0 if nxt == GOAL and the scores grow without limit, because value keeps flowing back out of a state that should have stopped the episode.

Comparing single runs. Change default_rng(0) to default_rng(7) and the numbers move. One run tells you nothing in this field. Average over several seeds before believing anything.

Try it yourself

Set N_ROOMS to 30 and keep 30 episodes. Watch it fail. Then work out whether the fix is more episodes, a slower exploration decay, or a better reward — and check your guess by changing one thing at a time. That habit is worth more than any algorithm on this site.

What to learn next

Researcher — Mathematics and papers.

The formal setting

Reinforcement learning studies sequential decision making under uncertainty, usually formalised as a Markov decision process $(\mathcal{S}, \mathcal{A}, P, R, \gamma)$:

  • $\mathcal{S}$ — the state space.
  • $\mathcal{A}$ — the action space.
  • $P(s' \mid s, a)$ — the transition kernel, the probability of landing in $s'$ after taking $a$ in $s$.
  • $R(s, a)$ — the reward function, often a distribution with mean $r(s,a)$.
  • $\gamma \in [0, 1)$ — the discount factor.

The agent is a policy $\pi(a \mid s)$, a distribution over actions given a state. The quantity being maximised is the expected discounted return:

$$ J(\pi) = \mathbb{E}{\tau \sim \pi}\left[ \sum{t=0}^{\infty} \gamma^t R_{t+1} \right], \qquad G_t = \sum_{k=0}^{\infty} \gamma^k R_{t+k+1} $$

Here $\tau = (s_0, a_0, r_1, s_1, \dots)$ is a trajectory drawn by running $\pi$ in the environment, and $G_t$ is the return from time $t$ onwards.

$\gamma$ does more work than it appears to. It keeps the sum finite for continuing tasks, it encodes a preference for sooner rewards, and it acts as a variance-reduction knob: a smaller $\gamma$ shortens the effective horizon to roughly $1/(1-\gamma)$ and makes credit assignment easier at the cost of bias. See the MDP lesson for the Bellman machinery.

Why this is not supervised learning

The distinction is not cosmetic. Four properties break the standard statistical learning setup.

  1. No supervision, only evaluation. A label states what you should have done. A reward states how well what you did turned out. Evaluative feedback carries strictly less information per sample.
  2. The data distribution depends on the parameters. $\pi_\theta$ generates the states it then trains on. The i.i.d. assumption behind most generalisation bounds does not hold, and improving $\theta$ shifts the distribution under which $\theta$ is next evaluated.
  3. Delayed credit. The learning signal for an action at time $t$ may only arrive at $t + 1000$. The variance of Monte-Carlo return estimates grows with the horizon.
  4. Exploration is part of the objective. A supervised learner cannot improve its own dataset. An RL agent must, and doing so optimally trades against exploiting what it already knows.

Sample complexity, honestly

For a tabular MDP with $|\mathcal{S}|$ states and $|\mathcal{A}|$ actions, obtaining an $\varepsilon$-optimal policy given a generative model requires on the order of $\tilde{\Theta}!\left( \frac{|\mathcal{S}||\mathcal{A}|}{(1-\gamma)^3 \varepsilon^2} \right)$ samples (Azar, Munos and Kappen, 2013, with a matching lower bound in Azar et al., 2012). The $(1-\gamma)^{-3}$ term is the one that hurts: at $\gamma = 0.99$ that factor alone is $10^6$.

Deep RL has no comparable guarantee. Published Atari results routinely use 200 million frames per game — roughly 38 days of continuous human play — to reach human-level scores. Sample efficiency, not asymptotic performance, is the binding constraint in almost every real deployment.

The reward hypothesis, and its critics

Sutton's reward hypothesis states that every goal can be described as maximisation of expected cumulative scalar reward. It is a hypothesis, not a theorem, and it is contested. Abel et al. (2021), On the Expressivity of Markov Reward, construct tasks — certain orderings and multi-objective preferences — for which no Markov reward function induces the desired set of optimal policies. Worth reading before assuming any objective can be compressed into one number.

Landmarks

  • Samuel (1959), Some Studies in Machine Learning Using the Game of Checkers — temporal-difference learning before it had the name.
  • Sutton (1988), Learning to Predict by the Methods of Temporal Differences — TD($\lambda$).
  • Watkins and Dayan (1992), Q-learning — the convergence proof for off-policy control.
  • Tesauro (1994), TD-Gammon — a neural network reaching world-class backgammon through self-play.
  • Mnih et al. (2015), Human-level control through deep reinforcement learning, Nature 518:529-533 — DQN. The earlier workshop version is Playing Atari with Deep Reinforcement Learning (2013), arxiv.org/abs/1312.5602.
  • Silver et al. (2017), Mastering the game of Go without human knowledge — AlphaGo Zero.
  • Ouyang et al. (2022), Training language models to follow instructions with human feedback — InstructGPT, arxiv.org/abs/2203.02155.

The standard text is Sutton and Barto, Reinforcement Learning: An Introduction (2nd edition, 2018). It is free online and it is the best book in the field.

Where the field actually is

Two honest observations about the current state.

Reproducibility is poor. Henderson et al. (2018), Deep Reinforcement Learning that Matters, showed that identical algorithms run with different random seeds produce non-overlapping performance distributions on standard benchmarks. Report your seeds, report variance across at least five of them, and treat any single-run curve with suspicion.

The largest deployed use of RL today is not robotics or games. It is preference optimisation on language models. RLHF and its offline relatives put more reinforcement learning into daily human use than every robot arm combined.

What to learn next