Reinforcement learning
In one sentence Reinforcement learning trains an agent by trial and error — it acts, receives rewards or penalties, and gradually prefers what worked.
Updated
Reinforcement learning, or RL, trains a decision-maker through trial and error: try actions, collect rewards and penalties, and shift toward what earned more reward.
It is how you actually learned to ride a bicycle. Nobody handed you labelled examples of correct handlebar angles. You wobbled, fell, adjusted, stayed up a little longer, and the staying-up itself was the teacher. Actions that preceded success got repeated; actions that preceded scraped knees did not.
The formal loop has four pieces: an agent (the learner) observes a state (the situation), takes an action, and receives a reward. Then the loop repeats.
state → action → reward → new state → action → ...
maximise total reward over timeTwo things make RL genuinely hard. Rewards are delayed — the losing chess move was twenty moves ago, so which action deserves the blame? And the agent must balance exploration (trying new things) against exploitation (repeating what works). This is a different setting from supervised-learning, which is told the right answer for every single input.
Landmark results include AlphaGo and game-playing agents, robotics, and datacentre cooling control. For LLM readers, the load-bearing application is RLHF — reinforcement learning from human feedback — the technique that turns a raw text predictor into a helpful assistant.
Where to go next
- Full lesson: What is reinforcement learning?
- Related terms: rlhf, supervised-learning, agent, reward-model