Deep Learning Section 036
Reinforcement Learning
How an AI learns by trying things and getting rewarded — the method behind game-playing AI and RLHF.
13 of 13 lessons published Three reading levels on every lesson
Lessons in order
Work top to bottom. Each lesson assumes the one above it.
- What is reinforcement learning?
- Agents, environments and rewards
- Markov decision processes
- Exploration vs exploitation
- Multi-armed bandits
- Q-learning
- Deep Q-networks (DQN)
- Policy gradients
- Actor-critic methods
- PPO
- Reward shaping
- RLHF — learning from human feedback
- Gymnasium and training environments