Agents, environments and rewards
Every reinforcement learning problem is one agent, one environment and one number called reward, connected in a loop that never changes shape.
- 14 min read
- 3 reading levels
- Updated
Read these first
On this page 10
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
An agent chooses, an environment reacts, and a reward says how that went.
Every reinforcement learning system ever built is those three things in a loop.
The analogy you have already lived
Picture yourself learning to make chai for the first time. You are the one deciding: more milk, less sugar, stove up, stove down. The kitchen is everything you do not control — how hot the gas burns, how old the milk is, when it boils over.
The reward is the face of whoever tastes it.
You are the agent. The kitchen is the environment. The face is the reward. Nothing else is needed to describe what is happening.
Why the split matters
Beginners often draw the line in the wrong place, and it ruins everything downstream.
The agent is only the decision-maker. In a self-driving car, the agent is the software choosing steering and braking. The car's engine, tyres and weight are the environment — the agent does not control them, it acts through them.
The environment includes the agent's own body. A robot's arm is part of the environment, because the arm can be broken, slow, or heavier than expected. The agent sends a command and finds out what happened.
Get this line wrong and you will spend weeks debugging an agent that is being asked to control something it cannot see.
The four things that pass between them
Every single step, the same four things move.
agent ---------------- action ----------------> environment
<--------------- observation ------------
<--------------- reward -----------------
<--------------- "is it over?" ----------Action — what the agent chose. Turn left. Buy. Press A.
Observation — what the agent gets to see afterwards. Note the word: it is not "the state of the world". It is only the part of the world the agent can see. A poker agent sees its own cards, never yours.
Reward — one number. That is the whole of it. Not a paragraph of feedback, not a label. One number per step.
Is it over? — a yes or no. One complete run from start to finish is called an episode. One game of chess is an episode. One delivery run is an episode.
Reward is smaller than you think
This trips up almost everyone. The reward is a single number, and it arrives on a schedule you chose.
In chess, most people give no reward at all for every move. One point arrives at the end for a win, and a penalty for a loss. Sixty moves, and only the last one carries any information.
That kind of reward is called sparse — it stays silent almost all the time. Sparse rewards are honest but very hard to learn from.
The temptation is to hand out small rewards along the way to help. That temptation causes more disasters than any other habit in this field, and reward shaping is the lesson about why.
Return, not reward
The agent is not trying to get a big reward on this step. It is trying to get a big total over the whole episode. That total has a name: the return.
The difference is everything. Sacrificing your queen looks terrible this move and wins the game four moves later.
reward this step: 0 0 0 0 +1
return from the start: the whole row added upMost agents also count later rewards slightly less than sooner ones. A sweet today beats the promise of a sweet next year. That preference is called discounting, and it also stops the total running away to infinity in a task that never ends.
Where you have already seen this loop
- A thermostat that learns your flat's heating. Action: heater on or off. Observation: the temperature. Reward: comfortable, and a small penalty for the electricity bill.
- An ad system. Action: which ad to show. Observation: who is looking. Reward: whether they clicked.
- A warehouse robot. Action: motor commands. Observation: camera and wheel sensors. Reward: parcels delivered, penalty for collisions.
- A chatbot being tuned. Action: the next word. Observation: the conversation so far. Reward: how a human rated the finished reply.
What is honestly hard here
Choosing what the agent gets to observe is a design decision, and it is often the whole problem. Give a robot too little and it cannot possibly succeed. Give it too much and it drowns.
Writing the reward is harder than writing the agent. You have to express what you actually want as a single number. That number gets read thousands of times a second, by something that will exploit any gap you leave. Teams routinely spend more time on the reward than on the algorithm.
Remember this
- The agent decides, the environment reacts, the reward scores it. Nothing else.
- The agent sees an observation, not the world. These are different, and the gap matters.
- Agents maximise the return — the whole episode's total — never one step's reward.
What to learn next
- Markov decision processes — the mathematics this loop turns into.
- Gymnasium and training environments — spaces, wrappers, and writing your own environment properly.
- Reward shaping — why the reward is the hardest thing on this page.
Developer — Code and libraries.
Setup
pip install gymnasium numpyGymnasium is the standard Python interface for reinforcement learning environments. It is the maintained successor to OpenAI Gym, which is no longer updated. Install size is a few megabytes, and everything here runs on a CPU.
The loop, in real code
Gymnasium gives every environment the same two methods, and the whole field is built on that agreement.
import gymnasium as gym
env = gym.make("FrozenLake-v1", map_name="4x4", is_slippery=False)
print("what the agent can see :", env.observation_space) # 16 squares, numbered 0..15
print("what the agent can do :", env.action_space) # 4 moves
print()
obs, info = env.reset(seed=42)
env.action_space.seed(42)
MOVES = ["left", "down", "right", "up"]
total = 0.0
for t in range(6):
action = env.action_space.sample() # a random flail, on purpose
obs, reward, terminated, truncated, info = env.step(action)
total += reward
print(f"step {t}: moved {MOVES[action]:<5} -> square {obs:2d} reward {reward} over={terminated}")
if terminated or truncated:
print(" episode ended, resetting")
obs, info = env.reset()
break
print()
print("reward collected:", total)
env.close()what the agent can see : Discrete(16) what the agent can do : Discrete(4) step 0: moved left -> square 0 reward 0 over=False step 1: moved up -> square 0 reward 0 over=False step 2: moved right -> square 1 reward 0 over=False step 3: moved down -> square 5 reward 0 over=True episode ended, resetting reward collected: 0.0
The random agent walked into a hole on step 3 and the episode ended with nothing. That is the honest baseline every agent has to beat.
The five values step returns
env.step(action) hands back exactly five things, always in this order.
obs — the observation. Here it is the number of the square, 0 to 15. Discrete(16) means "one integer from a set of sixteen".
reward — a float. FrozenLake pays 1.0 for reaching the goal and 0.0 for everything else, including falling in a hole. Notice that the hole is not punished. Falling in is bad only because it ends the episode without the point.
terminated — the episode genuinely finished. The goal was reached, or the agent died.
truncated — the episode was cut off from outside, usually by a step limit. Keeping this separate from terminated is not pedantry. A truncated episode still has future value that you should bootstrap from; a terminated one does not. Merging them is a real and common bug.
info — a dictionary of extras for debugging. Never train on it.
env.reset(seed=42) seeds the environment, and env.action_space.seed(42) seeds the random action sampler separately. Both are needed for a reproducible run, and forgetting the second one is why your "seeded" experiment still moves.
Writing your own environment, with no library at all
There is nothing magic in gym.make. An environment is any object with reset and step. Here is one, and two agents to run inside it.
import random
class RoadCrossing:
"""The environment. It owns the truth. The agent only sees what step() hands back."""
WAIT, WALK = 0, 1
def __init__(self, seed=0):
self.rng = random.Random(seed)
def reset(self):
self.light = "red"
return self.light # the observation
def step(self, action):
if action == self.WALK:
reward = 5.0 if self.light == "green" else -20.0
return self.light, reward, True # observation, reward, episode over
if self.light == "red" and self.rng.random() < 0.4:
self.light = "green"
return self.light, -0.5, False # standing still costs a little
def careful(observation):
return RoadCrossing.WALK if observation == "green" else RoadCrossing.WAIT
def reckless(observation):
return RoadCrossing.WALK
def run(agent, name, seed=0):
env = RoadCrossing(seed=seed)
obs, total, t = env.reset(), 0.0, 0
print(f"--- {name} ---")
while True:
action = agent(obs)
obs, reward, done = env.step(action)
total += reward
print(f"t={t} did {'WALK' if action else 'WAIT'} light now {obs:<5} reward {reward:+.1f}")
t += 1
if done:
break
print(f"return for this episode: {total:+.1f}\n")
run(careful, "careful agent")
run(reckless, "reckless agent")--- careful agent --- t=0 did WAIT light now red reward -0.5 t=1 did WAIT light now red reward -0.5 t=2 did WAIT light now red reward -0.5 t=3 did WAIT light now green reward -0.5 t=4 did WALK light now green reward +5.0 return for this episode: +3.0 --- reckless agent --- t=0 did WALK light now red reward -20.0 return for this episode: -20.0
What that output is actually teaching
The careful agent earned four penalties of half a point and still finished ahead. Return, not immediate reward. Every one of those waits looked like a loss at the time.
Now look at the numbers you would have to choose as a designer. Waiting costs -0.5 per step and crossing on green pays +5.0. Raise the waiting cost to -3.0 and the careful agent's return drops below zero after two waits — at which point walking into traffic starts to look attractive to it.
You did not change the algorithm. You changed three numbers, and the agent's whole personality changed with them. Reward design is engineering, not decoration.
Notice also that the agent never sees self.rng. It has an observation, not the state. It cannot know when the light will change, only that it has not yet.
Common mistakes
Treating truncated as terminated. When a step limit cuts an episode short, the future is not worthless — the pole was still balanced. Bootstrap from the final observation for truncated, and use zero for terminated. This mistake silently caps how good your agent can get.
Forgetting env.reset() before the first step. Gymnasium raises a clear error for this, which is a kindness. Your own environments should do the same.
Rewarding the agent for something it cannot observe. If the reward depends on wind speed and the agent never sees wind speed, you have handed it a noise generator and asked it to find a pattern.
Putting the goal in info and the score in reward, then training on info. info is for humans. If a value matters to learning, it belongs in the observation or the reward.
Try it yourself
In RoadCrossing, change the waiting penalty from -0.5 to -3.0 and re-run. Predict what the careful agent's return becomes before you look. Then find the exact penalty at which waiting for green stops being worth it, and you will have derived a real reward-design boundary by hand.
What to learn next
- Markov decision processes — the mathematics this loop turns into.
- Gymnasium and training environments — spaces, wrappers, and writing your own environment properly.
- Reward shaping — why the reward is the hardest thing on this page.
Researcher — Mathematics and papers.
The agent-environment interface, formally
At each discrete time step $t$ the environment is in state $s_t \in \mathcal{S}$. The agent samples $a_t \sim \pi(\cdot \mid s_t)$. The environment produces $s_{t+1} \sim P(\cdot \mid s_t, a_t)$ and a scalar $r_{t+1} \sim R(\cdot \mid s_t, a_t)$.
The return from time $t$ is $$ G_t = \sum_{k=0}^{T-t-1} \gamma^{k} r_{t+k+1} $$ with $T = \infty$ for continuing tasks (requiring $\gamma < 1$ for convergence when rewards are bounded) and $T$ finite for episodic ones.
Two boundaries are worth stating precisely.
The agent-environment boundary is a modelling choice, not a physical one. Sutton and Barto's rule: anything the agent cannot change arbitrarily by choice is part of the environment. A robot's motors are environment; the controller's memory is agent. Reward computation always sits on the environment side, because an agent able to edit its own reward function has a degenerate optimal policy.
The observation-state distinction is the difference between an MDP and a POMDP. When $o_t = s_t$ the process is fully observed. Otherwise you have a partially observable MDP, $(\mathcal{S}, \mathcal{A}, P, R, \Omega, O, \gamma)$, with observation set $\Omega$ and observation function $O(o \mid s', a)$. In a POMDP, no memoryless policy $\pi(a \mid o)$ is optimal in general — the sufficient statistic is the belief state $b_t(s) = \Pr(s_t = s \mid o_{1:t}, a_{1:t-1})$, and the belief-MDP over $\Delta(\mathcal{S})$ is continuous even when $\mathcal{S}$ is finite. Solving POMDPs exactly is PSPACE-complete for finite horizons (Papadimitriou and Tsitsiklis, 1987).
In practice this is why deep RL agents stack frames or use recurrent policies. A single Atari frame does not reveal velocity; four stacked frames approximately do.
Termination versus truncation
This has a precise consequence for the bootstrap target. For a transition ending at $s_{t+1}$:
$$ y_t = \begin{cases} r_{t+1} & \text{terminated (absorbing state, } V(s_{t+1}) = 0) \ r_{t+1} + \gamma V(s_{t+1}) & \text{truncated (time limit, value is not zero)} \ r_{t+1} + \gamma V(s_{t+1}) & \text{ongoing} \end{cases} $$
Treating a time-limit truncation as a termination injects a systematic negative bias into $V$ near the horizon. Pardo et al. (2018), Time Limits in Reinforcement Learning, quantifies this and shows measurable performance loss on standard benchmarks. Gymnasium's five-tuple API exists precisely to make the distinction unavoidable; the older four-tuple done flag conflated them.
Reward specification
The reward function is part of the problem statement, not the solution. Three properties are worth checking before training anything.
Scale. Value targets inherit the reward scale. Rewards in the thousands produce large TD errors and unstable gradients. Clipping to $[-1, 1]$ (as DQN did) stabilises training but changes the objective: it makes the agent indifferent between a reward of 10 and a reward of 1000.
Sparsity. With reward only at a goal state, the probability that undirected exploration reaches it decays exponentially in the horizon. This is the motivation for reward shaping, curiosity bonuses (Pathak et al., 2017), and hindsight relabelling (Andrychowicz et al., 2017).
Alignment. The reward defines the optimal policy exactly. Any gap between the reward and the designer's intent is not a bug the agent will forgive. Krakovna et al. maintain a catalogue of specification-gaming examples that is worth an hour of anyone's time.
Discounting is not only a mathematical convenience
Three separate readings of $\gamma$, all correct and sometimes in tension:
- Convergence. It bounds $|G_t| \le r_{\max}/(1-\gamma)$ for bounded rewards.
- Preference. It encodes genuine impatience, the analogue of an interest rate.
- Termination probability. A $\gamma$-discounted infinite-horizon problem is equivalent to an undiscounted problem with a $(1-\gamma)$ chance of ending at each step.
The effective horizon is $\approx 1/(1-\gamma)$: 100 steps at $\gamma = 0.99$, 1000 at $\gamma = 0.999$. Raising $\gamma$ increases the variance of return estimates and typically slows learning, even when the true task is undiscounted. Jiang et al. (2015), The Dependence of Effective Planning Horizon on Model Accuracy, shows that a smaller-than-true $\gamma$ can be the better choice under model error — a real bias-variance trade-off, not a hack.
References
- Sutton and Barto (2018), Reinforcement Learning: An Introduction, chapter 3 — the canonical treatment of this interface.
- Kaelbling, Littman and Cassandra (1998), Planning and acting in partially observable stochastic domains.
- Pardo et al. (2018), Time Limits in Reinforcement Learning — arxiv.org/abs/1712.00378.
- Towers et al. (2024), Gymnasium: A Standard Interface for Reinforcement Learning Environments — arxiv.org/abs/2407.17032.
What to learn next
- Markov decision processes — the mathematics this loop turns into.
- Gymnasium and training environments — spaces, wrappers, and writing your own environment properly.
- Reward shaping — why the reward is the hardest thing on this page.