Reinforcement Learning

Agents, environments and rewards

Every reinforcement learning problem is one agent, one environment and one number called reward, connected in a loop that never changes shape.

On this page 10
  1. The short answer
  2. The analogy you have already lived
  3. Why the split matters
  4. The four things that pass between them
  5. Reward is smaller than you think
  6. Return, not reward
  7. Where you have already seen this loop
  8. What is honestly hard here
  9. Remember this
  10. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

An agent chooses, an environment reacts, and a reward says how that went.

Every reinforcement learning system ever built is those three things in a loop.

The analogy you have already lived

Picture yourself learning to make chai for the first time. You are the one deciding: more milk, less sugar, stove up, stove down. The kitchen is everything you do not control — how hot the gas burns, how old the milk is, when it boils over.

The reward is the face of whoever tastes it.

You are the agent. The kitchen is the environment. The face is the reward. Nothing else is needed to describe what is happening.

Why the split matters

Beginners often draw the line in the wrong place, and it ruins everything downstream.

The agent is only the decision-maker. In a self-driving car, the agent is the software choosing steering and braking. The car's engine, tyres and weight are the environment — the agent does not control them, it acts through them.

The environment includes the agent's own body. A robot's arm is part of the environment, because the arm can be broken, slow, or heavier than expected. The agent sends a command and finds out what happened.

Get this line wrong and you will spend weeks debugging an agent that is being asked to control something it cannot see.

The four things that pass between them

Every single step, the same four things move.

   agent  ---------------- action ---------------->  environment
          <--------------- observation ------------
          <--------------- reward -----------------
          <--------------- "is it over?" ----------

Action — what the agent chose. Turn left. Buy. Press A.

Observation — what the agent gets to see afterwards. Note the word: it is not "the state of the world". It is only the part of the world the agent can see. A poker agent sees its own cards, never yours.

Reward — one number. That is the whole of it. Not a paragraph of feedback, not a label. One number per step.

Is it over? — a yes or no. One complete run from start to finish is called an episode. One game of chess is an episode. One delivery run is an episode.

Reward is smaller than you think

This trips up almost everyone. The reward is a single number, and it arrives on a schedule you chose.

In chess, most people give no reward at all for every move. One point arrives at the end for a win, and a penalty for a loss. Sixty moves, and only the last one carries any information.

That kind of reward is called sparse — it stays silent almost all the time. Sparse rewards are honest but very hard to learn from.

The temptation is to hand out small rewards along the way to help. That temptation causes more disasters than any other habit in this field, and reward shaping is the lesson about why.

Return, not reward

The agent is not trying to get a big reward on this step. It is trying to get a big total over the whole episode. That total has a name: the return.

The difference is everything. Sacrificing your queen looks terrible this move and wins the game four moves later.

   reward this step:  0    0    0    0    +1
   return from the start:  the whole row added up

Most agents also count later rewards slightly less than sooner ones. A sweet today beats the promise of a sweet next year. That preference is called discounting, and it also stops the total running away to infinity in a task that never ends.

Where you have already seen this loop

  • A thermostat that learns your flat's heating. Action: heater on or off. Observation: the temperature. Reward: comfortable, and a small penalty for the electricity bill.
  • An ad system. Action: which ad to show. Observation: who is looking. Reward: whether they clicked.
  • A warehouse robot. Action: motor commands. Observation: camera and wheel sensors. Reward: parcels delivered, penalty for collisions.
  • A chatbot being tuned. Action: the next word. Observation: the conversation so far. Reward: how a human rated the finished reply.

What is honestly hard here

Choosing what the agent gets to observe is a design decision, and it is often the whole problem. Give a robot too little and it cannot possibly succeed. Give it too much and it drowns.

Writing the reward is harder than writing the agent. You have to express what you actually want as a single number. That number gets read thousands of times a second, by something that will exploit any gap you leave. Teams routinely spend more time on the reward than on the algorithm.

Remember this

  • The agent decides, the environment reacts, the reward scores it. Nothing else.
  • The agent sees an observation, not the world. These are different, and the gap matters.
  • Agents maximise the return — the whole episode's total — never one step's reward.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install gymnasium numpy

Gymnasium is the standard Python interface for reinforcement learning environments. It is the maintained successor to OpenAI Gym, which is no longer updated. Install size is a few megabytes, and everything here runs on a CPU.

The loop, in real code

Gymnasium gives every environment the same two methods, and the whole field is built on that agreement.

the_loop.py
import gymnasium as gym

env = gym.make("FrozenLake-v1", map_name="4x4", is_slippery=False)

print("what the agent can see :", env.observation_space)   # 16 squares, numbered 0..15
print("what the agent can do  :", env.action_space)        # 4 moves
print()

obs, info = env.reset(seed=42)
env.action_space.seed(42)

MOVES = ["left", "down", "right", "up"]
total = 0.0

for t in range(6):
    action = env.action_space.sample()                     # a random flail, on purpose
    obs, reward, terminated, truncated, info = env.step(action)
    total += reward
    print(f"step {t}: moved {MOVES[action]:<5} -> square {obs:2d}  reward {reward}  over={terminated}")
    if terminated or truncated:
        print("  episode ended, resetting")
        obs, info = env.reset()
        break

print()
print("reward collected:", total)
env.close()
Output
what the agent can see : Discrete(16)
what the agent can do  : Discrete(4)

step 0: moved left  -> square  0  reward 0  over=False
step 1: moved up    -> square  0  reward 0  over=False
step 2: moved right -> square  1  reward 0  over=False
step 3: moved down  -> square  5  reward 0  over=True
  episode ended, resetting

reward collected: 0.0

The random agent walked into a hole on step 3 and the episode ended with nothing. That is the honest baseline every agent has to beat.

The five values step returns

env.step(action) hands back exactly five things, always in this order.

obs — the observation. Here it is the number of the square, 0 to 15. Discrete(16) means "one integer from a set of sixteen".

reward — a float. FrozenLake pays 1.0 for reaching the goal and 0.0 for everything else, including falling in a hole. Notice that the hole is not punished. Falling in is bad only because it ends the episode without the point.

terminated — the episode genuinely finished. The goal was reached, or the agent died.

truncated — the episode was cut off from outside, usually by a step limit. Keeping this separate from terminated is not pedantry. A truncated episode still has future value that you should bootstrap from; a terminated one does not. Merging them is a real and common bug.

info — a dictionary of extras for debugging. Never train on it.

env.reset(seed=42) seeds the environment, and env.action_space.seed(42) seeds the random action sampler separately. Both are needed for a reproducible run, and forgetting the second one is why your "seeded" experiment still moves.

Writing your own environment, with no library at all

There is nothing magic in gym.make. An environment is any object with reset and step. Here is one, and two agents to run inside it.

road_crossing.py
import random

class RoadCrossing:
    """The environment. It owns the truth. The agent only sees what step() hands back."""
    WAIT, WALK = 0, 1

    def __init__(self, seed=0):
        self.rng = random.Random(seed)

    def reset(self):
        self.light = "red"
        return self.light                       # the observation

    def step(self, action):
        if action == self.WALK:
            reward = 5.0 if self.light == "green" else -20.0
            return self.light, reward, True     # observation, reward, episode over
        if self.light == "red" and self.rng.random() < 0.4:
            self.light = "green"
        return self.light, -0.5, False          # standing still costs a little


def careful(observation):
    return RoadCrossing.WALK if observation == "green" else RoadCrossing.WAIT

def reckless(observation):
    return RoadCrossing.WALK


def run(agent, name, seed=0):
    env = RoadCrossing(seed=seed)
    obs, total, t = env.reset(), 0.0, 0
    print(f"--- {name} ---")
    while True:
        action = agent(obs)
        obs, reward, done = env.step(action)
        total += reward
        print(f"t={t}  did {'WALK' if action else 'WAIT'}  light now {obs:<5}  reward {reward:+.1f}")
        t += 1
        if done:
            break
    print(f"return for this episode: {total:+.1f}\n")

run(careful, "careful agent")
run(reckless, "reckless agent")
Output
--- careful agent ---
t=0  did WAIT  light now red    reward -0.5
t=1  did WAIT  light now red    reward -0.5
t=2  did WAIT  light now red    reward -0.5
t=3  did WAIT  light now green  reward -0.5
t=4  did WALK  light now green  reward +5.0
return for this episode: +3.0

--- reckless agent ---
t=0  did WALK  light now red    reward -20.0
return for this episode: -20.0

What that output is actually teaching

The careful agent earned four penalties of half a point and still finished ahead. Return, not immediate reward. Every one of those waits looked like a loss at the time.

Now look at the numbers you would have to choose as a designer. Waiting costs -0.5 per step and crossing on green pays +5.0. Raise the waiting cost to -3.0 and the careful agent's return drops below zero after two waits — at which point walking into traffic starts to look attractive to it.

You did not change the algorithm. You changed three numbers, and the agent's whole personality changed with them. Reward design is engineering, not decoration.

Notice also that the agent never sees self.rng. It has an observation, not the state. It cannot know when the light will change, only that it has not yet.

Common mistakes

Treating truncated as terminated. When a step limit cuts an episode short, the future is not worthless — the pole was still balanced. Bootstrap from the final observation for truncated, and use zero for terminated. This mistake silently caps how good your agent can get.

Forgetting env.reset() before the first step. Gymnasium raises a clear error for this, which is a kindness. Your own environments should do the same.

Rewarding the agent for something it cannot observe. If the reward depends on wind speed and the agent never sees wind speed, you have handed it a noise generator and asked it to find a pattern.

Putting the goal in info and the score in reward, then training on info. info is for humans. If a value matters to learning, it belongs in the observation or the reward.

Try it yourself

In RoadCrossing, change the waiting penalty from -0.5 to -3.0 and re-run. Predict what the careful agent's return becomes before you look. Then find the exact penalty at which waiting for green stops being worth it, and you will have derived a real reward-design boundary by hand.

What to learn next

Researcher — Mathematics and papers.

The agent-environment interface, formally

At each discrete time step $t$ the environment is in state $s_t \in \mathcal{S}$. The agent samples $a_t \sim \pi(\cdot \mid s_t)$. The environment produces $s_{t+1} \sim P(\cdot \mid s_t, a_t)$ and a scalar $r_{t+1} \sim R(\cdot \mid s_t, a_t)$.

The return from time $t$ is $$ G_t = \sum_{k=0}^{T-t-1} \gamma^{k} r_{t+k+1} $$ with $T = \infty$ for continuing tasks (requiring $\gamma < 1$ for convergence when rewards are bounded) and $T$ finite for episodic ones.

Two boundaries are worth stating precisely.

The agent-environment boundary is a modelling choice, not a physical one. Sutton and Barto's rule: anything the agent cannot change arbitrarily by choice is part of the environment. A robot's motors are environment; the controller's memory is agent. Reward computation always sits on the environment side, because an agent able to edit its own reward function has a degenerate optimal policy.

The observation-state distinction is the difference between an MDP and a POMDP. When $o_t = s_t$ the process is fully observed. Otherwise you have a partially observable MDP, $(\mathcal{S}, \mathcal{A}, P, R, \Omega, O, \gamma)$, with observation set $\Omega$ and observation function $O(o \mid s', a)$. In a POMDP, no memoryless policy $\pi(a \mid o)$ is optimal in general — the sufficient statistic is the belief state $b_t(s) = \Pr(s_t = s \mid o_{1:t}, a_{1:t-1})$, and the belief-MDP over $\Delta(\mathcal{S})$ is continuous even when $\mathcal{S}$ is finite. Solving POMDPs exactly is PSPACE-complete for finite horizons (Papadimitriou and Tsitsiklis, 1987).

In practice this is why deep RL agents stack frames or use recurrent policies. A single Atari frame does not reveal velocity; four stacked frames approximately do.

Termination versus truncation

This has a precise consequence for the bootstrap target. For a transition ending at $s_{t+1}$:

$$ y_t = \begin{cases} r_{t+1} & \text{terminated (absorbing state, } V(s_{t+1}) = 0) \ r_{t+1} + \gamma V(s_{t+1}) & \text{truncated (time limit, value is not zero)} \ r_{t+1} + \gamma V(s_{t+1}) & \text{ongoing} \end{cases} $$

Treating a time-limit truncation as a termination injects a systematic negative bias into $V$ near the horizon. Pardo et al. (2018), Time Limits in Reinforcement Learning, quantifies this and shows measurable performance loss on standard benchmarks. Gymnasium's five-tuple API exists precisely to make the distinction unavoidable; the older four-tuple done flag conflated them.

Reward specification

The reward function is part of the problem statement, not the solution. Three properties are worth checking before training anything.

Scale. Value targets inherit the reward scale. Rewards in the thousands produce large TD errors and unstable gradients. Clipping to $[-1, 1]$ (as DQN did) stabilises training but changes the objective: it makes the agent indifferent between a reward of 10 and a reward of 1000.

Sparsity. With reward only at a goal state, the probability that undirected exploration reaches it decays exponentially in the horizon. This is the motivation for reward shaping, curiosity bonuses (Pathak et al., 2017), and hindsight relabelling (Andrychowicz et al., 2017).

Alignment. The reward defines the optimal policy exactly. Any gap between the reward and the designer's intent is not a bug the agent will forgive. Krakovna et al. maintain a catalogue of specification-gaming examples that is worth an hour of anyone's time.

Discounting is not only a mathematical convenience

Three separate readings of $\gamma$, all correct and sometimes in tension:

  1. Convergence. It bounds $|G_t| \le r_{\max}/(1-\gamma)$ for bounded rewards.
  2. Preference. It encodes genuine impatience, the analogue of an interest rate.
  3. Termination probability. A $\gamma$-discounted infinite-horizon problem is equivalent to an undiscounted problem with a $(1-\gamma)$ chance of ending at each step.

The effective horizon is $\approx 1/(1-\gamma)$: 100 steps at $\gamma = 0.99$, 1000 at $\gamma = 0.999$. Raising $\gamma$ increases the variance of return estimates and typically slows learning, even when the true task is undiscounted. Jiang et al. (2015), The Dependence of Effective Planning Horizon on Model Accuracy, shows that a smaller-than-true $\gamma$ can be the better choice under model error — a real bias-variance trade-off, not a hack.

References

  • Sutton and Barto (2018), Reinforcement Learning: An Introduction, chapter 3 — the canonical treatment of this interface.
  • Kaelbling, Littman and Cassandra (1998), Planning and acting in partially observable stochastic domains.
  • Pardo et al. (2018), Time Limits in Reinforcement Learning — arxiv.org/abs/1712.00378.
  • Towers et al. (2024), Gymnasium: A Standard Interface for Reinforcement Learning Environments — arxiv.org/abs/2407.17032.

What to learn next