Gymnasium and training environments
Gymnasium is the standard way to describe a reinforcement learning problem in Python, so any agent can be plugged into any world without either knowing about the other.
- 17 min read
- 3 reading levels
- Updated
Read these first
On this page 10
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
Gymnasium is an agreement about how a reinforcement learning world talks to an agent, so that any agent works with any world.
Two methods, five return values, and that is the entire agreement.
The analogy you have already lived
Think about cricket nets.
You want to face a thousand deliveries this month. You cannot play a thousand matches. So you go to the nets: a safe, cheap, repeatable place to make the same mistake fifty times without anyone minding.
A training environment is the nets. A robot learning to walk falls over ten thousand times. Doing that with a real robot means ten thousand repairs. Doing it in a simulator costs an afternoon of electricity.
Almost all reinforcement learning happens in the nets. That fact shapes everything about how the field works, and it is also its biggest weakness.
Why a standard was needed
Before there was one, every research group wrote its own worlds with its own function names. An agent written for one lab's code could not be run on another's without a rewrite. Comparing two published results honestly was close to impossible.
In 2016 a common interface appeared, called Gym. In 2022 maintenance moved to a group who renamed it Gymnasium. That is the one to use. The older gym package is no longer maintained.
The value is boring and enormous: your agent talks to a socket, not to a specific world.
your agent -> reset() -> the world starts fresh
-> step(a) -> what you see, your score, is it over?Change the world from a balancing pole to a robot arm to an Atari game. The agent code does not change one line.
The two methods
reset() puts the world back to a starting position and hands you the first thing you can see. You call it at the beginning, and again after every ending.
step(action) does one action and hands back four things. What you see now. Your score for that step. Whether the episode genuinely ended, and whether it was instead cut short by a time limit.
That final distinction sounds fussy and it matters. An episode that ended because the pole fell is different from one that stopped because a stopwatch ran out. In the second case the pole was still balanced, and an agent that treats the two the same learns something false. The agents and rewards lesson goes into why.
Spaces: describing what is possible
Every world declares what an observation looks like and what actions are allowed. These declarations are called spaces.
- Discrete — a whole number from a fixed set. Four moves. Sixteen squares.
- Box — one or more decimal numbers, each with a lowest and highest value. A pole's angle. A joint's position.
- Dict — several of the above, named, for worlds that show you more than one kind of thing at once.
Because the world declares this, your agent can build itself to fit. That is why the same script can be pointed at a new world and still work.
Wrappers: changing a world without editing it
A wrapper is a thin layer you put around a world to change one thing about it, leaving the original untouched.
[ your agent ]
|
[ wrapper: keep a running total of each episode's score ]
|
[ wrapper: cut every episode off after 500 steps ]
|
[ the actual world ]Each layer looks like a world to whatever is above it. Common jobs: enforcing a time limit. Scaling observations into a sensible range. Stacking the last four frames so the agent can perceive motion. Recording a video.
Wrappers are the reason you rarely need to edit an environment. Wrap it instead.
Running many copies at once
Reinforcement learning is slow mostly because it needs an enormous amount of experience. A convenient fix is to run several copies of the world side by side, each with a different starting position, and step all of them together.
Eight copies gather experience eight times faster, and the mix of situations makes learning steadier as well. Every serious training run does this.
What is honestly hard here
The nets are not the match. A policy that walks beautifully in a simulator often falls over on a real robot. The simulator's friction, motor delay and sensor noise were all slightly wrong. This gap is called the reality gap, and closing it is an entire research area of its own. The common defence is to randomise the simulator's physics during training, so the agent has to work across a range of worlds rather than exactly one.
A benchmark score is not a capability. An agent that scores well on a standard test environment has learned that environment. Published results are often reported on the same handful of worlds the method was developed on.
Making your own environment is where most of the work is. In real projects, writing a correct, fast simulator of your actual problem takes far longer than choosing an algorithm. Nobody warns you about this, and it is nearly always true.
Remember this
- Gymnasium is an agreement:
resetandstep, so agents and worlds are interchangeable. - Spaces describe what can be seen and done. Wrappers change a world without editing it.
- The reality gap is real: doing well in the nets is not the same as doing well in the match.
What to learn next
- Agents, environments and rewards — the contract this lesson implements.
- PPO — the algorithm to point at your new environment first.
- Reward shaping — because the reward you wrote into your environment is the thing you will get.
Developer — Code and libraries.
Setup
pip install gymnasium numpyInstall gymnasium, not gym. The old package is unmaintained and its step returns four values instead of five, which breaks every modern example including the ones on this site.
Extra environment families are separate installs, so you download only what you use:
pip install "gymnasium[classic-control]" # CartPole, MountainCar, Pendulum
pip install "gymnasium[box2d]" # LunarLander, BipedalWalker
pip install "gymnasium[atari]" ale-py # Atari, needs ROMs accepted separately
pip install "gymnasium[mujoco]" # continuous-control roboticsEverything below needs only the base install.
Spaces, wrappers and vector environments
import gymnasium as gym
import numpy as np
env = gym.make("CartPole-v1")
env.observation_space.seed(0) # spaces have their own generator
print("observation space:", env.observation_space) # four numbers, each with a range
print("action space :", env.action_space) # two choices
print("a random observation:", np.round(env.observation_space.sample(), 3))
print("wrappers stacked on it:", env)
print()
# RecordEpisodeStatistics quietly totals the reward for you, so you stop writing that loop.
env = gym.wrappers.RecordEpisodeStatistics(gym.make("CartPole-v1"))
env.action_space.seed(0)
obs, info = env.reset(seed=0)
while True:
obs, reward, terminated, truncated, info = env.step(env.action_space.sample())
if terminated or truncated:
print("episode summary from the wrapper:", info["episode"])
break
env.close()
print()
# Vector environments: several copies stepping together, which is where the speed comes from.
vec = gym.make_vec("CartPole-v1", num_envs=4)
obs, info = vec.reset(seed=0)
print("batched observation shape:", obs.shape)
obs, reward, terminated, truncated, info = vec.step(np.array([0, 1, 0, 1]))
print("rewards from four copies :", reward)
print("terminated flags :", terminated)
vec.close()observation space: Box([-4.8 -inf -0.41887903 -inf], [4.8 inf 0.41887903 inf], (4,), float32)
action space : Discrete(2)
a random observation: [-4.407 0.126 -0.405 -0.132]
wrappers stacked on it: <TimeLimit<OrderEnforcing<PassiveEnvChecker<CartPoleEnv<CartPole-v1>>>>>
episode summary from the wrapper: {'r': 18.0, 'l': 18, 't': 0.000192}
batched observation shape: (4, 4)
rewards from four copies : [1. 1. 1. 1.]
terminated flags : [False False False False]The 't' value is wall-clock seconds and will differ on your machine. The rest should match.
Three things worth noticing in that output
gym.make already wrapped your environment three deep. TimeLimit<OrderEnforcing<PassiveEnvChecker<CartPoleEnv>>>. TimeLimit is what produces truncated at 500 steps. OrderEnforcing raises a clear error if you call step before reset. PassiveEnvChecker validates types on the first few calls and is worth keeping while developing.
The observation space has infinities in it. Cart velocity and pole angular velocity are unbounded. The sampled observation shows a cart position of -4.407 and a pole angle of -0.405 — both near the edges of the declared range, and both states in which the episode would already have ended. Sampling an observation space gives you a valid shape, never a plausible state. Do not use it to build test data.
The vector environment's shape is (4, 4). Four copies, four numbers each. Vector environments auto-reset a copy when it finishes, so the batch never has holes in it — check your library's version for how the final observation of a finished episode is reported, since this detail changed between Gymnasium releases.
Writing your own environment
This is the part you will actually need. A thermostat: keep a room at 22 degrees, with the heater costing a little to run.
import gymnasium as gym
import numpy as np
from gymnasium import spaces
class Thermostat(gym.Env):
"""Keep a room at 22 degrees. Turning the heater on costs a little electricity."""
metadata = {"render_modes": []}
TARGET = 22.0
def __init__(self):
super().__init__()
self.observation_space = spaces.Box(low=0.0, high=40.0, shape=(1,), dtype=np.float32)
self.action_space = spaces.Discrete(2) # 0 = heater off, 1 = heater on
def reset(self, seed=None, options=None):
super().reset(seed=seed) # this seeds self.np_random for you
self.temp = float(self.np_random.uniform(15.0, 29.0))
self.t = 0
return np.array([self.temp], dtype=np.float32), {}
def step(self, action):
self.temp += 1.0 if action == 1 else -0.7 # heating up, or cooling toward the street
self.temp += float(self.np_random.normal(0.0, 0.2)) # draughts
self.temp = float(np.clip(self.temp, 0.0, 40.0))
self.t += 1
reward = -abs(self.temp - self.TARGET) - 0.1 * action # comfort first, then the bill
terminated = False # nothing here ever "ends"
truncated = self.t >= 100 # so we stop on a clock instead
return np.array([self.temp], dtype=np.float32), reward, terminated, truncated, {}
from gymnasium.utils.env_checker import check_env
check_env(Thermostat(), skip_render_check=True)
print("check_env passed: the environment obeys the Gymnasium contract")
gym.register(id="Thermostat-v0", entry_point=Thermostat)
env = gym.make("Thermostat-v0")
def simple_agent(observation):
return 1 if observation[0] < Thermostat.TARGET else 0
obs, _ = env.reset(seed=3)
total = 0.0
print(f"\nstarting temperature: {obs[0]:.1f}")
for t in range(100):
obs, reward, terminated, truncated, _ = env.step(simple_agent(obs))
total += reward
if t in (0, 4, 20, 99):
print(f" t={t:3d} temperature {obs[0]:5.1f} reward {reward:+.2f}")
if terminated or truncated:
break
print(f"return over the episode: {total:.1f}")
env.close()check_env passed: the environment obeys the Gymnasium contract starting temperature: 16.2 t= 0 temperature 16.7 reward -5.41 t= 4 temperature 20.5 reward -1.58 t= 20 temperature 23.0 reward -1.08 t= 99 temperature 22.7 reward -0.79 return over the episode: -68.4
The details that make it a correct environment
super().reset(seed=seed) — this one line gives you self.np_random, a properly seeded generator. Use it for every random draw in the environment. Reaching for np.random or the random module directly makes your environment unseedable, and you will discover this the day you try to reproduce a result.
dtype=np.float32 on both the space and the returned array — they must match. Returning a float64 array from a float32 space is the most common check_env failure. It also silently doubles memory in a replay buffer.
terminated versus truncated. This thermostat never ends on its own; a room does not finish. So terminated is permanently False and the episode stops on a clock, which is a truncation. Getting this backwards teaches a critic that surviving to step 100 is worth nothing thereafter.
The reward is a design choice, and you can see it in the numbers. -abs(temp - 22) - 0.1 * action says a degree of discomfort costs ten times as much as running the heater for a step. Change that 0.1 to 5.0 and the best policy becomes "leave the heater off and be cold", because the bill now dominates. The algorithm would not change at all. The behaviour would change completely.
check_env — run it on every environment you write, once, before you train anything. It catches wrong dtypes, spaces that do not contain the observations you return, and reset signatures that ignore seed. Ten seconds now against a day of confusion later.
Common mistakes
Installing gym instead of gymnasium. Symptom: ValueError: too many values to unpack (expected 4), or the reverse. The old API returned four values from step; the new one returns five.
Returning a Python float where a Box space was declared. A Box with shape=(1,) needs an array of shape (1,), not a scalar. check_env catches this immediately.
Using np.random inside the environment. Your environment is then unseedable regardless of what you pass to reset(seed=...).
Mutating and returning the same array object every step. A replay buffer holding references to one array will find every entry showing the latest state. Return a fresh array, as the code above does.
Forgetting env.close() on vector or video-recording environments. Leaked subprocesses and truncated video files.
Believing a simulator that has never been checked against reality. If the plan is to deploy on hardware, measure the real system's response early and compare. Discovering a wrong friction coefficient after a month of training is a painful and very common way to lose a month.
Try it yourself
Change the thermostat's cooling rate from -0.7 to -2.5, so the room loses heat much faster, and run the same agent. Its simple rule — heat when below target — now overshoots and undershoots badly, and the return gets much worse. Then write a better rule by hand before reaching for any algorithm. Knowing what a hand-written policy scores is the baseline every reinforcement learning result should be compared against, and it is skipped far too often.
What to learn next
- Agents, environments and rewards — the contract this lesson implements.
- PPO — the algorithm to point at your new environment first.
- Reward shaping — because the reward you wrote into your environment is the thing you will get.
Researcher — Mathematics and papers.
What the interface actually standardises
Gymnasium (Towers et al., 2024) formalises the agent-environment interface for a POMDP as a five-tuple return from step: $(o_{t+1}, r_{t+1}, \text{terminated}, \text{truncated}, \text{info})$.
The separation of terminated and truncated is the substantive change from the original Gym API and it is not cosmetic. Bootstrapping requires
$$
y_t = r_{t+1} + \gamma V(s_{t+1}) \cdot \mathbb{1}[\neg \text{terminated}]
$$
Time-limit truncation must not zero the bootstrap. Pardo et al. (2018) quantify the bias this introduces and demonstrate measurable performance loss on standard benchmarks when the two are conflated. The old single done flag made the correct implementation impossible without reading info, which most code did not.
A second correctness issue: a fixed time limit makes the process non-Markov unless the remaining time is part of the observation. For a genuinely time-limited task, include $t$ or $T-t$ in the observation. For a continuing task that you cut short only for practical reasons, do not — and bootstrap at the cut.
Determinism and seeding
Reproducibility requires seeding three independent sources: the environment's transition and initial-state randomness (env.reset(seed=...), exposed as self.np_random), the action-space sampler (env.action_space.seed(...)), and your framework's global generators.
Even then, exact reproduction across machines is not guaranteed. Floating-point reduction order varies with thread count and BLAS build; cuDNN kernel selection is non-deterministic by default. Report seeds, and report distributions across seeds rather than single runs. Henderson et al. (2018) showed that identical algorithms on identical benchmarks produce non-overlapping performance distributions across seed groups — one of the more uncomfortable results in the field's recent history, and still largely unaddressed in practice.
Benchmark suites, and what each measures
| Suite | Domain | What it actually tests |
|---|---|---|
| Classic control | Low-dimensional continuous | Debugging; solvable in minutes on CPU |
| ALE / Atari | Pixels, discrete actions | Representation learning; deterministic without sticky actions |
| MuJoCo / Gymnasium-Robotics | Continuous control | Locomotion; heavily overfit by the field |
| Procgen | Procedural 2D | Generalisation to unseen levels |
| MiniGrid / BabyAI | Gridworlds, instructions | Sparse reward, compositional generalisation |
| Meta-World | Robotic manipulation | Multi-task and meta-learning |
| D4RL / Minari | Static datasets | Offline RL |
| PettingZoo | Multi-agent | Non-stationarity from other learners |
Two evaluation cautions. Atari is deterministic unless you enable sticky actions (repeat the previous action with probability 0.25) — without them, memorising an action sequence scores well without any policy at all. Machado et al. (2018) is the reference on evaluating on ALE properly. And MuJoCo results are not comparable across versions: solver, contact model and default parameters changed between MuJoCo 1.5, 2.x and 3.x, and between -v2, -v3 and -v4 task definitions.
Vectorisation and throughput
SyncVectorEnv runs copies in one process; AsyncVectorEnv uses subprocesses with shared memory. Asynchronous helps when the environment step is expensive relative to inter-process communication, and hurts for very cheap environments where the serialisation dominates.
The more significant recent development is GPU-resident simulation. Isaac Gym, Brax and Madrona run thousands of environment copies inside the accelerator, eliminating the host-device transfer that dominates classical pipelines. Freeman et al. (2021) report throughput improvements of several orders of magnitude for rigid-body tasks. This changes which algorithms are practical: with effectively free samples, on-policy methods like PPO become preferable to sample-efficient off-policy ones, reversing the usual advice.
The sim-to-real gap
The mismatch between simulator and reality has several independent components: unmodelled dynamics (friction, backlash, cable drag), actuator delay and bandwidth, sensor noise and latency, and calibration error. Standard responses:
- Domain randomisation (Tobin et al., 2017; Peng et al., 2018) — randomise physical parameters during training so the policy must be robust across a distribution of worlds. Effective, and it costs performance on any single world.
- System identification — fit simulator parameters to real trajectories, narrowing the distribution to randomise over.
- Real-world fine-tuning — a short calibration phase on hardware, usually with a conservative offline method to bound the damage.
- Actuator networks (Hwangbo et al., 2019) — learn the motor's real response from hardware data and substitute it into the simulator. The ANYmal results made this a standard technique for legged robots.
Offline RL and datasets
When interaction is unsafe or expensive, learn from a fixed dataset. The core difficulty is distribution shift: a learned $Q$ extrapolates optimistically on out-of-distribution actions, and nothing in the data can correct it. Conservative Q-Learning (Kumar et al., 2020) penalises Q-values on unseen actions; Implicit Q-Learning (Kostrikov et al., 2021) avoids querying them entirely with expectile regression.
D4RL (Fu et al., 2020) was the standard benchmark; Minari is its maintained Gymnasium-native successor. Note the recurring critique that several D4RL tasks are solvable by behaviour cloning on the best trajectories in the dataset, so a strong filtered-BC baseline belongs in any offline RL comparison.
References
- Towers et al. (2024), Gymnasium: A Standard Interface for Reinforcement Learning Environments — arxiv.org/abs/2407.17032.
- Brockman et al. (2016), OpenAI Gym — arxiv.org/abs/1606.01540.
- Machado et al. (2018), Revisiting the Arcade Learning Environment — arxiv.org/abs/1709.06009.
- Henderson et al. (2018), Deep Reinforcement Learning that Matters — arxiv.org/abs/1709.06560.
- Cobbe et al. (2020), Leveraging Procedural Generation to Benchmark Reinforcement Learning (Procgen) — arxiv.org/abs/1912.01588.
- Freeman et al. (2021), Brax — A Differentiable Physics Engine for Large Scale Rigid Body Simulation — arxiv.org/abs/2106.13281.
- Agarwal et al. (2021), Deep Reinforcement Learning at the Edge of the Statistical Precipice — how to report RL results honestly, arxiv.org/abs/2108.13264.
What to learn next
- Agents, environments and rewards — the contract this lesson implements.
- PPO — the algorithm to point at your new environment first.
- Reward shaping — because the reward you wrote into your environment is the thing you will get.