How to Set Up OpenAI Gymnasium
Install Gymnasium in an isolated Python environment, run a verified CartPole example, understand the reset and step API, and add optional environment families safely.
Gymnasium is the maintained successor to OpenAI Gym. To set it up reliably, create a virtual environment, install the core package plus the extra for the environment family you want, and run a small program that checks the modern reset() and step() return values. The most important compatibility detail is that Gymnasium separates episode termination from time-limit truncation.
What Gymnasium provides
Gymnasium gives reinforcement-learning code a common interface for environments, action spaces, observation spaces, rewards, and episode boundaries. It includes classic toy environments for learning and an API for building custom environments.
Older tutorials often import gym and expect this pattern:
observation = env.reset()
observation, reward, done, info = env.step(action)
The maintained API is different:
observation, info = env.reset()
observation, reward, terminated, truncated, info = env.step(action)
terminated means the environment reached its natural terminal state. truncated means the episode ended for an outside reason, commonly a time limit. A training algorithm may want to treat those cases differently.
Create an isolated environment
Use a virtual environment so Gymnasium and its optional dependencies do not interfere with other Python projects.
mkdir gymnasium-demo
cd gymnasium-demo
python -m venv .venv
Activate it on macOS or Linux:
source .venv/bin/activate
Activate it in PowerShell on Windows:
.venv\Scripts\Activate.ps1
Upgrade packaging tools and install the classic-control extra, which includes the dependencies used by the CartPole rendering and environment family:
python -m pip install --upgrade pip
python -m pip install "gymnasium[classic-control]"
Check that the shell is using the virtual environment:
python -c "import gymnasium; print(gymnasium.__version__)"
Using python -m pip instead of a bare pip makes it less likely that packages are installed into a different Python interpreter.
Run a first environment
Create test_gym.py:
import gymnasium as gym
env = gym.make("CartPole-v1")
observation, info = env.reset(seed=42)
env.action_space.seed(42)
for episode_step in range(1000):
action = env.action_space.sample()
observation, reward, terminated, truncated, info = env.step(action)
if terminated or truncated:
print(f"episode ended at step {episode_step + 1}")
observation, info = env.reset()
break
env.close()
Run it with the same interpreter used for installation:
python test_gym.py
The random action is only a smoke test. A reinforcement-learning agent would replace env.action_space.sample() with its policy while still handling the two episode-ending flags.
Understand reset() and step()
reset() starts an episode and returns both the first observation and an information dictionary. Passing a seed makes the initial state reproducible when the environment supports deterministic seeding:
observation, info = env.reset(seed=123)
step(action) returns five values:
observation: the next state seen by the agent.reward: the scalar feedback for the action.terminated: the task’s natural terminal condition was reached.truncated: an external limit ended the episode.info: diagnostic data supplied by the environment.
Always reset after either terminated or truncated. Continuing to call step() after an episode has ended is an environment error, not a useful way to collect more data.
Render CartPole
For a visible window, request human rendering when creating the environment:
import gymnasium as gym
env = gym.make("CartPole-v1", render_mode="human")
observation, info = env.reset(seed=42)
for _ in range(500):
action = env.action_space.sample()
observation, reward, terminated, truncated, info = env.step(action)
if terminated or truncated:
observation, info = env.reset()
env.close()
Rendering is useful for debugging, but it adds a display dependency and is slower than running without a window. For video or image-based evaluation, use the environment’s supported rgb_array mode instead.
Install other environment families
Install only the extra your project needs. Examples include:
python -m pip install "gymnasium[classic-control]"
python -m pip install "gymnasium[box2d]"
Some families require system packages or additional tools. Box2D installations, for example, can need SWIG depending on the platform. If a package fails to build, first check the optional-dependency instructions for that environment family rather than reinstalling the core package repeatedly.
Check spaces before writing an agent
An agent needs to know the shape and type of its actions and observations:
import gymnasium as gym
env = gym.make("CartPole-v1")
print(env.observation_space)
print(env.action_space)
print(env.observation_space.shape)
print(env.action_space.n)
env.close()
CartPole-v1 has a discrete action space, while other environments may expose continuous vectors, images, dictionaries, or tuples. The policy and neural-network input must match the space.
Common setup problems
ModuleNotFoundError: gymnasium— activate.venv, then runpython -m pip show gymnasiumandpython -c "import sys; print(sys.executable)"to confirm that installation and execution use the same interpreter.Environment CartPole-v1 doesn't exist— install the appropriate environment extra, such asgymnasium[classic-control], in the active virtual environment.- A rendering window fails to open — run the non-rendered smoke test first, then check the platform’s display and rendering dependencies.
- Old code expects
done— update the loop to handleterminated or truncated; do not simply rename one variable without understanding the distinction. - An optional package fails to compile — check the extra’s platform requirements and install its system-level prerequisites.
For a focused diagnosis of import, dependency, and API failures, see How to Fix Common Gymnasium Setup Problems.
What to do next
Once the smoke test works, replace random actions with an algorithm, log episode rewards, and keep the environment version and Python dependencies recorded. When building a custom environment, validate its spaces and run it through Gymnasium’s environment checks before training.
Gymnasium is the environment interface, not a complete reinforcement-learning algorithm. You still need to choose an agent implementation, define how observations become model inputs, and decide how terminated and truncated episodes affect learning.
Comments
One comment per thread every 30 minutes · edits are unlimited.