CLI tool
PWhiddy/PokemonRedExperiments avatar
PWhiddy/PokemonRedExperiments

PokemonRedExperiments: Teaching a Reinforcement Learning Agent to Play Pokemon Red

Playing Pokemon Red with Reinforcement Learning

7,908 stars786 forksJupyter NotebookMIT

At a glance

What is it?
PokemonRedExperiments is an open-source project that trains reinforcement learning agents to play Pokemon Red using Stable Baselines 3 and the PyBoy Game Boy emulator. It includes a pretrained model, two training scripts, TensorBoard integration, and a live training broadcast system.
Who is it for?
PokemonRedExperiments is suited for RL researchers and hobbyists who want a working, documented starting point for training agents on a real Game Boy game, and who understand the reinforcement learning fundamentals well enough to interpret the training metrics. It requires Python 3.10 or later, ffmpeg, a legally obtained Pokemon Red ROM file, and enough compute to run a training session to meaningful depth.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 18 days ago.
What is it written in?
Mainly Jupyter Notebook, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What PokemonRedExperiments Is and Who It Is For

PokemonRedExperiments is a research project that applies reinforcement learning to a complete Game Boy game. The agent runs inside a PyBoy emulator instance, receives the game's screen state as observations, and outputs button presses as actions. Stable Baselines 3 provides the RL algorithm infrastructure.

The project targets machine learning researchers, students, and hobbyists who want to study RL applied to a non-trivial sequential decision problem. Pokemon Red is a useful benchmark because it has a clear progression structure, sparse rewards, and a large state space. The agent must explore maps, manage inventory, win battles, and navigate story sequences, which tests a wide range of exploration and generalization challenges.

A companion paper was published in 2025 with the arxiv identifier 2502.19920, and the README documents follow-up projects including a multiplayer live training broadcast, PufferAI/pokegym, and an academic PokeRL site. The project has produced real research outputs, not just a demonstration.

How the RL Agent Explores Pokemon Red

The V1 training baseline used a KNN-based frame exploration reward: the agent received a bonus for visiting game states that differed from previously seen frames, which pushed it to explore new areas. This approach proved effective for early exploration but had limitations as the game state space grew.

V2, which the README describes as the recommended version, replaces the frame KNN with a coordinate-based exploration reward. The agent receives rewards based on map coordinates it has not visited before, which is more memory-efficient and more semantically meaningful for navigation tasks. V2 also trains faster, uses less memory, and streams to the training broadcast map by default.

The Stable Baselines 3 library provides the RL algorithm. The training scripts configure the environment, wrap it with the broadcast system if needed, and run the standard SB3 training loop. The game environment is implemented using PyBoy's game wrapper interface, which exposes game state data such as player coordinates and game memory.

Setting Up the Environment and Running the Pretrained Model

Python 3.10 or later is recommended. ffmpeg must be installed and available from the command line. A legally obtained Pokemon Red ROM file is required and must be named PokemonRed.gb and placed in the base directory. The README provides a SHA1 checksum to verify the correct ROM version:

code
shasum PokemonRed.gb

The expected SHA1 is ea9bcae617fdf159b045185467ae58b2e4a48b9a. After placing the ROM, move into the appropriate directory and install dependencies:

code
cd baselines
pip install -r requirements.txt

To run the pretrained model interactively:

code
python run_pretrained_interactive.py

The interactive mode lets the user control the game with arrow keys and the `a` and `s` keys while the AI runs alongside. Pausing the AI input during the game is done by editing `agent_enabled.txt`. For macOS, the V2 directory provides a separate `macos_requirements.txt` because the standard requirements file may not work on that platform.

Training Your Own Agent with V2

V2 is in the `v2` directory and uses the same setup pattern as the baselines directory. The V2 training script is:

code
python baseline_fast_v2.py

V2 reaches Cerulean City according to the README, which is the third major landmark in Pokemon Red and requires defeating the first gym, navigating routes, and entering Cerulean. Reaching Cerulean demonstrates that the coordinate-based exploration reward successfully drives meaningful progression through the early game.

The training broadcast is enabled by default in V2. It wraps the environment with a StreamWrapper that sends real-time training data to a shared global game map viewable at pwhiddy.github.io/pokerl-map-viz/. The wrapper accepts optional metadata:

python
env = StreamWrapper(
            env,
            stream_metadata = {
                "user": "super-cool-user",
                "env_id": id,
                "color": "#0033ff",
                "extra": "",
            }
        )

This system lets multiple training sessions from different users appear simultaneously on a shared map visualization, which is a useful tool for comparing exploration coverage across different runs.

Monitoring Training Progress

Two monitoring tools are available. TensorBoard is built into the training workflow. After starting a training session, move into the session directory and run:

code
tensorboard --logdir .

Navigate to localhost:6006 to view loss curves, reward metrics, and other training statistics. The game state at each step is also rendered to image files in the session directory, providing a visual audit trail of what the agent was doing.

Weights and Biases integration is available as a toggle. The README instructs setting `use_wandb_logging` to `True` in the training script. wandb is listed as a dependency but requires a wandb account for the cloud logging features. The local TensorBoard setup works without any external accounts.

The visualization directory contains static map visualization code for generating images of explored map areas from completed training runs.

Limitations and Legal Requirements

The most significant operational constraint is the ROM requirement. The README states that a legally obtained Pokemon Red ROM is necessary, defines the correct SHA1 hash, and provides no download source. The ROM must be sourced by the user. This is a legal requirement, not a technical one, and it applies in all jurisdictions regardless of whether the user owns the original cartridge.

Training to meaningful depth requires sustained compute. The README does not document how many training steps are needed to reach Cerulean, how long that takes on typical hardware, or what GPU memory the training requires. Developers should treat the project as a research starting point that requires tuning and iteration rather than a push-button solution.

V2 is the recommended version, but the V1 baselines directory remains in the repository for reference. Running V1 code on current dependencies may produce compatibility issues since the README note about the updated V2 script implies V1 was written for an earlier dependency state.

PokemonRedExperiments vs. PufferAI/PokeGym

PufferAI/pokegym is a follow-up project explicitly linked in the README as a successor environment. PokeGym is a packaged gymnasium environment for Pokemon, built on top of the PokemonRedExperiments work. It is designed for research that needs a standardized environment interface compatible with the broader ecosystem of RL libraries and benchmarking tools.

PokemonRedExperiments is the original experimental codebase. It is less abstracted and more directly tied to specific training scripts and the PyBoy emulator configuration. PufferAI/pokegym packages that into a cleaner environment class.

For researchers who want to run experiments and compare results with others using standardized tooling, pokegym is likely the more appropriate starting point. For developers who want to understand the implementation at the lowest level or adapt the training setup in non-standard ways, PokemonRedExperiments provides more direct access to the components.

Editorial conclusion

PokemonRedExperiments is suited for RL researchers and hobbyists who want a working, documented starting point for training agents on a real Game Boy game, and who understand the reinforcement learning fundamentals well enough to interpret the training metrics. It requires Python 3.10 or later, ffmpeg, a legally obtained Pokemon Red ROM file, and enough compute to run a training session to meaningful depth. It is not a beginner project: the setup requires comfort with Python environments, command-line tools, and basic RL concepts. Developers who want a more polished gymnasium environment for Pokemon should evaluate PufferAI/pokegym, which is listed as a follow-up project in the README.

Frequently asked questions

Do I need to own the original Pokemon Red cartridge to use PokemonRedExperiments?

The README requires a legally obtained Pokemon Red ROM file named PokemonRed.gb. It does not document how to obtain the ROM legally. Users are responsible for ensuring their ROM is legally sourced according to the laws in their jurisdiction.

How does the reward system work in PokemonRedExperiments V2?

V2 uses a coordinate-based exploration reward: the agent receives bonuses for visiting map coordinates it has not previously reached. This replaces the V1 approach of rewarding frame novelty based on KNN distance, which is more memory-efficient and more semantically aligned with navigation progress.

How do I track training progress in PokemonRedExperiments?

Run `tensorboard --logdir .` from the session directory and navigate to localhost:6006 to view training metrics. Alternatively, set `use_wandb_logging` to `True` in the training script to log to Weights and Biases. Game state images are also written to the session directory at each step.

Official sources

  1. Issues
  2. License: MIT
  3. PWhiddy/PokemonRedExperiments on GitHub
  4. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/pwhiddy-pokemonredexperiments.svg)](https://hysenlabs.com/projects/pwhiddy-pokemonredexperiments)