Library / SDK
facebookresearch/Pearl avatar
facebookresearch/Pearl

Pearl: Meta's Modular Reinforcement Learning Agent Library for Production

A Production-ready Reinforcement Learning AI Agent Library brought by the Applied Reinforcement Learning team at Meta.

3,033 stars207 forksJupyter NotebookMIT

At a glance

What is it?
Pearl is a production-oriented reinforcement learning library from Meta's Applied RL team that lets engineers assemble agents from interchangeable components: policy learners, exploration modules, replay buffers, and safety constraints. It targets real-world deployment scenarios with sparse feedback, dynamic action spaces, and high variability, which are the exact areas where standard tutorial-style RL frameworks fall short.
Who is it for?
Pearl is the right starting point for engineers who need a production-oriented RL agent that handles dynamic action spaces, offline learning, or safety constraints without building those pieces from scratch. It requires PyTorch, pip version 21.3 or later, and setuptools 64 or later.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 43 days ago.
What is it written in?
Mainly Jupyter Notebook, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 27, 2026, and from our analysis. They are not legal advice.

Editorial analysis

Production RL Scenarios Gymnasium Alone Does Not Cover

Most reinforcement learning tutorials assume static action spaces, dense reward signals, and full observability. Pearl is designed around the opposite: environments with limited observability, sparse feedback, high stochasticity, and action spaces that change over time. The README describes the motivating use cases as recommender systems, auction bidding systems, and creative selection, each of which has a different mix of these constraints.

Where a Gymnasium CartPole example has a fixed four-dimensional observation and two discrete actions, a production recommender system may add or remove candidate items at each step. Pearl handles this through a dedicated dynamic action spaces feature. The README positions Pearl as the tool for practitioners who need these properties without writing each one from scratch, and for researchers who want a codebase that reflects production constraints rather than benchmark environments.

The README presents a table showing which Pearl features each of the three real-world application types uses. Recommender systems use dynamic action spaces, history summarization, and offline learning. Auction bidding systems use offline learning and intelligent neural exploration. Creative selection uses safety constraints and data augmentation. Not every application needs the full feature set, and Pearl's modular design means unused components are simply not instantiated.

The Modular Component Architecture

Pearl agents are assembled from interchangeable components. The core assembly point is `PearlAgent`, which takes a `policy_learner`, a replay buffer, and optional modules for exploration, safety, history summarization, and data augmentation. Each component is a separate class, and the README states explicitly that practitioners can select any subset and combine them for their specific use case.

The policy learners available include Deep Q-Learning (DQN), Double DQN, and actor-critic algorithms. The replay buffer interface accepts a `BasicReplayBuffer` for standard cases. Contextual bandit algorithms are also supported: the tutorial notebooks demonstrate SquareCB, LinUCB, and LinTS on UCI datasets. Because every component implements a shared interface, swapping one policy learner for another requires changing one constructor argument rather than restructuring the agent. The README identifies a benchmark configuration file at `utils/scripts/benchmark_config.py` where predefined agent configurations are listed.

The design also separates concerns that other RL libraries bundle together. The exploration module is independent of the policy learner, which means you can swap neural exploration strategies without changing the policy algorithm. The safety module is similarly independent, allowing safety constraints to be applied to any policy learner without modifying its code. This separation is unusual in RL libraries and is the main architectural property that distinguishes Pearl from single-class algorithm implementations.

Installing Pearl and Running the CartPole Example

Pearl is not on PyPI. Installation requires cloning the repository and running an editable pip install:

bash
git clone https://github.com/facebookresearch/Pearl.git
cd Pearl
pip install -e .

The README specifies pip version 21.3 or later and setuptools version 64 or later as minimum requirements. The pyproject.toml lists `torch`, `torchvision`, `torchaudio`, `gym`, `gymnasium` with mujoco and atari extras, `numpy`, `matplotlib`, `pandas`, and `mujoco` as dependencies, so the install pulls a full scientific Python stack including deep learning and simulation libraries. The README provides a quick-start code fragment for CartPole-v1 using `GymEnvironment`, `DeepQLearning`, and `BasicReplayBuffer`. Users can replace `GymEnvironment` with any custom environment class that implements the expected interface.

Saving and Loading Agent State Dicts

A January 2025 update added serialization support. Pearl components now produce state dicts using the same interface as PyTorch modules, which means `torch.save` and `torch.load` work directly:

python
agent = PearlAgent(...)
torch.save(agent.state_dict(), 'agent_state.pth')

agent2 = PearlAgent(...)
agent2.load_state_dict(torch.load('agent_state.pth'))

assert agent2.compare(agent) == ""

The README notes two constraints. First, the `agent2` instance must have the same component structure as the original. Second, component attributes that are not PyTorch parameters, buffers, or sub-modules are not included in the state dict automatically; implementing `get_extra_state` and `set_extra_state` methods on those components is required to capture them. Each component must also implement a `compare` method that returns a string describing differences between two instances; this method supports testing and validation.

Where Pearl's Flexibility Creates Friction

Pearl's modular design imposes assembly overhead that simpler libraries avoid. Configuring a working agent requires understanding which combination of policy learner, replay buffer, and exploration module is appropriate for a given environment. The README addresses this by providing tutorial notebooks and a benchmark configuration file, but the initial learning curve is steeper than a single-line API.

Stable Baselines 3 is a widely used alternative that provides clean, documented implementations of standard RL algorithms for Gym-compatible environments. The key difference is scope: Stable Baselines 3 targets standard benchmark environments and does not natively handle dynamic action spaces, offline learning, or safety constraints. Pearl adds those features at the cost of more configuration. Teams that need standard DQN or PPO on a fixed-action-space environment will find Stable Baselines 3 easier to start with. Teams deploying RL in production with changing catalogs or safety constraints will find Pearl's component model more relevant.

The README also notes a specific dependency constraint: the `compare` method required on each custom component is new as of the January 2025 update, which means custom components written before that update will need the method added before they work correctly with the current serialization system. This is a source-level breaking change for anyone who extended Pearl before that date.

Maintenance, Tutorial Coverage, and License

The last push to the Pearl repository was on 2026-08-19. The repository is not archived. Pearl was introduced at NeurIPS 2023 and is developed by Meta's Applied Reinforcement Learning team. The repository provides five tutorial notebooks covering a single-item recommender system using the MIND dataset, contextual bandits, Frozen Lake with DQN, actor-critic with safety constraints, and standard DQN and Double DQN on CartPole. The README notes that more tutorials are in progress.

The version in pyproject.toml is 0.1.0, which matches the README's description of this as a beta release. The MIT license permits unrestricted commercial use, modification, and redistribution. The dependencies include `gymnasium` with atari and mujoco extras, which carry their own license terms; in particular, the atari extra requires accepting the ALE ROM license at install time via the `accept-rom-license` flag in the pyproject.toml dependencies list.

The repository structure separates the library code in `pearl/` from tutorials in `tutorials/` and tests in `test/`. Developer documentation is in `dev_docs/`. Because there are no GitHub releases, tracking changes requires reading commit history or the README's News section directly. The upgrade path from one commit to another is not documented beyond what the News section describes.

Editorial conclusion

Pearl is the right starting point for engineers who need a production-oriented RL agent that handles dynamic action spaces, offline learning, or safety constraints without building those pieces from scratch. It requires PyTorch, pip version 21.3 or later, and setuptools 64 or later. Before adopting, verify that you need at least one of Pearl's production features: dynamic action spaces, safe decision-making, or offline learning. Teams whose workloads fit standard Gymnasium environments without these constraints should evaluate Stable Baselines 3 first, as it offers cleaner documented implementations with less assembly required. The last push to the Pearl repository was on 2026-08-19, and the MIT license permits unrestricted commercial use.

Frequently asked questions

What policy learners does Pearl include?

The README and tutorials describe Deep Q-Learning (DQN), Double DQN, and actor-critic algorithms for sequential decision-making, as well as contextual bandit algorithms including SquareCB, LinUCB, and LinTS. A benchmark configuration file at utils/scripts/benchmark_config.py lists available predefined agent configurations with their component combinations.

Does Pearl require PyTorch?

Yes. The pyproject.toml lists torch, torchvision, and torchaudio as required dependencies. Installing Pearl via pip install -e . pulls the full PyTorch stack alongside numpy, gymnasium, and mujoco.

Can Pearl components be saved and reloaded between training runs?

Yes, as of a January 2025 update. Pearl components produce state dicts compatible with torch.save and torch.load. Components with non-PyTorch attributes must implement get_extra_state and set_extra_state to include those attributes in the saved dict. The receiving agent must have the same component structure as the original for load_state_dict to succeed.

Official sources

  1. facebookresearch/Pearl on GitHub
  2. Issues
  3. License: MIT
  4. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/facebookresearch-pearl.svg)](https://hysenlabs.com/projects/facebookresearch-pearl)