OfflineRL-Kit: A PyTorch Library for Offline Reinforcement Learning Research
An elegant PyTorch offline reinforcement learning library for researchers.
At a glance
- What is it?
- OfflineRL-Kit is a pure PyTorch library that implements nine state-of-the-art offline reinforcement learning algorithms, including CQL, IQL, MOPO, and MOBILE, with a modular design that lets researchers build new algorithms by composing existing components. Installation requires MuJoCo, D4RL, and Python setup tools.
- Who is it for?
- OfflineRL-Kit is a good starting point for researchers who want a clean PyTorch baseline for offline RL experiments and need to run or compare CQL, TD3+BC, IQL, EDAC, MCQ, MOPO, COMBO, RAMBO, or MOBILE on D4RL benchmarks. It is the wrong choice for practitioners who need online RL, continuous deployment, or a framework that abstracts away environment setup.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 53 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What Offline RL Is and Why a Dedicated Library Helps
Online reinforcement learning trains agents by collecting fresh experience from an environment in real time. Offline reinforcement learning trains entirely from a fixed dataset of previously collected transitions, without any further environment interaction during training. This distinction matters because many real-world applications, such as robotics, autonomous driving, and healthcare decision support, cannot afford the trial-and-error cost of online exploration. The agent must learn from historical data alone.
The challenge is that offline RL introduces distribution shift: the learned policy will encounter state-action pairs that are underrepresented or absent in the training data, and standard Q-learning overestimates their value. Algorithms like CQL, IQL, and MOPO address this with different strategies, and researchers comparing those strategies need a common implementation baseline.
OfflineRL-Kit provides that baseline. All nine algorithms share the same data loading, logging, and evaluation infrastructure, which means differences in benchmark results reflect genuine algorithmic variation rather than implementation details. The README describes the framework design as elegant and clear, with a code structure intended to be easy to understand and extend. New algorithms can be built by composing the library's existing components with minimal additional code.
Nine Supported Algorithms Across Model-Free and Model-Based Categories
OfflineRL-Kit separates its algorithms into two families. Model-free methods operate directly on the data without learning a dynamics model. Model-based methods learn a model of the environment dynamics and use it to generate synthetic transitions for policy training.
The model-free algorithms are: Conservative Q-Learning (CQL), TD3+BC, Implicit Q-Learning (IQL), Ensemble-Diversified Actor Critic (EDAC), and Mildly Conservative Q-Learning (MCQ). CQL adds a conservative regularizer to the Q-function to penalize out-of-distribution actions. IQL avoids querying out-of-distribution actions entirely by approximating the value function with expectile regression. EDAC uses an ensemble of critics with diversity regularization to reduce overestimation. MCQ applies mild conservatism rather than the strong penalty in the original CQL.
The model-based algorithms are: MOPO, COMBO, RAMBO, and MOBILE. MOPO uses uncertainty estimates from an ensemble of dynamics models to penalize high-uncertainty transitions. COMBO combines model rollouts with a conservative critic objective without requiring explicit uncertainty quantification. RAMBO trains a model that is adversarially pessimistic, resisting policy exploitation. MOBILE uses model-Bellman inconsistency as a penalty signal.
The README includes a benchmark table across nine D4RL tasks covering halfcheetah, hopper, and walker2d at medium, medium-replay, and medium-expert dataset levels, with results averaged over four seeds for all eight algorithms.
Installing OfflineRL-Kit
Installation requires three steps: MuJoCo, D4RL, and OfflineRL-Kit itself.
First, install the MuJoCo physics engine from mujoco.org and then install mujoco-py matching your MuJoCo version.
Second, install D4RL from source:
git clone https://github.com/Farama-Foundation/d4rl.git
cd d4rl
pip install -e .Third, clone and install OfflineRL-Kit:
git clone https://github.com/yihaosun1124/OfflineRL-Kit.git
cd OfflineRL-Kit
python setup.py installThe setup.py requires gym between versions 0.15.4 and 0.24.1. This is a narrow version range: gym 0.26 and later introduced API changes that break compatibility. The dependency is fixed at this older range because D4RL itself targets it. Users working with newer gym versions will need a separate virtual environment for OfflineRL-Kit experiments.
The package name installed is offlinerlkit. Dependencies declared in setup.py include gym, matplotlib, numpy, pandas, torch, tensorboard, and tqdm. Ray is commented out in setup.py but is used in the tune_example scripts.
Training a CQL Policy and Tuning Hyperparameters
The README shows a CQL training example that walks through creating the environment, loading the D4RL dataset into a ReplayBuffer, defining actor and critic networks, configuring the CQLPolicy, setting up logging, and running the trainer.
The ReplayBuffer is constructed with the dataset size and observation and action specifications. It is loaded with the D4RL qlearning_dataset. Actor and critic networks are MLP backbones with a TanhDiagGaussian distribution for the actor. The CQLPolicy combines all components into a single object that handles the conservative Q-learning update.
Logging is configured through the Logger class with a dictionary specifying output types: consoleout_backup writes to stdout, policy_training_progress writes to CSV, and tb writes TensorBoard event files. The log directory structure encodes the task name, algorithm name, and seed.
For hyperparameter tuning, the library integrates with Ray Tune. A tuning script initializes ray.init(), constructs a config dictionary with grid_search over hyperparameter values, and calls tune.run() with GPU resource allocation per trial:
ray.init()
config["real_ratio"] = tune.grid_search(real_ratios)
config["seed"] = tune.grid_search(seeds)
analysis = tune.run(run_exp, name="tune_mopo", config=config, resources_per_trial={"gpu": 0.5})The full tuning script for MOPO is in tune_example/tune_mopo.py. Run examples for individual algorithms are in run_example/.
Limitations: Gym Version Lock, Online RL, and Scope
The most significant practical limitation is the gym version lock. The setup.py pins gym to the 0.15.4 to 0.24.1 range. Researchers who use modern gym (0.26+) for other work will need to maintain a separate environment for OfflineRL-Kit. D4RL itself has not been updated to support newer gym versions, so this constraint is upstream from OfflineRL-Kit.
OfflineRL-Kit is exclusively an offline RL library. It has no online training loop, no real-time environment interaction mechanism, and no infrastructure for replay buffer collection during training. Researchers who need online, off-policy algorithms like SAC or TD3 in their comparison study will need a separate library.
The repository is also D4RL-centric. Its benchmark table uses only D4RL continuous control tasks. Researchers working with Atari, discrete action spaces, or custom environments will need to adapt the data loading infrastructure, which the modular design should support but which is not demonstrated in the existing examples.
The alternative most directly comparable to OfflineRL-Kit is d3rlpy, which is a separate offline RL library that also runs on PyTorch and covers many of the same algorithms. d3rlpy maintains an active release schedule, provides a higher-level API, and supports both offline and online modes. OfflineRL-Kit's advantage is a cleaner, more research-oriented code structure that is easier to modify at the algorithm level.
Maintenance, License, and the VLARLKit Connection
OfflineRL-Kit is MIT-licensed. The last push was on 2026-08-09, and the repository has no GitHub releases, so version tracking requires monitoring commits. The setup.py sets the version to 0.0.1.
The README mentions a follow-on project called VLARLKit (github.com/VLARLKit/VLARLKit) focused on VLA (vision-language-action) reinforcement learning. The mention suggests the original author has moved toward VLA RL research as a successor direction, which is worth noting when evaluating whether OfflineRL-Kit will continue to receive updates.
Detailed training logs for the benchmark results are hosted on Google Drive at the link provided in the README, covering all nine D4RL tasks and all eight algorithms at four seeds each. These logs are available for comparison without running experiments locally.
Editorial conclusion
OfflineRL-Kit is a good starting point for researchers who want a clean PyTorch baseline for offline RL experiments and need to run or compare CQL, TD3+BC, IQL, EDAC, MCQ, MOPO, COMBO, RAMBO, or MOBILE on D4RL benchmarks. It is the wrong choice for practitioners who need online RL, continuous deployment, or a framework that abstracts away environment setup. Before using it, install MuJoCo and D4RL and verify that your gym version is between 0.15.4 and 0.24.1 as required by setup.py. The repository is MIT-licensed and the last push was on 2026-08-09.
Frequently asked questions
What are some examples of offline reinforcement learning algorithms implemented in OfflineRL-Kit?
OfflineRL-Kit implements CQL, TD3+BC, IQL, EDAC, MCQ, MOPO, COMBO, RAMBO, and MOBILE. Model-free methods include CQL and IQL, which penalize or avoid out-of-distribution actions. Model-based methods include MOPO and MOBILE, which use dynamics model uncertainty or Bellman inconsistency as additional training signals.
What is the difference between offline and online reinforcement learning?
Online RL collects new experience from an environment at training time, allowing the policy to improve through direct interaction. Offline RL trains entirely from a fixed dataset of previously collected transitions with no further environment interaction, which avoids exploration costs but introduces distribution shift problems that require algorithms like CQL or IQL to handle.
What Python and environment dependencies does OfflineRL-Kit require?
OfflineRL-Kit requires MuJoCo (from mujoco.org), mujoco-py, D4RL installed from source, and gym between versions 0.15.4 and 0.24.1. Python dependencies include torch, numpy, pandas, matplotlib, tensorboard, and tqdm. Ray is used for hyperparameter tuning but is not listed in setup.py.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/yihaosun1124-offlinerl-kit)