# OfflineRL-Kit gives you nine offline RL algorithms and a ReplayBuffer

> OfflineRL-Kit is a pure PyTorch offline reinforcement learning library that puts CQL, TD3+BC, IQL, EDAC, MCQ, MOPO, COMBO, RAMBO and MOBILE behind the same components, plus Ray tuning and a four-file logger. Its benchmark table is labelled ongoing, and several of its model-based results have a standard deviation as large as the score it is reporting.

**yihaosun1124/OfflineRL-Kit** — An elegant PyTorch offline reinforcement learning library for researchers.

- Repository: https://github.com/yihaosun1124/OfflineRL-Kit
- Stars: 393 · Forks: 49
- Language: Python
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/yihaosun1124-offlinerl-kit

## Four seeds, and error bars as wide as the scores

The benchmark table is the most useful thing in the repository, partly because of what it shows and partly because of what it quietly admits.

It is labelled four seeds and ongoing. Nine environments across three difficulty bands, eight algorithms, every cell reported as a mean with a spread.

Some cells are tight and the ranking is trustworthy. On halfcheetah-medium, the model-free results cluster between the high forties and the mid sixties with spreads under one and a half, while the model-based results sit higher, with the adversarial model-based variant leading at 78.7.

Other cells are not tight at all. On hopper-medium-expert, one model-based entry reports 74.6 with a spread of 44.2. On walker2d-medium-expert, another reports 78.4 with a spread of 45.4. A number that varies by more than half its own value across four seeds is not evidence that one algorithm beats another on that environment, whatever the table's best column says.

The ensemble-diversified model-free variant and the model-bellman one are the most consistent entries across the whole table, with spreads mostly under two. If you are choosing by stability rather than by peak score, the table already points at them.

Read as a whole, the table is doing what a research library's table should do: showing you the spread so you can see which comparisons survive it.

## Two critics, a squashed Gaussian, and a misspelled keyword

The quick start is a single algorithm assembled from named pieces, and the composition tells you what the library is really for.

The data path is an environment, the dataset loader, and a replay buffer constructed with an explicit shape and dtype for observations and actions plus a device. Nothing is inferred.

The models are three multilayer perceptrons. The actor takes only the observation; each critic takes the observation concatenated with the action, and there are two of them. Clipped double Q is the standard conservative trick and here it is expressed as two critic objects rather than a list.

The policy distribution is a squashed diagonal Gaussian with two switches: one to allow an unbounded support, one to make the standard deviation a learned function of the state rather than a free parameter. That constructor is the thing you would replace if you were implementing a different actor.

The policy object then takes the actor, both critics, all three optimisers, the action space, the usual target and discount arguments, and the algorithm's own hyper-parameters. The last of those is where a small blemish sits: the keyword is spelled with the two letters transposed, and it is spelled that way in the constructor call you are meant to copy.

That is the whole point of this library in miniature. You can read every layer between the dataset and the loss, and where the author left a typo, you can see it, fix it in your copy, and know what you changed.

## Installing means building D4RL and coupling a physics version

The installation section has three ordered steps, and none of them is a single command.

First, a physics engine, downloaded from its own site, plus a Python binding whose version has to match whichever engine version you installed. The instruction is explicit that the binding's version depends on the engine, which means there is no combination that is simply current.

Second, the benchmark datasets, cloned from a separate repository and installed in editable mode:

```shell
git clone https://github.com/Farama-Foundation/d4rl.git
cd d4rl
pip install -e .
```

Third, the library itself:

```shell
git clone https://github.com/yihaosun1124/OfflineRL-Kit.git
cd OfflineRL-Kit
python setup.py install
```

That last command is the problem. Direct setup script installation was deprecated years ago and removed from current packaging tooling, so on a modern environment this step fails or warns rather than working.

The editable install used for the dataset dependency is also worth noticing, because it means the benchmark code is not a released package but a checkout on your machine. Any behaviour you get from the datasets is behaviour of whatever commit you cloned.

None of this is unusual for a research library from a few years ago. It is the cost of the pure PyTorch design: no wheel, no bundled datasets, no vendor abstraction over the engine.

## A malformed requirement and a dependency commented out

The packaging metadata is small enough to read line by line, and it contains two things that will cost you time.

The first is the gym requirement. It is written as a single string containing a comma, rather than a version specifier, which means it is not a parseable requirement expression and will not behave the way it looks like it should. The floor it expresses is old, and the ceiling is also old, so even read charitably the library is asking for a gym generation that predates the current successor.

The second is Ray. The dependency is present in the file and commented out, while the tuning documentation tells you to use Ray for hyper-parameter search. So the tuning example depends on something the install does not pull in.

Three more facts from the same file. The version is 0.0.1, which is consistent with there being no published releases at all: the versions you will find are not on the package index but in the repository's own history. The declared runtime dependencies are the ordinary set, matplotlib, numpy, pandas, torch, tensorboard and tqdm. And the author and maintainer fields point at an academic address rather than a personal one.

At the repository level there is no separate requirements file, no test directory and no continuous-integration configuration among the top-level entries. The library presents two example directories instead, one for training runs and one for tuning.

## Ray tuning runs at half a GPU per trial

The tuning section is short and the resource line is the interesting part.

It calls for Ray, initialises it, loads the default arguments, then declares a grid search over two axes: the proportion of real data in the buffer, taking two values including a small one, and the seed, taking two values. Then it runs the experiment function and hands Ray a resource specification.

That specification is half a GPU per trial. For a library whose algorithms are small multilayer perceptrons and whose datasets are modest, half a GPU is a sensible allocation rather than an aspirational one, and it means several configurations share one device.

The named example tunes a model-based algorithm, and the full script is one of the two example directories in the repository, so the pattern is meant to be copied rather than reconstructed.

Two observations about the design. The grid is declared as configuration data rather than as a nested loop, which is what makes it reproducible: the same axes produce the same set of trials. And seeds being one of the grid axes rather than a loop variable means the sweep reports variance across seeds as part of its output, which is the same instinct that produced the spreads in the benchmark table.

The compatibility claim of parallel tuning is therefore modest and concrete: it is Ray, a grid search, and half a GPU per trial.

## Four log files because an experiment outlives its terminal

The logging system is described as clear and powerful, and the specification is a list of file types with a reason for each.

A plain text file is a backup of everything printed to standard output, so you have the console transcript even if the process died. A comma-separated file holds the training progress, meaning loss and performance metrics rather than log lines. A TensorBoard event file holds the curve you actually plot. And a JSON file holds the hyper-parameters, as a backup.

The directory naming is keyed by three things plus the full argument namespace: the task, the algorithm name and the seed, along with the arguments themselves. That is a naming scheme you can sort, and sorting by seed is what you want when you are comparing across four of them.

The constructor takes a small output configuration rather than a boolean set of flags, mapping three sinks to their type, and there is a separate call that records the hyper-parameters into the logger before training starts.

Then the trainer object receives the policy, an evaluation environment, the buffer and the logger, along with the epoch count, steps per epoch, batch size and evaluation episode count, and one call starts training.

The design point is that nothing about the pipeline is implicit. There is no global state, no hidden default directory and no auto-detected device. Every choice is an argument, which is what makes the library readable by someone auditing a result six months later.

## One trainer shape, two algorithm families

The scalability claim is that you can build a new algorithm in a few lines from the components, and the way to evaluate that claim is to look at what the components actually are.

There are five visible: a multilayer perceptron, a distribution class, an actor wrapper that pairs a backbone with a distribution and a device, a critic wrapper, and the replay buffer. On top of those sit the policy classes and the trainers.

The trainer in the example is named for the model-free family. That naming implies a parallel one for the model-based family, which is consistent with the nine algorithms being split into two groups of five and four.

The split is also a statement about what the algorithms share. Every model-free entry here is an actor-critic method with a conservative or distribution-correcting modification, so they differ in the loss and in how the policy is regularised, not in the network shapes. That is what lets one trainer and one buffer serve all of them.

The model-based entries add a learned dynamics or adversarial component, and those results are also the ones with the widest spreads in the benchmark table, which is a reasonable place to be careful.

The project is also pointing somewhere. The README promotes a separate repository for reinforcement learning on vision-language-action models, and this one is the general-purpose predecessor.

## Against an installable offline RL library

The realistic alternative is a library that installs from a package index and does not ask you to clone a dataset repository first. That is the comparison a new user will actually make, and the difference is not algorithmic.

The algorithms here are the standard conservative and distribution-correcting set, each with its own paper linked from the README, and a reader who knows offline reinforcement learning will recognise all nine. Nothing here is proprietary or unusual. What this library offers is visibility: every network, every distribution, every hyper-parameter and every log destination is an argument in front of you, and the quick start is the training script.

Against that, a packaged library gives you working environment and dataset integration, pinned dependency resolution, and a current gym generation. You give up the ability to read the trainer in one sitting, and you accept the framework's own opinions about how an actor, a critic and a buffer should be shaped.

The specific costs here are named rather than hidden. A physics engine and a binding that must agree on version. A dataset repository cloned in editable mode. A packaging command that current tooling no longer supports. A dependency commented out that the tuning example needs. And a package version of 0.0.1 with no releases behind it.

For a paper you need to re-run in a year with the same seeds, that list is a nuisance. For a codebase you intend to extend, the same list is why you would choose it.

## Conclusion

OfflineRL-Kit suits a researcher who wants to read an algorithm's policy class, compose it from visible components, and sweep hyper-parameters across seeds without adopting a framework's abstractions. It does not suit someone who wants a clean install today, because the documented path clones D4RL from source, couples a physics engine version to a binding, and finishes with a retired packaging command. Verify two things before you start: that the model-based results you care about are separated by more than their own error bars in the benchmark table, because several are not, and that your gym and physics engine versions sit inside the range the install instructions assume. The package version is 0.0.1, the licence is MIT, and the last push on the repository was on 2026-08-09.

## FAQ

### Which offline reinforcement learning algorithms does OfflineRL-Kit implement?

Five model-free entries: CQL, TD3+BC, IQL, EDAC and MCQ. Four model-based entries: MOPO, COMBO, RAMBO and MOBILE. Each is linked to its paper from the algorithm list in the README.

### How do I install OfflineRL-Kit?

In three steps: install the MuJoCo engine and the matching Python binding, clone the D4RL dataset repository and pip install it in editable mode, then clone this repository and install it. The last step uses a direct setup script install command that current packaging tooling has retired, so expect friction on a modern environment.

### How do I tune a hyper-parameter in OfflineRL-Kit?

With Ray. Declare a grid search over the axes you care about, such as the proportion of real data and the seed, and pass a resource specification per trial; the documented example asks for half a GPU. A full tuning script is included in the repository's tuning example directory.

### What log files does the OfflineRL-Kit logger produce?

Four types. A text file backing up standard output, a comma-separated file of training progress such as loss and performance, a TensorBoard event file for the training curve, and a JSON file backing up the hyper-parameters. Log directories are named from the task, algorithm and seed.

### How reliable are the OfflineRL-Kit benchmark numbers?

They are reported as four seeds and the table is labelled ongoing, and the spreads vary a lot. Several model-based cells report a standard deviation larger than half the mean, such as 74.6 with a spread of 44.2 on one expert environment, so those rows do not separate the algorithms.

## Sources

- [Issues](https://github.com/yihaosun1124/OfflineRL-Kit/issues)
- [License: MIT](https://github.com/yihaosun1124/OfflineRL-Kit/blob/main/LICENSE)
- [README](https://github.com/yihaosun1124/OfflineRL-Kit/blob/main/README.md)
- [yihaosun1124/OfflineRL-Kit on GitHub](https://github.com/yihaosun1124/OfflineRL-Kit)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/yihaosun1124-offlinerl-kit
