# skrl implements three backends and its manifest describes two

> A modular reinforcement learning library for PyTorch, JAX and NVIDIA Warp, where the details worth reading are a four-dependency core with Gymnasium hard required, an optional extra group written out by hand instead of composed, and a citation that points at a 2023 paper.

**Toni-SM/skrl** — Modular Reinforcement Learning (RL) library (implemented in PyTorch, JAX, and NVIDIA Warp) with support for Gymnasium/Gym, NVIDIA Isaac Lab, MuJoCo Playground and other environments

- Repository: https://github.com/Toni-SM/skrl
- Website: https://skrl.readthedocs.io/
- Stars: 1,098 · Forks: 157
- Language: Python
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/toni-sm-skrl

## The citation is a 2023 paper and the current release is 2.1.0

The README's citation section points at a Journal of Machine Learning Research article: skrl, Modular and Flexible Library for Reinforcement Learning, volume 24, number 254, pages 1 to 9, published in 2023, authored by Antonio Serrano-Muñoz, Dimitrios Chrysostomou, Simon Bøgh and Nestor Arana-Arexolaleiba.

The release history has moved considerably since. The tags are 1.4.3 from 2025-03-30, 2.0.0 from 2026-04-08 and 2.1.0 from 2026-05-11. So there is a 2.0 release that post-dates the paper by three years, and a 2.1 on top of it.

That gap matters for two different acts that the single citation section conflates. Citing the methodology is what the paper is for, and a methods section that describes an algorithm implementation should cite it. Citing the software is a separate question, and the answer there is the release you actually ran. A review that reproduces results against 2.1.0 is not reproducing the artifact described in a 2021-vintage paper.

The project is not archived and the default branch received a commit on 2026-10-02, so this is a moving target rather than an abandoned one, which makes the distinction more important rather than less.

## The manifest description names two backends and the project ships three

The project metadata describes skrl as a modular and flexible library for reinforcement learning on PyTorch and JAX. Warp is absent from that sentence.

It is present everywhere else. The README states the library is implemented in PyTorch, JAX and NVIDIA Warp. The optional dependency groups are `torch`, `jax` and `warp`, plus an `all` group. Continuous integration runs four workflows, of which three are test runs named for the backends they cover: one for Torch, one for JAX, one for Warp.

So Warp is a first-class backend with its own test job and its own extra, and the one place a stranger would look first describes a two-backend library. That line is what package indexes, search results and dependency resolvers show, which makes it the most consequential string in the file and the one most out of date.

The rest of the metadata is consistent and modest. The version is 2.1.0, matching the newest tag. The Python floor is 3.10. The licence is MIT and the classifiers say so. The audience is declared as science and research, the topics as artificial intelligence under scientific and engineering, and the operating system as independent.

## Four core dependencies, one of which is unavoidable

The unconditional dependency list has four entries: gymnasium, packaging, tensorboard and tqdm. Everything else, including all three deep learning frameworks, is optional.

That is a genuinely small core, and the structure is sensible. Install the `torch` extra and you get `torch>=1.11`. The `jax` extra brings jax and jaxlib at 0.4.31 or later, flax at 0.9.0 or later, and optax. The `warp` extra brings warp-lang at 1.15 or later and warp-nn at 0.4 or later. A `tests` extra collects pytest, pytest-html, pytest-cov, hypothesis and pyyaml.

The gymnasium entry is the one to note. It is not optional, so it is installed whether you are running JAX on a TPU, Warp on a GPU cluster, or nothing but a unit test. For a library whose stated selling point is that you pick your own framework, the environment interface arrives whether you asked for it or not. That is a defensible choice, since Gymnasium is the common denominator across the supported environments, and it is also the dependency most likely to be unwanted in a locked-down environment.

The presence of hypothesis in the tests extra is the other detail worth recording. Property-based testing is not a default choice, and its presence says something about how the test suite is written.

## The combined extra repeats the others by hand instead of composing them

The `all` extra is not built from the other three. It is a fourth, independent list containing the same entries again: torch at 1.11 or later, jax and jaxlib at 0.4.31 or later, flax at 0.9.0 or later, optax, warp-lang at 1.15 or later, and warp-nn at 0.4 or later.

Duplication in a manifest like this is not stylistic. Raising the torch floor in the `torch` extra does not raise it in `all`, so the two extras can disagree, and the only symptom would be an environment that installed one and not the other. The same applies to every entry in all three lists. Nothing enforces that they stay aligned.

The floors themselves are worth comparing. `torch>=1.11` is a much older release than jax at 0.4.31 or warp-lang at 1.15, and the three backends are therefore claiming compatibility with quite different eras of their respective ecosystems. That may be accurate, since a thin wrapper over a framework's array API may not need the latest framework, but it means the three extras are not interchangeable statements about what is supported.

A separate small detail sits in the tool configuration: codespell is configured to skip `pyproject.toml`, alongside the docs static and build directories. The file that carries the two-backend description is therefore outside the spell checker's coverage, which is a reasonable exclusion for a manifest full of URLs and is also the file where that sentence lives.

## Continuous integration runs the same suite three times, once per backend

Four workflows are linked from the README badge row. One is a pre-commit check. The other three are test runs named for Torch, JAX and Warp.

Running the same test suite once per backend is the whole point of supporting three, because the failure mode of a multi-backend library is a test that passes on one and fails on another. A backend that is only compiled and imported is not a supported backend. Three jobs is the minimum that makes the claim checkable, and having them visible as badges rather than buried in a status page is the useful part.

The hooks side is configured in the repository rather than left to each contributor. A pre-commit configuration file sits at the top level, and one of the four workflows enforces it in CI, which means the same formatting and lint rules apply to a local commit and to a pull request.

The package definition also constrains what gets shipped. A setuptools find directive includes only packages matching `skrl*`, so the `docs/`, `examples/` and `tests/` directories at the top level are not installed. That is the right default, and it is worth noticing because the example tree is large enough to be shipped by accident under a looser rule.

Formatting is handled by Black with a line length of 120 and the docs directory excluded.

## Examples are organised by environment, and one supported environment has none

The example tree is keyed to where the environments come from, not to which algorithm is being demonstrated. There are directories for Gym, Gymnasium, Isaac Lab, ManiSkill, MuJoCo Playground, real-world hardware, shimmy and shared utilities, plus a shell script at the top of the tree.

That is a coherent choice for a library whose hard part is the environment interface, and the two non-obvious directories are the interesting ones. The real-world directory implies running an agent against actual hardware rather than a simulator. The shimmy directory implies compatibility with older Gym environments through an adapter layer, which is what you would expect given that both OpenAI Gym and Gymnasium are named as supported.

The gap is PettingZoo. It is named in the README alongside Gym, Gymnasium and ManiSkill as an environment interface the library supports, and it is the one supported interface with no directory in the example tree. The multi-agent side of the library is therefore the part with the least worked material next to it.

The sibling of the real-world directory is the scope mechanism, which is the one genuinely distinct idea in the description: an agent can train simultaneously across scopes, which are subsets of the available environments that may or may not share resources, all within a single run. That is a different axis from parallelism over a single environment, and it is the feature to look at first if you are evaluating the library.

## The default branch is develop and the README asks you to use it

The repository's default branch is `develop`, not `main`, and the README carries a note saying the project is under active continuous development, asking readers to make sure they have the latest version and pointing at the develop branch or its documentation for the latest updates that will be released.

That is a candid instruction and an unusual default. Most projects keep the default branch as the stable line and put development elsewhere. Here the development branch is what you land on, and the documentation build has a matching path for it alongside the stable one, so the rendered docs can be read per branch.

The practical effect is that two users who both say they are using skrl may be months apart. One installed 2.1.0 from a package index in May 2026. The other cloned develop, which received a commit on 2026-10-02, and whose manifest still reads 2.1.0 because unreleased work does not bump the version. The about-equivalent version string is identical in both cases.

The readme also carries a machine-readable citation for the library, a link to a Hugging Face organisation alongside the package index badge, and links to the documentation, the discussions area and the issue tracker. None of that changes the version question, which is the one to settle before comparing anyone's results.

## Conclusion

The design here is the interesting part: a small core, a backend chosen at install time rather than at import time, and a scope mechanism that lets several environment subsets train in one run. Two things to check before you build on it. The library's citation is a 2023 paper and the current release is 2.1.0 from May 2026, so cite the release and the paper as different things. And the default branch is develop, which the README asks you to use for the latest code, meaning the version you get depends on whether you install from a release or clone. If you maintain a fork, note that the combined extra repeats the per-backend extras by hand.

## FAQ

### What is skrl and which frameworks does it implement?

It is an open-source modular reinforcement learning library written in Python, focused on modularity, readability, simplicity and transparency of algorithm implementation. The README states it is implemented in PyTorch, JAX and NVIDIA Warp, with each available as a separate optional dependency extra and each covered by its own test workflow in CI.

### How do I install skrl for a specific framework?

The core install pulls only gymnasium, packaging, tensorboard and tqdm. A backend is added with an optional extra: torch, jax, warp or all. The test tooling is a fifth extra called tests, which brings pytest, pytest-html, pytest-cov, hypothesis and pyyaml. Python 3.10 or later is required.

### Which environments does skrl support?

OpenAI Gym, Farama Gymnasium, PettingZoo and ManiSkill among other environment interfaces, plus loading and configuring NVIDIA Isaac Lab and MuJoCo Playground environments. The example tree ships directories for Gym, Gymnasium, Isaac Lab, ManiSkill, MuJoCo Playground, real-world hardware, shimmy and shared utilities, with no PettingZoo directory among them.

### What does training by scopes mean in skrl?

It is the multi-agent capability described in the README: agents can train simultaneously by scopes, which are subsets of environments among all available environments, and those subsets may or may not share resources, all within the same run.

### Should I use the develop branch or a released version of skrl?

The repository's default branch is develop, and the README asks readers to visit that branch or its documentation for the latest updates to be released. The newest tag is 2.1.0 from 2026-05-11, while the manifest on develop still reads 2.1.0, so the version string does not distinguish released code from unreleased code.

### How should I cite skrl in a publication?

The README points at a Journal of Machine Learning Research paper from 2023, volume 24, number 254, pages 1 to 9, by Serrano-Muñoz, Chrysostomou, Bøgh and Arana-Arexolaleiba. That describes the methodology rather than a specific release, and the library has since had a 2.0 release in April 2026 and 2.1 in May 2026.

## Sources

- [License: MIT](https://github.com/Toni-SM/skrl/blob/develop/LICENSE)
- [Project website](https://skrl.readthedocs.io/)
- [README](https://github.com/Toni-SM/skrl/blob/develop/README.md)
- [Releases](https://github.com/Toni-SM/skrl/releases)
- [Toni-SM/skrl on GitHub](https://github.com/Toni-SM/skrl)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/toni-sm-skrl
