Library / SDK
kengz/SLM-Lab avatar
kengz/SLM-Lab

SLM-Lab: deep reinforcement learning as a set of JSON files

Modular Deep Reinforcement Learning framework in PyTorch. Companion library of the book "Foundations of Deep Reinforcement Learning".

1,361 stars289 forksPythonMIT

At a glance

What is it?
The companion library for the book Foundations of Deep Reinforcement Learning, rebuilt around Gymnasium and uv, where an experiment is a spec file rather than a training script.
Who is it for?
SLM-Lab earns its keep on reproducibility and on the spec file, not on being the fastest RL framework available. A run saves its spec and its git SHA, which is a stronger claim than a benchmark table, and the v5 line keeps the dependency story current with Gymnasium 1.3.0 and uv.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 51 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 20, 2026, and from our analysis. They are not legal advice.

Editorial analysis

An experiment is a JSON file, not a script

The core claim of SLM-Lab is that an experiment should be fully described by data. The README's feature table puts it plainly: JSON spec files fully define experiments, with no code changes needed. That is the whole architecture in one sentence, and it is the right call for a research framework, because the alternative, a training script with hyperparameters edited in place, produces results nobody can reproduce three months later.

The reproducibility claim has teeth. Every run saves its spec and its git SHA. Combined with the benchmark table that links to docs/BENCHMARKS.md, you get a setup where a number in the README can be traced back to a specific configuration and a specific commit.

The spec system is also why the framework needs so much surrounding machinery. If nothing is hardcoded, something has to validate the spec, resolve environment names, build the network architectures, and write out plots. That machinery lives in the slm_lab package, with bin/ for entry points, docs/ for the gitbook, and a test/ directory alongside it.

The three commands that matter

The quick start is short, and the middle command is the one you will actually use.

bash
uv sync
uv tool install --editable .

slm-lab run
slm-lab run --render
slm-lab run spec.json spec_name train
slm-lab run spec.json spec_name search
slm-lab --help

Running with no spec gives you PPO on CartPole, which is a sanity check rather than an experiment. Adding --render opens a visualisation. The third form is the real interface: a spec file, a spec name inside it, and a subcommand. The CLI is built on Typer, and slm-lab --help plus slm-lab run --help are the documented ways to see the rest.

Install tooling is uv rather than pip, which shows up throughout the project. pyproject.toml declares requires-python of 3.12 or newer, and the dependency split is unusually clean: the core list contains only CLI orchestration, plotting and sync libraries, with the heavy learning dependencies isolated in a dependency group named ml that includes ale-py, gymnasium with the box2d, classic-control and mujoco extras, optuna, ray and pygame.

That split has a concrete payoff, described in the README as a minimal install for a box that only dispatches remote runs and generates plots. uv sync --no-default-groups skips the ML dependencies entirely, which means a machine that orchestrates experiments does not need CUDA or a GPU driver installed to look at a training curve.

Seven algorithms and the environments they were checked against

The algorithm table is the part of the README worth reading closely, because it states what each algorithm is for and where it has been validated rather than implying everything works everywhere.

| Algorithm | Type | Best For | Validated Environments | | REINFORCE | On-policy | Learning and teaching | Classic | | SARSA | On-policy | Tabular-like | Classic | | DQN and DDQN with PER | Off-policy | Discrete actions | Classic, Box2D, Atari | | A2C | On-policy | Fast iteration | Classic, Box2D, Atari | | PPO | On-policy | General purpose | Classic, Box2D, MuJoCo (11), Atari (54) | | SAC | Off-policy | Continuous control | Classic, Box2D, MuJoCo | | CrossQ | Off-policy | Sample-efficient control | Classic, Box2D, MuJoCo |

The header claims 70 or more validated environments, and PPO carries most of that weight with 11 MuJoCo environments and 54 Atari ones. CrossQ is the newest entry and the least settled, which shows up in the v5.2.0 release notes: several commits exist purely to align CrossQ specs and reproduce tables with actual benchmark run data, including one that caps its max_frame at SAC levels for a fair comparison. Correcting your own published table is a good sign, and it is also a reminder that the numbers move.

Environments come from Gymnasium, the maintained fork of OpenAI Gym, and the README is explicit that any compatible environment works as long as you name it in the spec. Difficulty is graded by family, with classic control easy, Box2D medium, MuJoCo hard and Atari varied.

Dstack for GPUs, HuggingFace for results

The cloud story is a separate command rather than a flag, which keeps local and remote runs from sharing a code path.

bash
cp .env.example .env
uv tool install dstack
slm-lab run-remote spec.json spec_name train
slm-lab run-remote --gpu spec.json spec_name train
slm-lab pull spec_name
slm-lab list

The .env.example file is two lines, an HF_REPO pointing at a development repository with a comment to change it for a public release, and an HF_TOKEN placeholder. Results sync to HuggingFace, and slm-lab pull brings them back.

Remote runs are shaped by four files under .dstack/, one each for GPU training, GPU search, CPU training and CPU search. The distinction between train and search matters: search is ASHA, the asynchronous hyperparameter search method, run on CPU by default, while --gpu is what you add for image environments.

Two release notes give a sense of what breaks in practice. v5.3.0 upgraded gymnasium to 1.3.0 to fix a box2d-py build failure during uv sync, and added a MuJoCo Playground dependency group along with 54 PPO benchmarks. v5.1.0 added TorchArc YAML network architectures with full benchmark validation. Both are the unglamorous work of keeping a dependency tree buildable, which is most of what maintaining a research framework actually consists of.

The book and the framework have split into two timelines

SLM-Lab is the companion library for Foundations of Deep Reinforcement Learning, and the README handles the version drift directly rather than hoping nobody notices. v5.0 moved the project to Gymnasium, uv tooling and modern dependencies with ARM support, and anyone following along with the book is told to check out v4.1.1 for the code as printed.

That instruction is the most useful line in the README for a new reader. It means the repository is not one thing, and the version you install depends on whether you are learning from a book or working on current environments.

The release history supports that reading. v5.1.0 in February 2026, v5.2.0 in March, v5.3.0 in June, and a last push on 2026-08-20. The gap between 2022 and 2026 in the git history is not visible in the tags, which is a little misleading if you assume the release list is the whole story.

The repository also contains a CLAUDE.md, an AGENTS-style setup, a Dockerfile built on ubuntu 22.04 with uv installed from astral.sh, and a .dstackignore. The Dockerfile installs system packages for gymnasium including swig and the GL libraries, installs Python 3.12 through uv, and defaults its command to printing the CLI help.

Where SLM-Lab is better than writing the loop yourself

The honest comparison is not against other RL libraries, it is against the script you would otherwise write. If your experiment fits a spec and an environment Gymnasium already provides, SLM-Lab gives you plots, TensorBoard logging, saved specs and git SHAs with almost no code. The README lists automatic analysis, with training curves, metrics and TensorBoard logging out of the box.

If you need something the spec system cannot express, the abstraction becomes a cost. Custom environments, unusual loss functions and anything requiring surgical changes to the training loop all mean reading the framework source rather than your own script, and the layout with slm_lab, bin and test directories means a moderate amount of navigation.

The teaching angle is the other differentiator and the one the project is built around. An algorithm implemented inside a spec-driven framework is readable in a way that a research codebase usually is not, and the fact that REINFORCE and SARSA sit in the same table as SAC is a deliberate signal about the intended audience.

For anyone coming from Stable Baselines or CleanRL, the practical question is whether spec-driven reproducibility is worth giving up the direct control loop. If your research involves many runs across many configurations, the answer tends to be yes. If it involves one unusual environment, probably not.

Editorial conclusion

SLM-Lab earns its keep on reproducibility and on the spec file, not on being the fastest RL framework available. A run saves its spec and its git SHA, which is a stronger claim than a benchmark table, and the v5 line keeps the dependency story current with Gymnasium 1.3.0 and uv. The cost is an opinionated layout that will feel slower than writing a training loop directly, and the book and the framework have now diverged enough that the README tells book readers to check out v4.1.1. Start with slm-lab run for the PPO CartPole demo, read one existing spec under slm_lab, and only then decide whether the abstraction earns its keep for your own environment.

Frequently asked questions

What does SLM stand for in SLM-Lab?

The project describes itself as a modular deep reinforcement learning framework in PyTorch and as the companion library for the book Foundations of Deep Reinforcement Learning. The abbreviation is not expanded anywhere in the repository documentation, so it is best read as the project name rather than decoded.

How do you run a reinforcement learning experiment in SLM-Lab?

With a JSON spec file, a spec name and a subcommand, for example slm-lab run spec.json spec_name train. Running slm-lab with no spec gives you PPO on CartPole as a demo. The CLI is built on Typer, so slm-lab --help lists every command.

Which reinforcement learning algorithms does SLM-Lab implement?

Seven: REINFORCE, SARSA, DQN and DDQN with prioritized experience replay, A2C, PPO, SAC and CrossQ. The README states which environments each was validated on, with PPO carrying the widest coverage at 11 MuJoCo environments and 54 Atari environments.

Can I use SLM-Lab for environments other than Gymnasium ones?

The README states that any Gymnasium-compatible environment works, as long as you name it in the spec. Environments are grouped by family, with classic control, Box2D, MuJoCo and Atari listed as the supported categories, and the project depends on gymnasium 1.3.0 with the box2d, classic-control and mujoco extras.

Which version of SLM-Lab should I use with the book?

The book version, which the README says is v4.1.1. The current line is v5, which moved the framework to Gymnasium, uv tooling and ARM support, so the printed code and the current default branch are not the same codebase.

Official sources

  1. kengz/SLM-Lab on GitHub
  2. License: MIT
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/kengz-slm-lab.svg)](https://hysenlabs.com/projects/kengz-slm-lab)