AgileRL reads as one library and installs as two, and platform markers decide what you get
Streamlining reinforcement learning with RLOps. State-of-the-art RL algorithms and tools, with 10x faster training through evolutionary hyperparameter optimization.
At a glance
- What is it?
- AgileRL bundles LLM post-training, evolutionary hyperparameter tuning and classic deep RL into one Python package, but the install extras are split by platform and by package, and the two published wheels have to reach the index in a fixed order. The code is one library. The installation is not.
- Who is it for?
- Adopt AgileRL when you want SFT, DPO and agentic RL plus evolutionary hyperparameter tuning under one dependency tree, and when you are on Linux with NVIDIA hardware, since the two libraries that make local quantized runs practical carry Linux-only markers.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository received new commits within the last day.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 4, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Two speed claims sit in different places, and the benchmark gives no numbers
Two different speed claims apply to the same package. The repository description promises 10x faster training through evolutionary hyperparameter optimization. The README benchmark section promises something narrower: over 4x more tokens per second than ART and TRL, and higher reward on half the GPU memory. Neither figure can be recomputed from the repository.
What the benchmark section does give is the setup wrapped around the claim. AgileRL's CISPO is compared with ART and TRL on the GEM Sudoku Hard task, at 32k-token context and up to 50 turns per rollout. AgileRL ran on A100 40GB nodes while ART and TRL needed A100 80GB. The async and hyperparameter runs happened on Arena rather than on local hardware, and all three systems started from the same hyperparameters. No throughput or reward values appear in the prose at all, so the only quantitative comparison a reader can extract is about card memory and turn count.
That distinction matters when the two claims get quoted together. One is a claim about wall-clock training speed attributed to evolutionary tuning, the other is a claim about tokens per second and reward for a single agentic algorithm on a single puzzle task.
The core dependency list is the classic RL stack, whether or not you want it
pip install agilerl is described as classic RL plus the Arena CLI, yet the base dependency list is what a post-training user also ends up with. Before any extra applies, agilerl declares accelerate, gymnasium, hydra-core, omegaconf, flatten_dict, dill, h5py, google-cloud-storage, minari[all], SuperSuit, pettingzoo, mpe2, redis, wandb, pygame-ce, pymunk and jax[cpu]. A CPU build of JAX sits in that base list rather than in an optional group, and SuperSuit, PettingZoo and mpe2 are multi-agent tooling with no role in an LLM run. The extras begin only after all of that:
pip install "agilerl[llm]" # LLM post-training
pip install agilerl # classic RL + the Arena CLIExtras are additive, so the post-training path sheds none of the classic stack. agilerl-arena is a core dependency at agilerl-arena>=1.0.0,<2.0, which means the hosted Arena client and its CLI arrive even if you never sign in. That upper bound is also the ceiling on the whole install: arena 2.0 cannot be picked up until agilerl itself is released. The pettingzoo pin carries a comment worth reading, since the MPE environments were split into their own package at pettingzoo 1.25 and now arrive as mpe2>=1.1.0.
Platform markers make the cpu-llm extra smaller than the extras table implies
The extras table calls agilerl[cpu-llm] the same Hugging Face stack as agilerl[llm], minus vLLM. The packaging metadata narrows that claim. Inside the cpu-llm group, bitsandbytes==0.49.2 and liger-kernel==0.8.2 both carry sys_platform == 'linux' markers, while transformers==5.14.1, peft==0.21.0, datasets==5.0.1 and openenv>=0.4.1,<0.5 carry none. On a Mac or a Windows box the group resolves to the transformers stack with no quantization library and no fused kernels, which are the same two components named when the README explains that QLoRA lets bigger models fit on a colocated local run.
The same marker style shows up in the physics extra, where swig>=4.4.1 and gymnasium[box2d] are both gated on sys_platform != 'win32'. On Windows the agilerl[box2d] extra installs neither package: the name resolves, the dependency never appears. The cpu extra takes the opposite route and pins torch>=2.11.0,<2.12, with an inline comment explaining that this matches the 2.11 line of the default Linux CUDA wheel while the CPU index also publishes newer torch. So the platform you are on decides both which optional packages exist and which torch version you land on.
One YAML manifest, and an agent path that never opens a browser
Every hosted run starts from a YAML manifest, and the CLI around it is short:
arena login
arena models list # supported base models
arena experiments submit my_manifest.yaml --project my-project
arena agent deploy my-experiment # deploy the best checkpoint
arena agent run my-deployment
arena agent generate --prompt "What is 17 * 23?"The same sequence runs without a terminal. A personal access token stands in for the interactive login, the manifest schema can be printed, and validation returns one JSON verdict with stable exit codes:
export ARENA_API_KEY="arena_pat_..." # no interactive login
arena manifest schema # JSON Schema the agent writes against
arena manifest validate my_manifest.yaml --json # one JSON verdict, stable exit codes
arena experiments metrics my-experiment # download metrics to judge the runThe README points Cursor, Claude Code or Codex at precisely this loop: write a manifest, check it, launch it, pull metrics, deploy. Its last sentence in that walkthrough breaks off mid-clause right after naming ArenaClient as the Python equivalent, so the SDK call shape has to come from the documentation site rather than from this page. Note what leaves your machine in the process. Arena executes the manifest on managed GPU clusters, so async rollouts, population-based tuning across nodes and support for base models up to Nemotron 3.5 Super VL 120B-A12B are hosted capabilities with a changing model list, not local flags.
Publishing two packages in dependency order out of one shared dist directory
Build and publish live in a justfile, and its comments name a hazard most projects never write down. This is a uv workspace, so every uv build writes to the workspace-root dist/ directory and both packages' artifacts land side by side. They are then separated by name: dist/agilerl_arena-* selects the Arena SDK, dist/agilerl-[0-9]* selects the core package. The character between the two names is what decides that split, since both prefixes share the agilerl string.
Upload order is fixed and explained. The arena recipe has to reach the index first, because agilerl depends on agilerl-arena and the combined release is briefly uninstallable if the dependency lands second. Both upload recipes depend on check-dist, which runs check-extras and then build and finishes with a dry run against https://pypi.org/simple. check-extras fails the publish when a workspace pin drifts or when the all extra fails to list one of the extras it should carry.
The --check-url flag carries a second consequence the comment spells out: the two packages are versioned independently, so a release often reuses one unchanged version, and the flag skips anything already present on the index. Practically this means a green just publish can upload one wheel and deliberately skip the other.
Local LLM training shares one distributed config and one generation engine
A local LLM run is described with the same primitives as the classic algorithms. One FSDPConfig object is reused across every LLM algorithm, launched with torchrun, over either data parallel or FSDP2 sharding. Memory is attacked from several directions at once: chunked fused log-probs, activation checkpointing and offload, CPU offload for optimizer state or parameters, and trainer offload during rollout. Generation runs through vLLM, with the option to colocate the trainer and vLLM on a single GPU, which is the arrangement the QLoRA recommendation for fitting larger models locally depends on.
Kernels are selected per model family across dense, MoE and hybrid Mamba architectures. MoE LoRA runs per expert without materializing full tensors, and hybrid models such as Nemotron-H receive fused kernels and fixes for Mamba2. Environments arrive through OpenEnv in three shapes: a Python class, a library entrypoint such as GEM, or a remote env server, and a dataset plus a reward function substitutes for an environment entirely. Evolutionary hyperparameter tuning then adjusts the learning rate and the KL penalty during the run rather than before it, which is the mechanism behind the tuning claim in the repository description.
Exact pins beside wide ranges, Python 3.10 to 3.13, and two releases on one day
Version policy mixes hard pins with generous ranges. tensordict is held at exactly 0.13.0. The cpu-llm group pins bitsandbytes==0.49.2, datasets==5.0.1, liger-kernel==0.8.2, peft==0.21.0 and transformers==5.14.1, and vLLM is pinned at 0.25.1. Elsewhere the ranges are wide: numpy>=2.0.0,<3.0, redis>=4.4.4,<8.2.0, gymnasium>=1.0.0, accelerate>=1.7.0, wandb>=0.18.0, pymunk>=6.2,<7.4. hydra-core and omegaconf sit in narrow minor bands while flatten_dict sits beside them for config plumbing. Supported Python is >=3.10,<3.14. A uv.lock file, a pre-commit configuration and a justfile share the repository root with a sitecustomize.py, which Python imports at interpreter start whenever the repository root lands on the path, as it does for an editable install from a clone.
The release titles say what actually changed. v2.42.0, published on 2026-10-02, folded vLLM into agilerl[llm] and led the README with LLM post-training, so the extra list and the first screen of the README both moved that day. v2.42.1, published the same day, re-tied the output head and recomputed RoPE after a meta load, which is a load-time correctness fix rather than a change to the training loop. agilerl-arena v1.10.0 went out on 2026-10-01 and added gpt-oss-20b with flex attention and packed-expert LoRA. The last push landed on 2026-10-02 and the repository is not archived.
Editorial conclusion
Adopt AgileRL when you want SFT, DPO and agentic RL plus evolutionary hyperparameter tuning under one dependency tree, and when you are on Linux with NVIDIA hardware, since the two libraries that make local quantized runs practical carry Linux-only markers. Before installing anything else, run arena models list to confirm your base model is served, then read the extras in pyproject.toml yourself: the box2d extra installs nothing on Windows, and the plain agilerl install already carries jax[cpu], minari, SuperSuit and the Arena client. Judge the speed claims on your own hardware, because the comparison against ART and TRL gives no throughput or reward figures in text.
Frequently asked questions
In AgileRL, what does a plain agilerl install pull in that an LLM post-training user never touches?
The base list carries jax[cpu], minari[all], SuperSuit, pettingzoo, mpe2, redis, wandb, google-cloud-storage and h5py before any extra is applied, plus agilerl-arena>=1.0.0,<2.0 as a required dependency. Extras are additive, so agilerl[llm] keeps all of it.
Which parts of the AgileRL Hugging Face stack are missing on macOS or Windows?
Inside the cpu-llm group, bitsandbytes==0.49.2 and liger-kernel==0.8.2 are marked sys_platform == 'linux', while transformers, peft, datasets and openenv are unmarked and install anywhere. The quantization library and the fused kernels therefore drop out on those platforms.
Does the agilerl[box2d] extra install anything on Windows?
No. Both entries in that group, swig>=4.4.1 and gymnasium[box2d]>=1.0.0, carry sys_platform != 'win32' markers, so the extra resolves to nothing on Windows even though the name resolves fine.
How does AgileRL publish its two packages without breaking installs?
The justfile uploads agilerl-arena before agilerl, because agilerl depends on agilerl-arena and the release would be briefly uninstallable otherwise. Both packages are built into one shared workspace-root dist/ directory and then selected by name, and the two are versioned independently, so --check-url skips a version already on the index.
Which algorithms does the AgileRL LLM post-training path cover?
Supervised fine-tuning and DPO come first, followed by reinforcement learning with GRPO, CISPO, GSPO, REINFORCE or PPO. LoRA and QLoRA are available for training adapters on a frozen base.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/agilerl-agilerl)