Open-AgentRL: Three ICML 2026 Reinforcement Learning Frameworks for LLM Agents
RLAnything (ICML 2026) & AutoTool (ICML 2026), DemyAgent: Open-Source RL for LLMs and Agentic Scenarios
At a glance
- What is it?
- Open-AgentRL is a Python repository from Gen-Verse that bundles three ICML 2026 research frameworks for training LLM-based agents with reinforcement learning: RLAnything for closed-loop joint optimization of policy, reward model, and environment; DemyAgent for studying which data and algorithm choices drive performance; and AutoTool for teaching agents to generalize tool selection to unseen tools. The underlying infrastructure is the verl distributed training engine, installable via pip.
- Who is it for?
- Researchers who want to reproduce the RLAnything or DemyAgent ICML 2026 results will find this repository ready to use; the 3K SFT and 30K RL datasets for DemyAgent are already public on HuggingFace. AutoTool replication is incomplete until the full 200K tool-use dataset is released.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 111 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Three Research Projects in One Repository
Open-AgentRL is not a single tool. It is a monorepo that houses three distinct research projects, all accepted to ICML 2026, sharing the verl distributed training infrastructure as their base. The package installs under the name verl (as declared in pyproject.toml) and requires Python 3.10 or newer.
RLAnything addresses a design problem in standard RL pipelines: policy, reward model, and environment are typically trained independently, which limits what each can learn from the others. DemyAgent takes a different angle and asks which practical choices in data, exploration algorithms, and reasoning modes actually matter for agentic RL, rather than proposing new architecture. AutoTool moves past the fixed-toolset assumption and trains agents that can select from a library of 1,346 tools, including tools not seen during training.
The three projects share a codebase directory tree (train/, reward/, configs/, examples/, etc.) and a common dependency set, but each has its own entry points, configuration files, and released artifacts. Teams evaluating the repository should decide which of the three projects they intend to use before installing, because the dependency surface is large and some components such as sglang are optional extras pinned to specific versions.
RLAnything: Closing the Loop Between Policy, Reward Model, and Environment
Standard reinforcement learning for LLMs treats the policy, the reward model, and the training environment as three separate components that are fixed at the start of a run. RLAnything proposes a closed-loop design where all three are updated together during training.
The policy receives feedback from both outcome signals (whether the final answer is correct) and step-wise signals from the reward model, which the paper reports is better than using outcome signals alone. The reward model is not frozen; it is jointly optimized through consistency feedback during training, which in turn improves the quality of signals the policy receives. The environment adapts automatically based on critic feedback from both the reward model and the policy, a mechanism the authors describe as theory-motivated automatic environment adaptation.
According to the README, step-wise signals from the optimized reward model outperform outcome signals that rely on human labels. This claim is backed by ICML 2026 peer review, though independent reproduction of the full benchmark numbers requires downloading the released policy and reward model checkpoints from HuggingFace (collection: Gen-Verse/open-agentrl).
DemyAgent: What Data Quality and Algorithm Choices Actually Move the Needle
DemyAgent investigates three dimensions of agentic reinforcement learning: data, algorithms, and reasoning modes. Its central contribution is empirical: the README summarizes which choices matter and by how much, based on experiments the authors ran before ICML 2026 submission.
On data: real end-to-end trajectories and high-diversity datasets significantly outperform synthetic alternatives. The team released 3K SFT and 30K RL training examples on HuggingFace (collection: Gen-Verse/open-agentrl-68eda4c05755ca5a8c663656) to support reproducibility.
On algorithms: reward clipping and entropy maintenance are identified as exploration-friendly techniques that boost training efficiency. These are standard RL modifications, but the paper provides empirical evidence that they matter specifically in agentic settings.
On reasoning modes: deliberative reasoning with selective tool calls outperforms both frequent tool invocation and verbose self-reasoning. The practical implication is that model size matters less than training recipe: the DemyAgent-4B model, trained on the released dataset, outperforms 32B models on AIME2024/2025, GPQA-Diamond, and LiveCodeBench-v6 according to the README. That claim depends on the specific evaluation configuration, which the repository's eval scripts document.
AutoTool: Generalizing Tool Selection to 1,346 Tools
Existing agentic RL work typically fixes the toolset at training time and does not address what happens when the agent needs a tool it has never seen. AutoTool constructs a 200K tool-use trajectory dataset spanning 1,346 tools and 120 task types across math, science, search QA, code generation, and multimodal reasoning.
Training uses a dual-phase pipeline. Phase I stabilizes tool-integrated reasoning trajectories using both SFT and RL. Phase II refines multi-step tool selection with a KL-regularized Plackett-Luce ranking objective, which directly optimizes ordering within large tool lists rather than independent binary classification.
The model is trained on only 460 of the 1,346 tools. At inference time it generalizes to the full 1,346-tool library, including 886 tools it never saw in training. The README reports benchmark gains of +6.4% on math and science tasks, +4.5% on search QA, +7.7% on code, and +6.9% on multimodal tasks compared to the baselines it was evaluated against.
The training code and example data format for the dual-phase pipeline are available in the autotool/ directory. The full 200K dataset was not yet published as of the last push on 2026-06-12; the README describes it as coming soon.
Installing Open-AgentRL and Running the First Experiment
The repository requires Python 3.10 or newer, as specified in pyproject.toml. The core dependencies are listed in requirements.txt and include accelerate, datasets, hydra-core, ray[default], tensordict, transformers, wandb, and vllm (via optional extras). Flash-attn is listed in requirements.txt and requires a CUDA-capable GPU; there is no CPU fallback.
To install all development dependencies at once:
pip install -r requirements.txtAlternatively, to install only the core verl package, which pulls in the required dependencies declared in pyproject.toml:
pip install -e .The examples/ directory contains ready-to-run trainer scripts organized by algorithm: examples/grpo_trainer/, examples/ppo_trainer/, examples/reinforce_plus_plus_trainer/, and others. Each subfolder contains configuration files and a script that can be launched directly. Ray must be running before starting a distributed training job; the scripts in examples/ray/ show the cluster setup.
Users running on a single GPU can attempt single-process training by bypassing the Ray launcher, but this is not the intended workflow and the README does not document a single-GPU path explicitly.
Where Open-AgentRL Falls Short
The most significant gap is data availability. AutoTool's 200K training dataset was not released as of the last push on 2026-06-12; the README describes the release as coming soon. Anyone attempting to reproduce AutoTool's reported benchmark numbers cannot do so without that data.
The dependency stack is heavy. The requirements.txt includes flash-attn, which must be compiled or installed from a pre-built wheel matching your CUDA version and Python version exactly. The ray[default] requirement means distributed setup adds operational overhead. Teams without existing Ray infrastructure should budget time for cluster configuration before the first training run.
The project has no GitHub releases. Versioning follows the main branch directly, so there is no stable tag to pin in a production environment. The pyproject.toml version field is read dynamically from a file, which means pip install from a commit hash is the only reproducible install path.
Finally, the README does not provide end-to-end benchmarking scripts that independently verify the reported numbers; the harness-bench results referenced in the DemyAgent section point to a separate evaluation repository (Gen-Verse/second-brain-evals), which teams should read before accepting the claims at face value.
Comparing with TRL and Considering the License
TRL (Transformer Reinforcement Learning) from HuggingFace is the standard library for applying RL to language models and covers PPO, DPO, GRPO, and several other methods through a trainer API that integrates with HuggingFace Datasets and Accelerate. TRL is designed for single-machine or simple distributed training and does not natively address Ray-based multi-node orchestration or the joint multi-component optimization that RLAnything introduces.
Open-AgentRL, built on verl, is intended for GPU clusters and adds the joint policy-reward-environment loop, the agentic trajectory datasets, and the tool-selection ranking objective that TRL does not cover. The tradeoff is complexity: verl's Ray-based architecture requires more infrastructure knowledge to operate, while TRL's trainer API works with a standard HuggingFace training loop.
Both projects are licensed under Apache 2.0, which permits commercial use, modification, and redistribution with attribution. The CONTRIBUTING.md file in Open-AgentRL documents the contribution process for those who want to extend the framework.
Editorial conclusion
Researchers who want to reproduce the RLAnything or DemyAgent ICML 2026 results will find this repository ready to use; the 3K SFT and 30K RL datasets for DemyAgent are already public on HuggingFace. AutoTool replication is incomplete until the full 200K tool-use dataset is released. Engineers who need a distributed RL training framework for LLMs and can run a Ray cluster will find the verl infrastructure practical. Those on a single consumer GPU should verify flash-attn and Ray compatibility before committing to a training run, since neither has a CPU fallback. The Apache 2.0 license permits commercial research use without restriction.
Frequently asked questions
What training algorithms does Open-AgentRL support?
The examples/ directory contains trainers for GRPO, PPO, REINFORCE++, RLOO, ReMax, and SFT, among others. Each algorithm has its own subdirectory under examples/ with configuration files.
Where are the DemyAgent datasets available?
The README points to two HuggingFace collections: Gen-Verse/open-agentrl-68eda4c05755ca5a8c663656 for the 3K SFT and 30K RL data, and Gen-Verse/DemyAgent-4B for the trained 4B model checkpoint.
Does Open-AgentRL require a GPU cluster to run?
The framework is built on Ray for distributed training and requires flash-attn, which needs a CUDA-capable GPU. The README does not document a single-GPU or CPU-only path, so a multi-GPU setup is the expected baseline.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/gen-verse-open-agentrl)