# OpenManus-RL: a live-streamed RL tuning project for LLM agents

> OpenManus-RL is an extension of OpenManus, run jointly by Ulab-UIUC and MetaGPT, that explores reinforcement learning tuning for LLM agents. The README describes a roadmap and methods; it does not describe a finished, installable agent training pipeline.

**OpenManus/OpenManus-RL** — A live stream development of RL tunning for LLM agents

- Repository: https://github.com/OpenManus/OpenManus-RL
- Stars: 4,160 · Forks: 593
- Language: Python
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/openmanus-openmanus-rl

## What OpenManus-RL is trying to solve, and who it is for

OpenManus-RL starts from a specific gap. The original OpenManus project gives you an agent that plans and calls tools, but it does not give you a way to improve that agent's decisions from experience. The README frames the project as an exploration of "new paradigms for RL-based LLM agent tuning", inspired by RL tuning work on reasoning models such as Deepseek-R1 and QwQ-32B. The intended audience is narrow: researchers and engineers who already have an agent loop and want to post-train the model that drives it, rather than prompt it better.

The README is explicit that this is a collaborative effort between Ulab-UIUC and MetaGPT, and that it is an extended version of the OpenManus initiative. It also states the project will be updated "in a dynamic, live-streaming fashion". That framing matters for adoption: the repository is closer to a public research log than to a released library. The last push was on 2026-05-05, and the most recent release listed is OpenManus 0.0.2 from 2025-05-09. The README does not describe a stable API surface, and it does not claim one.

## The method: rollout strategies, reasoning formats and reward models

The README's Method section is the most concrete part of the project. It says the approach draws on RAGEN's Reasoning-Interaction Chain Optimization (RICO) and then explores alternative algorithmic structures. Four rollout strategies are named: Tree-of-Thoughts, Graph-of-Thoughts, DFSDT (depth-first search decision trees) and Monte Carlo Tree Search. Each is described in one line, for example MCTS "explores reasoning and decision paths probabilistically, balancing exploration and exploitation".

On the output side, the README compares ReAct, which interleaves reasoning and action, against outcome-based reasoning, which optimizes toward explicit outcome predictions. Post-training is split into Supervised Fine-Tuning, described as initializing reasoning capabilities from human-annotated instructions, and generalized reward-based methods. The roadmap lists four stages in order: agent environment support for online RL tuning, trajectory collection from reasoning models such as deepseek-r1 and QwQ-32B, an RL fine-tuning paradigm for customizing agent behavior, and evaluation on WebShop, GAIA, OSWorld and AgentBench.

What the README does not give is the wiring between those stages. There is no description of how a rollout is scored, how the reward model is trained, or how trajectories move from the environment into the training loop. The section on training an agent reward model is listed in the table of contents but the excerpt stops before its content. Treat the Method section as a research agenda, not as an architecture you can implement from the README alone.

## Installing OpenManus-RL and what a first run actually looks like

The README does not contain installation instructions. It links to a Hugging Face dataset at CharlieDreemur/OpenManus-RL and states that "Code and dataset are now available" and that the verl submodule has been integrated. Beyond that, the only installation-shaped artifacts are the repository files themselves.

The dependency list in requirements.txt is the closest thing to a setup contract. Note the pins: transformers==4.51.1, tensordict<=0.6.2, and a commented-out vllm==0.8.4. The pyproject.toml in the repository is actually verl's packaging metadata, not OpenManus-RL's; it declares the project name as "verl" with description "veRL: Volcano Engine Reinforcement Learning for LLM" and pins transformers<4.48 and vllm<=0.6.3, which conflicts with requirements.txt. That mismatch is a real thing to resolve before you install anything.

```bash
pip install -r requirements.txt
```

Running that installs the agent-side dependencies, including flash-attn and liger-kernel, which are GPU packages. If the install fails on those, the repository also ships requirements_docker.txt as an alternative list.

The environment file shows which credentials the agent side expects. Copy .env.example to .env and fill in the keys you have.

```bash
cp .env.example .env
```

```bash
OPENAI_API_BASE=
OPENAI_API_KEY=<your api key>
TOGETHER_API_KEY=<your api key>
GOOGLE_API_KEY=
GOOGLE_CX=
```

GOOGLE_API_KEY and GOOGLE_CX together suggest a Google Custom Search tool; the README does not document which tools read which variable. The verl submodule is a separate checkout, so training-side setup is not covered by the top-level requirements.txt. The README gives no example command for launching a rollout or a training job, and no expected output. Until you find an entry point under openmanus_rl/ or scripts/, the first real use is reading code, not running it.

## Where OpenManus-RL breaks down, and when it is the wrong tool

The clearest limitation is documentation depth. A reader arriving from the README learns the research questions but not the interfaces. There is no documented CLI, no configuration schema for a training run, no statement of which GPU count or memory a run needs, and no rollback procedure if a fine-tuned checkpoint degrades agent behavior. For an RL project, that last omission is notable: post-training can make an agent worse on tasks it previously handled, and the README offers no evaluation gate to catch it.

The dependency situation is the second problem. Two packaging files in the same repository disagree about the transformers ceiling, one at 4.51.1 and one below 4.48, and about vllm, one commented at 0.8.4 and one capped at 0.6.3. Whichever path you take, the other file will mislead you. This is the kind of drift that appears when a vendored submodule's metadata sits at the repository root.

Third, the project is scoped to a research loop. If your goal is to ship an agent that books meetings or fills forms, the RL tuning work here adds nothing you can use, and the OpenManus base project is the relevant code. If your goal is to fine-tune an agent on a benchmark and publish a number, the missing pieces are the ones you would have to build: environment wrappers, reward scoring, and a reproducible evaluation harness. The README names the benchmarks but does not say which ones are wired up.

## How OpenManus-RL differs from verl and from RAGEN

The nearest comparison is verl, the Volcano Engine RL framework whose code is vendored here as a submodule. verl is a general LLM RL training library: it handles distributed rollout, actor and critic updates, and inference backends such as vLLM and SGLang, which appear in its setup.py extras. Its unit of work is a prompt and a reward. OpenManus-RL's unit of work is an agent trajectory, meaning a sequence of reasoning steps and tool calls inside an environment. The two are not competitors; OpenManus-RL uses verl for the training mechanics and supplies the agent-specific layer. If you already run verl, the question is whether the environment and trajectory code here is worth adopting, and the README does not answer that.

The second comparison is RAGEN, which the README credits for the RICO idea. RAGEN's approach centers on optimizing the reasoning-interaction chain. OpenManus-RL's stated difference is breadth: it adds alternative rollout strategies (ToT, GoT, DFSDT, MCTS), a comparison of ReAct against outcome-based reasoning, and reward model training. Whether that breadth is realized in code is not something the README establishes.

A third option is to skip RL entirely and use the base OpenManus agent with a stronger reasoning model. Given the README's own observation that reasoning models like Deepseek-R1 and QwQ-32B already provide strong capabilities, that path is cheaper and has a documented codebase behind it.

## Licence, maintenance and the cost of keeping up

The repository is Apache-2.0, and the vendored verl code carries the same licence, with the copyright header in setup.py attributed to Bytedance Ltd. and/or its affiliates. Apache-2.0 permits commercial use and modification and requires that you retain notices and state changes. The README does not discuss licence implications for models you fine-tune, and that is a separate question from the code licence. Nothing here is legal advice; check the terms attached to any base model and dataset you use.

Maintenance is the harder judgement. The last push was on 2026-05-05, more than four months before today, and the newest release listed is OpenManus 0.0.2 from 2025-05-09. The README promises live-streamed updates, but the release history does not show a cadence you can plan around. The project also invites contributions to a fine-tuning codebase, tuning dataset, environment setup and computing resources, which suggests those pieces are still being assembled.

Upgrade cost has one concrete driver: the pinned dependency set. Because the repository vendors verl, updating OpenManus-RL may mean reconciling two dependency manifests with different transformers and vllm ceilings. Budget for that, and treat the pins in requirements.txt as the operative ones until the repository says otherwise.

## Conclusion

Adopt OpenManus-RL only if you are researching RL post-training for agents and are willing to work from the repository layout rather than a documented install path. The README describes a roadmap, not a stable tool: it lists environments such as GAIA, AgentBench, WebShop and OSWorld, and rollout strategies such as ToT, GoT, DFSDT and MCTS, but gives no commands, no version pinning guidance for the agent side, and no rollback instructions. If you need a working agent today, use OpenManus itself. Before committing, check the openmanus_rl/ and verl directories for a runnable entry point, read requirements.txt for the pinned transformers==4.51.1 and tensordict<=0.6.2, and confirm which benchmark you intend to reproduce.

## FAQ

### How does OpenManus work?

The README frames OpenManus-RL as an extended version of the original OpenManus initiative, which provides an agent that reasons and calls tools. OpenManus-RL adds RL tuning on top of that agent loop, collecting trajectories and post-training the driving model. The README does not describe OpenManus's internal loop in detail.

### Is Manus AI open-source?

The README describes OpenManus-RL as an open-source initiative led by Ulab-UIUC and MetaGPT, licensed Apache-2.0, and an extended version of the OpenManus project. It does not make any claim about Manus AI itself.

### What is OpenManus-RL?

It is a joint Ulab-UIUC and MetaGPT project that explores RL-based tuning for LLM agents, built as an extension of OpenManus. The README states it covers rollout strategies such as ToT, GoT, DFSDT and MCTS, post-training strategies including SFT, and evaluation on benchmarks such as GAIA, AgentBench, WebShop and OSWorld.

### Where do I download the OpenManus-RL dataset?

The README links to a Hugging Face dataset under CharlieDreemur/OpenManus-RL. A news entry dated 2025-03-09 says the Agent SFT dataset was collected and open-sourced there.

### Does OpenManus-RL use verl for training?

Yes. The README states that the verl submodule has been integrated for enhanced RL training capabilities, and a verl/ directory appears among the top-level repository entries. The repository's pyproject.toml is verl's own packaging metadata, which pins transformers<4.48 and vllm<=0.6.3.

## Sources

- [Issues](https://github.com/OpenManus/OpenManus-RL/issues)
- [License: Apache-2.0](https://github.com/OpenManus/OpenManus-RL/blob/main/LICENSE)
- [OpenManus/OpenManus-RL on GitHub](https://github.com/OpenManus/OpenManus-RL)
- [README](https://github.com/OpenManus/OpenManus-RL/blob/main/README.md)
- [Releases](https://github.com/OpenManus/OpenManus-RL/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/openmanus-openmanus-rl
