# AgentsMeetRL: A Curated Index of Open-Source Agentic RL Projects

> AgentsMeetRL is an actively maintained awesome list that catalogs open-source repositories training LLM agents with reinforcement learning, organized across 16 taxonomy categories with documented RL frameworks, algorithms, reward types, and environments. It also ships as a Claude Code skill for on-demand training guidance.

**thinkwee/AgentsMeetRL** — Awesome List for Agentic RL

- Repository: https://github.com/thinkwee/AgentsMeetRL
- Website: https://thinkwee.top/amr/
- Stars: 1,851 · Forks: 74
- Language: HTML
- License: not declared
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/thinkwee-agentsmeetrl

## What AgentsMeetRL Covers and Who Should Use It

Training LLM agents with reinforcement learning is a young and fragmented field. Dozens of repositories have appeared covering multi-turn reasoning, tool use, web browsing, code generation, memory, and multimodal interaction. Each project makes different choices about RL framework, algorithm, reward signal, and evaluation environment. Finding and comparing those choices requires reading many README files individually.

AgentsMeetRL solves that problem by maintaining a structured list of verified, code-released agentic RL repositories. The README makes the entry criteria explicit: a project qualifies as an agent if it involves multi-turn interactions or tool use. Projects whose papers have been published but whose code has not yet been released are tracked separately in an Under Review section and are excluded from the main tables until code appears publicly. The audience is researchers and engineers who need an overview before making a technical choice, not someone looking for a production-ready framework to install.

## Taxonomy: 16 Categories and What Each One Contains

The list organizes repositories into 16 categories:

- Base Framework: general-purpose RL training frameworks for LLM agents, such as veRL, OpenRLHF, and trl
- General/MultiTask: agents trained or evaluated across multiple tasks or environments
- Search and RAG: search-augmented reasoning agents that use retrieval tools
- Web and GUI: agents interacting with browsers, mobile GUIs, or operating systems
- Tool-Use: agents trained to call external tools (APIs, code executors, MCP)
- Code and SWE: software engineering and code generation agents
- Reasoning: reasoning agents with tool-integrated or multi-turn reasoning (math, QA, visual)
- Multi-Agent RL: multi-agent collaboration, negotiation, or credit assignment
- Memory: agents that learn to manage or retrieve memory
- Embodied: agents in embodied or physical simulation environments
- Domain-Specific: RL agents for specialized domains such as medical or OS tuning
- Reward and Training: process and outcome reward models, training methodologies
- Safety: RL for agent safety alignment, adversarial red-teaming, jailbreak defense
- VLM Agent: vision-language model agents trained with RL
- Self-Evolution: agents that self-evolve via RL feedback loops
- Environment: benchmarks, gyms, and sandbox environments

The README notes that the Self-Evolution category definition is still evolving in the community, which is a fair signal that classification there may shift as the field settles.

## What Each Table Entry Documents

For each project, AgentsMeetRL records the RL framework it depends on, the RL algorithm used, the reward type, and the environment. Reward types are standardized into four classes: External Verifier (such as a compiler or math solver), Rule-Based (such as a LaTeX parser with exact match scoring), Model-Based (such as a trained verifier LLM), and Custom. This classification lets a reader quickly filter for projects using a particular reward design without reading each individual paper.

The README adds a expandable "Click to view technical details" section under each table, providing deeper information about framework-level choices. This is the fastest path to answers like "which projects use veRL" or "which reward models use an external verifier" without leaving the list.

## How the Repository Is Organized and How to Browse It

The primary content lives in README.md and index.html. The README is the reference document; index.html is a rendered version hosted at thinkwee.top/amr/. A CITATION.cff file in the root supports citing the list in academic work using the button on the GitHub sidebar.

The skills/ directory holds a Claude Code skill named agents-meet-rl. Installing it involves placing the skills/ directory where Claude Code or Codex can find it. Once active, the skill turns the corpus into an on-demand assistant for agentic RL questions: reward not moving, KL or entropy blow-ups, GRPO or PPO or DAPO knobs, retokenization drift, tool-call parse failures, long-horizon credit assignment, and LLM-judge inconsistency are listed in the README as example query categories.

## Update Pace and What Gets Left Out

The README documents monthly updates. The August 2026 update added 23 repositories across 9 categories, the July 2026 update added 13 across 8, the June 2026 update added 43 across 11, and the April 2026 update added 67 across nearly every category. Each update entry explicitly names what was excluded and why: the August 2026 note, for example, flags that no qualifying Safety, Embodied, or Multi-Agent RL repositories appeared that window because the available options were uniformly SFT-only, inference-only, or had withheld code.

This transparency about omissions is a practical feature. A researcher who noticed that a specific project was not listed can check whether it falls into the Under Review section or was evaluated and found not to meet the agent criteria. The README's description of the entry criteria (at least one of multi-turn interaction or tool use) gives a clear standard to apply when assessing a candidate project.

## Limitations: What AgentsMeetRL Does Not Provide

AgentsMeetRL is a survey and index, not a framework. It does not provide code that runs, benchmark scores, or installation instructions for any of the listed projects. A reader looking to train an agent needs to follow the link to a listed repository and work from that repository's own documentation.

The README acknowledges that the list is based on code analysis by LLM coding agents, with manual review, and states plainly that "there may still be omissions." Classification of any given project into one of the 16 categories involves judgment calls, and the README invites corrections through issues and pull requests. Researchers in narrow sub-areas may find that their specific sub-field is underrepresented, particularly if relevant work appeared in a gap between update windows.

The Under Review section exists specifically for papers that describe qualifying agent behavior but have not yet released code publicly. Named examples from the August 2026 update include Qwen-UI-Agent, Qwen-CUA, UI-Mate, SearchMaster, RoMeRL, Agon, SINKFLEX-RL, GRASP, MAVEN, EviBack, and ChemWorld. These move to the main tables only once code appears, which means the list is intentionally incomplete at any given moment for the most recent research.

## License and Maintenance

The repository does not specify a license in the root listing. The CITATION.cff file supports formal academic citation via the GitHub sidebar button. The last push to the repository was on 2026-09-15, and the README's most recent documented update is dated 2026-08-26. The repository is not archived.

## Conclusion

AgentsMeetRL is the right starting point for researchers and engineers who need to survey the agentic RL ecosystem before choosing a framework or reward design. It covers verified, code-released projects only; papers whose code is still private are held in an Under Review section. Teams looking for a running framework rather than a survey should go to one of the Base Framework entries (veRL, OpenRLHF, trl are named in the README) rather than treating this list as a product itself. The last documented update was 2026-08-26, and the repository received its most recent push on 2026-09-15.

## FAQ

### What criteria does AgentsMeetRL use to classify a project as an agent?

The README states that a project must have at least one of the following to qualify: multi-turn interactions or tool use. Projects meeting this definition include Tool-Integrated Reasoning (TIR) projects. Repositories whose papers describe agent behavior but whose code has not been publicly released are placed in an Under Review section rather than the main tables.

### How do I use the AgentsMeetRL Claude Code skill?

The skill lives in the skills/ directory of the repository. The README describes it as a Claude Code skill named agents-meet-rl that turns the corpus into an on-demand assistant for agentic RL training, evaluation, and experiment design questions. The README does not document a specific installation command beyond pointing to the skills/ directory.

### How often is AgentsMeetRL updated?

The README documents updates in 2026 for March, April, May, June, July, and August, suggesting roughly monthly additions. Each update entry names newly added repositories and explicitly lists excluded projects with reasons.

## Sources

- [Issues](https://github.com/thinkwee/AgentsMeetRL/issues)
- [Project website](https://thinkwee.top/amr/)
- [README](https://github.com/thinkwee/AgentsMeetRL/blob/main/README.md)
- [thinkwee/AgentsMeetRL on GitHub](https://github.com/thinkwee/AgentsMeetRL)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/thinkwee-agentsmeetrl
