# AgentGym's fourteen environments are nine servers and eleven datasets

> Code for a 2024 arXiv paper, with a reinforcement learning framework that lives in a second repository, a trajectory set of 14,485 episodes and an evaluation set of 1,160 tasks. Five of the fourteen environments come from one upstream suite and three contribute no trajectories at all.

**WooooDyy/AgentGym** — Code and implementations for the ACL 2025 paper "AgentGym: Evolving Large Language Model-based Agents across Diverse Environments" by Zhiheng Xi et al.

- Repository: https://github.com/WooooDyy/AgentGym
- Website: https://arxiv.org/abs/2406.04151
- Stars: 850 · Forks: 117
- Language: Python
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/woooodyy-agentgym

## Fourteen environments are served by nine server directories

The suite table lists 14 environments and, in its own rightmost column, the server each one runs on. Counting the distinct names there gives nine: webshop, webarena, lmrlgym, alfworld, sciworld, babyai, textcraft, tool and sqlgym. Two environments, a maze task and a word game, share one server. Five more, weather, movie, academia, sheet and a to-do list, all share another, and all five trace to the same upstream project, a task suite from a Hong Kong university group. So the headline number counts tasks rather than integrations, which is a reasonable choice for a benchmark and a misleading one for judging how much code this repository contains. The environments are grouped by theme as web navigation, text games, house-holding tasks, digital games, embodied tasks, tool use and programming, and they all present the same ReAct format so a single agent interface can drive any of them.

## Three environments have evaluation tasks and no trajectories

The table has two count columns and they do not agree about coverage. Web browsing has 200 evaluation tasks and 0 trajectories. A spreadsheet task has 20 and 0, and an academic search task has 20 and 0. So eleven of the fourteen environments contribute training episodes and three contribute nothing but a test set. That is a meaningful gap rather than a rounding detail, because the point of a trajectory set is to show an agent what a successful interaction looks like in that environment. An agent trained on AgentTraj has no demonstration of the spreadsheet or academic search formats at all, and the first time it meets them is the evaluation. Anyone using the pair of artefacts together should treat those three as out-of-distribution rather than as training-and-test splits of the same task.

## Fourteen thousand episodes, and two environments are half of them

Adding the trajectory column gives 14,485 episodes in total, and the distribution is lopsided. The web shopping environment contributes 3,930 and the text-to-SQL one 3,000, which together are 6,930, just under half the set. Next come an embodied household environment at 2,420, a science simulation at 2,120, a word game at 955 and a gridworld task at 810. The tail is thin: a crafting environment at 374, a weather lookup at 311, a maze at 215, a movie lookup at 215 and a to-do list at 135. The evaluation side is far smaller and more even at 1,160 tasks, with 200 each for shopping, web, science, text-to-SQL and the household tasks. Training a generalist on this data means the agent's prior is dominated by two of fourteen task types.

## The reinforcement learning work is a second repository

The repository summary calls the underlying work an ACL 2025 paper, and the page's linked preprint is the June 2024 arXiv entry. Those are not the same thing to follow. Three of the news items, all dated 10 September 2025, announce that agents can now be developed for custom environments and trained with reinforcement learning, and that a second preprint on long-horizon decision making through multi-turn reinforcement learning is out, and that a framework is released. That framework lives in its own repository, and the custom environment tutorial the first item points at is a document inside this one. So this branch is the platform, the datasets and the evaluation code for the original paper, while the training loop people usually want is elsewhere. Anyone arriving for the reinforcement learning headline needs two repositories, and the page does not say which parts move between them.

## Ten environment directories and a submodule file

The root is a list of wrappers. Alongside the shared agentenv package there are ten environment directories, one per integration family: alfworld, babyai, lmrlgym, sciworld, searchqa, sqlgym, textcraft, tool, webarena and webshop. A file declaring submodules sits at the root as well, which is how the wrappers reach the ten upstream projects named in the table's original repository column, from a Princeton web shopping codebase to a Babylon open-world environment to a text-to-SQL benchmark inside a larger Alibaba research tree. Practical consequence: a plain clone does not fetch them, and an environment directory that looks like the project is usually a thin server in front of someone else's. One more directory is worth naming because it has no row in the visible table: a server that answers questions over a search corpus, which makes the on-disk count eleven.

## The visible page has no install command

There is not a single runnable command in the documentation this repository leads with. What a reader gets instead is a table of environments with links to each server directory, a set of artefacts on the model hub, three preprint links, a project page, and a note that contributions for more environments are welcome. A development tutorial is referenced for the custom environment workflow, and it is filed under a numbered tutorials directory, so the guidance exists but sits a level down from the entry page. There is also an interactive frontend directory at the root, added alongside a news item announcing trajectory replay, step-by-step inspection of agent decisions and behaviour analysis, with no install or launch instruction for it anywhere on the page. For a research artefact that has been in place for two years, the amount of runnable guidance a first-time reader gets is close to zero.

## Comparable to the state of the art, with the numbers absent

The introduction's result claim is that evolved agents can achieve results comparable to state-of-the-art models, and no table on the page supports it. The method is named, a trained 7B model is published on the model hub under the suite's organisation, and the original paper is linked, so the evidence is reachable, just not here. The timeline is worth reading next to it. Everything in the news list is dated between June 2024 and September 2025, the newest items all pointing at the separate reinforcement learning repository, while the last push to this branch is dated 2026-05-30. So there are roughly eight months of commits after the last announcement, with nothing on the page saying what they are. Two small slips also sit in the sentences a reader is most likely to quote: the environment table links a science simulation as SciWrold, and the introduction describes the suite as broad, real-time, uniformat and concurrent.

## Conclusion

Use the suite to compare agents on a common ReAct interface, and read the table's two count columns separately, because the training data and the evaluation data do not cover the same environments. Three environments have evaluation tasks and no trajectories, so an agent trained on this set has never seen them act, and two environments supply almost half the episodes. If your interest is reinforcement learning, this repository is the wrong starting point, because the framework the news items point at is a separate project and the code here predates it. If your interest is the released 7B model, note that it was published in June 2024 while the last commit to this branch is dated 2026-05-30, so the artefacts on the model hub and the code on the branch are not the same vintage.

## FAQ

### What is AgentGym and what does it contain?

A framework of 14 interactive environments presented in a unified ReAct format, covering web navigation, text games, household tasks, digital games, embodied tasks, tool use and programming, together with a trajectory set and an evaluation suite. The repository also ships an environment visualisation frontend and a set of documentation including a tutorial on building custom environments.

### How large is the AgentGym trajectory dataset?

The trajectory column across the 14 environments adds up to 14,485 episodes, of which 11 environments contribute any at all. Web browsing contributes 3,930 and text-to-SQL contributes 3,000, so those two make up just under half of the set. The evaluation suite adds up to 1,160 tasks.

### Which AgentGym environments have no training trajectories?

Three: web browsing, spreadsheet tasks and academic search have evaluation tasks, 200, 20 and 20 respectively, and zero trajectories. Every other environment in the table contributes episodes, so those three are evaluation-only and were never demonstrated in the released training data.

### Where is the AgentGym reinforcement learning framework?

In a separate repository. The news items dated 10 September 2025 announce a multi-turn reinforcement learning preprint and a released framework for training agents directly in interactive environments, and both point outside this repository. This branch holds the platform, datasets, benchmark and training implementations for the original paper.

### Which model does AgentGym release?

A 7B model published on the model hub under the suite's organisation, alongside two dataset artefacts, the trajectory set and the evaluation set. The announcement of the model is dated 6 June 2024, while the last push to the repository's default branch is dated 2026-05-30.

## Sources

- [Issues](https://github.com/WooooDyy/AgentGym/issues)
- [License: MIT](https://github.com/WooooDyy/AgentGym/blob/main/LICENSE)
- [Project website](https://arxiv.org/abs/2406.04151)
- [README](https://github.com/WooooDyy/AgentGym/blob/main/README.md)
- [WooooDyy/AgentGym on GitHub](https://github.com/WooooDyy/AgentGym)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/woooodyy-agentgym
