# AppWorld puts nine apps and 457 APIs behind a simulated population

> A controllable execution environment for benchmarking function calling and interactive coding agents, with a companion benchmark of autonomous tasks. It arrived as an ACL 2024 best resource paper and now also exposes the environment over MCP for evaluating terminal-based agents.

**StonyBrookNLP/appworld** — 🌍 AppWorld: A Controllable World of Apps and People for Benchmarking Function Calling and Interactive Coding Agent, ACL'24 Best Resource Paper.

- Repository: https://github.com/StonyBrookNLP/appworld
- Website: https://appworld.dev/
- Stars: 526 · Forks: 86
- Language: Python
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/stonybrooknlp-appworld

## Nine apps, 457 APIs and about a hundred synthetic residents

The engine is an execution environment for day-to-day applications, described as high fidelity and described as controllable, with nine apps behind it and 457 APIs to call. Those apps are populated with the digital activity of roughly one hundred people living in a simulated world, which is what turns it from a set of mock endpoints into something with state that has to be reasoned about.

The companion half is a benchmark of autonomous agent tasks, characterised as natural, diverse and challenging, and explicitly requiring rich and interactive coding rather than a single function call. That pairing is the contribution: a world to act in, and a set of goals to try to reach inside it.

The paper appeared at ACL 2024 and received a best resource paper award. The site around the code carries a task explorer for browsing the dataset, an API explorer for the interface surface, a leaderboard, videos, a blog write-up and the paper itself on arXiv.

The package is published on PyPI, requires Python 3.11 or newer, and the manifest marks the development status as alpha. It also declares Python 3.11 through 3.14 as supported and classifies itself for both a research and a developer audience.

## Three commands turn a wheel into a working world

Installation is separated from data download on purpose, which is unusual and worth knowing before you start:

```bash
# Install and download:
pip install appworld
appworld install
appworld download data
```

The middle command sets up the environment itself. The third fetches the content that makes it usable, and skipping it leaves you with a working import and nothing to act on.

Everything is then reachable through one command line interface, which advertises a help listing of its own commands. Two of those are described as exploratory: one to browse the dataset and one for an interactive playground, where you can poke at the environment without writing an agent at all.

That playground is the fastest way to form an intuition for the world before deciding how to attack the benchmark, and it costs nothing beyond the data download you already needed.

## Each task is a context manager wrapping one world

The programmatic entry point builds one world per task, as a context manager, so the lifecycle of a run is expressed in the shape of the code rather than in setup and teardown calls:

```python
from appworld import AppWorld, load_task_ids

for task_id in load_task_ids("test_challenge"): # Or train, dev, test_normal
    with AppWorld(task_id=task_id, experiment_name="sample") as world:
        world.task.instruction # To see task instruction.
        world.execute("""
        # ...
        response = apis.spotify.login(...)
        print(response)
        """)
        # => {"access_token": ...}
```

Three things in those lines carry the design. The instruction is available as data before you write any code, so an agent reads it rather than a human transcribing it. Code runs through an execute method rather than through direct API calls, which is what makes it interactive coding rather than function calling. And the reply is a normal value you can print and inspect.

The execution space is persistent across calls: variables from earlier code and output already printed are both available later, so a sequence of code blocks behaves like one continuous session against the same world rather than a series of disconnected requests.

## Four splits, and the difficulty is concentrated in one

Task identifiers are loaded by split name, and four are named in the documentation: train, dev, test_normal and test_challenge.

The naming implies the intended usage pattern rather than stating it. A normal split and a challenge split exist as separate difficulty tiers, with the challenge set being the one the worked example uses, and the presence of both a normal and a challenge test set means a reported number should always say which one it came from.

Runs are named too. Every world is constructed with an experiment name alongside the task id, and evaluation takes that experiment name and a dataset name as its two positional arguments:

```bash
# Evaluate:
appworld evaluate sample test_challenge
#    experiment name ^     ^ dataset name
```

So the two dimensions that define a result, which tasks and which agent, are both explicit arguments rather than ambient state. That is what makes it possible to evaluate a saved run later without re-running the agent, and it is the difference between a benchmark you can iterate on and one you can only re-run from scratch.

## An MCP server is now the newest route into the world

The most recent addition is an MCP server and client, which shifts the integration away from writing code against a Python package and towards connecting an existing agent.

Four connection routes are documented, and the differences between them are about who owns the agent loop. One is for a graphical MCP client. One is for any agent framework that already speaks MCP. One uses a client built into AppWorld itself, which is the shortest path if you have no agent tooling yet. The fourth is a standalone client shipped by the project, for the case where you want the protocol without adopting either side's harness.

There is also a dedicated guide for evaluating terminal agents, naming Codex and Gemini as examples, through the MCP path. That is the significant part for current practice: an agent that already works in a terminal does not need to be rewritten in Python to be scored, it needs to be pointed at the world.

The documentation also carries sections on agent development restrictions and on code execution safety, both of which are worth reading before you let an autonomous loop write and run code against a live environment.

## One run command covers five agents and a hundred models

Reference agent implementations are packaged separately from the engine, so the core install stays small:

```bash
pip install -e 'experiments[simplified]'  # from repo or 'pip install appworld-agents' once published.
appworld run auto --agent-name {AGENT_NAME} --model-name {MODEL_NAME} --dataset-name test_challenge
# Replace any/all {...} with "options" to see choices.
```

The claim is five or more agents against a hundred or more models. Substitution with an options keyword is how you discover the valid values rather than reading them off a list, which matters when the agent roster changes between releases.

Agent experiments have their own section of the documentation covering installation, experiment options, running and evaluating, downloading experiment outputs, adding your own base model and customizing agents. The last two are the extension points: adding a base model is how you test a new model without writing an agent, and customizing agents is how you change the scaffold rather than the world.

A minimal ReAct agent is also provided as a notebook, which is the reference implementation to read when you want to know what the harness expects of a solution rather than what the harness does.

## Guides cover the parts the maintainers expect you to extend

The repository is organised around the assumption that you will extend the environment, not only run it against, and five guides are provided for that.

One covers generating the base database and the tasks from scratch. One covers developing new applications inside the world, which is how you add capabilities without invalidating existing tasks. One covers writing new task generators, which is how you extend the benchmark rather than just the environment.

One covers evaluating terminal agents through the MCP path. The fifth covers parallelising worlds, and it names its motivation directly as reinforcement learning rollouts or simply faster runs, which tells you the environment was built to be run many times over rather than once.

Around that sits ordinary project structure: a source directory, an experiments directory, a generate directory, guides, notebooks including the minimal agent, tests with their own configuration, scripts, a dockerfile, and a private ignore file that keeps generated data out of version control. A release process document and a development document sit at the top level.

One versioning note. The manifest declares a development version while the newest published release is a post release from February 2026, and the last push to main is dated 2026-09-04, so the gap between the published package and the working tree is larger than the tag history alone suggests. There is also an explicit release disclaimer in the documentation, which is worth locating before you rely on any result.

## Conclusion

AppWorld suits research into agent tool use, because the environment is small enough to reason about while still being messy enough to be realistic, and the fixed API count and simulated population make two runs comparable in a way a live service never is. It suits you less if you want a realistic corpus of real human behaviour, since the population is synthetic by design, and a leaderboard number here is not a claim about your product. Before running anything, download the data after installing, since the engine is useless without it, and read the sections on agent development restrictions and code execution safety before pointing an autonomous agent at the sandbox.

## FAQ

### What is inside the AppWorld environment?

Nine day-to-day apps operable through 457 APIs, populated with the digital activity of roughly one hundred people in a simulated world, alongside a benchmark of autonomous agent tasks requiring interactive coding.

### Why does installing AppWorld need a separate data download?

Installation and data download are separate commands. After installing the package, an install command sets up the environment and a download command fetches the content, and without that second step the engine has nothing to act on.

### How do I write code against an AppWorld task?

Load task ids for a split, open a world per task as a context manager, read the instruction from the task object, and run your code through the world's execute method. State persists across calls, so earlier variables and printed output remain available.

### What dataset splits does AppWorld provide?

Four are named: train, dev, test_normal and test_challenge. Since normal and challenge test sets are separate difficulty tiers, any reported score should state which split it came from.

### Can I evaluate a terminal agent like Codex or Gemini with AppWorld?

Yes. An MCP server and client were added as the newest integration, with four documented connection routes, and there is a dedicated guide for evaluating terminal agents through that path. A minimal ReAct agent is also provided as a reference notebook.

### How do I add a new model to AppWorld without writing an agent?

The agent experiments section covers adding your own base model as one of its extension points, alongside customising agents. The run command takes an agent name, a model name and a dataset name, and substituting options reveals the valid choices.

## Sources

- [License: Apache-2.0](https://github.com/StonyBrookNLP/appworld/blob/main/LICENSE)
- [Project website](https://appworld.dev/)
- [README](https://github.com/StonyBrookNLP/appworld/blob/main/README.md)
- [Releases](https://github.com/StonyBrookNLP/appworld/releases)
- [StonyBrookNLP/appworld on GitHub](https://github.com/StonyBrookNLP/appworld)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/stonybrooknlp-appworld
