# Meta Agents Research Environments: running the Gaia2 benchmark on agents that have to keep up with a changing world

> ARE is Meta's Python platform for evaluating LLM agents in scenarios that evolve while the agent works. It ships the Gaia2 benchmark, a CLI, and a browser GUI, and it installs in one uvx command.

**facebookresearch/meta-agents-research-environments** — Meta Agents Research Environments is a comprehensive platform designed to evaluate AI agents in dynamic, realistic scenarios. Unlike static benchmarks, this platform introduces evolving environments where agents must adapt their strategies as new information becomes available, mirroring real-world challenges.

- Repository: https://github.com/facebookresearch/meta-agents-research-environments
- Website: https://facebookresearch.github.io/meta-agents-research-environments/
- Stars: 559 · Forks: 78
- Language: Python
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/facebookresearch-meta-agents-research-environments

## What ARE is for, and who ends up using it

Static benchmarks hand an agent a fixed prompt and score the answer. That works for question answering and breaks down for anything that takes minutes and ten or more steps, because in those tasks the world does not hold still. The README states the platform's purpose plainly: it introduces evolving environments where agents must adapt their strategies as new information becomes available. New mail arrives, a file changes, a calendar entry moves, and the agent has to notice and replan.

The intended audience is agent researchers and evaluation engineers, not application developers. The package is published as meta-agents-research-environments on PyPI and the project describes itself as a research platform. It runs the Gaia2 benchmark, described in the README as a follow-up to Gaia with 800 scenarios across 10 universes, and the repository also carries a multilingual extension, facebook/omnilingual-gaia2, on Hugging Face. If your job is to compare two agent architectures on tasks that require multi-step reasoning, the scenario model here is closer to the problem than a single-turn eval set is.

## Agents, apps, events and scenarios: the moving parts

The documentation names five core concepts: agents, environments, apps, events, and scenarios. The repository layout backs that up. The are/ directory holds simulation code, and the dependency list in requirements.txt points at the machinery: litellm sits in simulation/agents/llm/litellm_engine.py as the abstraction over model providers, mcp appears in simulation/apps/mcp/, and fsspec is used in simulation/apps/sandbox_file_system.py.

So the shape is this. An app is a simulated surface with a sandboxed file system behind it. An agent is a loop that calls a model through LiteLLM and invokes tools whose schemas are generated from docstrings, which is why docstring-parser is a core dependency. Events are what make the environment dynamic: they arrive during a run and change state. A scenario ties the apps, the initial state, and the event schedule together, and the benchmark scores the agent against it.

The GUI renders a scenario as a DAG, and the README shows a screenshot of that visualization in the Scenarios Mode. That is a useful design tell. A scenario is not a linear script; it is a dependency graph of steps, and the evaluation cares about how the agent traverses it when events perturb the order.

## Installing ARE and running your first Gaia2 scenario

The README's prerequisites section names uv, the Python package installer and resolver from Astral. Install that first, then you have two paths.

The fastest path needs no virtualenv at all. uvx runs the package straight from PyPI, so this command pulls the validation split of the hosted Gaia2 dataset and runs a single scenario:

```bash
uvx --from meta-agents-research-environments are-benchmark gaia2-run --hf meta-agents-research-environments/gaia2 --hf_split validation -l 1
```

The -l flag is the limit, so you should see one scenario execute rather than the whole split. If you would rather run something local and self-contained, the README gives a tutorial scenario for that:

```bash
uvx --from meta-agents-research-environments are-run -s scenario_tutorial -a default
```

For repeated work, install the package instead. The README distinguishes two dependency sets: the bare package for CLI benchmarking, and the gui extra, which it recommends for interactive exploration because it adds the web interface.

```bash
pip install "meta-agents-research-environments[gui]"
```

With the GUI extra installed, are-gui serves a browser interface that the README says typically runs at http://localhost:8080. The Dockerfile exposes that same port and sets ARE_SIMULATION_SERVER_PORT to 8080, so the port is consistent across both routes. The interface has a Playground Mode for chat-like interaction and a Scenarios Mode for structured execution with the DAG view.

Model access goes through LiteLLM, and the README shows an environment variable for the Llama API route plus a local endpoint option:

```bash
export LLAMA_API_KEY="your-api-key"
are-benchmark run --hf meta-agents-research-environments/gaia2 --hf_split validation \
  --model Llama-3.1-70B-Instruct --provider llama-api --agent default
```

The local variant uses --provider local with --endpoint pointing at your own server. The README points to a separate LLM Configuration Guide for the rest of the provider matrix.

## Python 3.10, pinned dependencies and the GUI build

pyproject.toml sets requires-python to >=3.10, and the README badge still advertises Python 3.8+. Those disagree, and the packaging metadata is the one that matters at install time. If you are on 3.9, pip will refuse the package.

The core dependency list is pinned exactly, not ranged: click 8.1.8, datasets 4.0.0, litellm 1.71.1, mcp[cli] 1.11.0, numpy 2.2.6, pillow 11.1.0, fsspec 2024.12.0, Jinja2 3.1.6. Exact pins make a research environment reproducible, which is the right call for a benchmark, but they also mean your existing environment will likely conflict. Use the uvx path or a dedicated virtualenv rather than installing into a shared one.

The GUI adds a second toolchain. The Dockerfile builds the frontend in a node:23 stage, runs npm install and npm run build inside are/simulation/gui/client, and copies the resulting build directory into the Python image. The wheel build explicitly excludes the client's node_modules, src, and Vite output, so a pip install of the gui extra does not ship prebuilt JavaScript through the wheel the way the container does. Expect the GUI path to be the heavier of the two to get working.

## Where this platform will not help you

ARE is a research platform, and that shows in what it does not promise. There are no retrieved releases for it, so version 1.2.0 in pyproject.toml is the packaging version rather than a tagged release you can pin against. If your team needs a frozen, versioned evaluation artifact for regression testing, this is the wrong shape of dependency.

The scenario authoring model is also a real cost. Scenarios live in Python under are/simulation/scenarios/, and the wheel build excludes are/simulation/datasets/ entirely, which tells you the datasets are fetched rather than bundled. Writing a new scenario means writing code against the app and event abstractions, not editing a YAML file. For a one-off evaluation of a prompt change, that is far more work than it is worth.

There is a live-model dependency baked into the evaluation loop. Agents reach models through LiteLLM, so a benchmark run needs credentials or a local endpoint, and the results are entangled with whichever model and provider you pointed at. The README's leaderboard is explicitly described as self-published results, which is an honest label and also a warning: cross-run comparison depends on matching model, provider, and split, and the platform does not enforce that for you.

One more boundary. The README documents the CLI commands, the GUI, and the model configuration, but it does not document rollback or resumption of a partially completed benchmark run. If you need to stop and continue a long evaluation, check the benchmark CLI's own --help output before assuming it exists.

## How ARE differs from a static agent benchmark

The obvious comparison is the original Gaia benchmark, since the README frames Gaia2 as its follow-up. Gaia presents a fixed set of questions with attached files and expects a final answer. The environment is inert between the question and the response. ARE's stated difference is that the state evolves and new information is continuously integrated, so an agent that plans once and executes will lose to one that re-reads its environment.

That difference has a practical consequence for how you write the agent. In a static benchmark, the tool surface is mostly retrieval. Here, apps are simulated systems with their own state, and events push changes into them mid-run. The MCP integration in simulation/apps/mcp/ matters for the same reason: it lets a tool live outside the process while still participating in the scenario. If your agent architecture assumes a read-only world, ARE will expose that assumption rather than measure it.

The trade-off is cost and determinism. Dynamic scenarios with live model calls take minutes per run and are harder to reproduce exactly than a static eval with cached responses. That is the price of testing adaptation, and it is worth paying only if adaptation is the thing you actually want to measure.

## Licence, maintenance and what an upgrade costs you

The licence is MIT, declared both in the README badge and in pyproject.toml as {text = "MIT License"}. For a benchmark harness that is about as permissive as it gets: you can vendor it, modify it, and ship derived evaluation code. The LICENSE file is at the repository root. Nothing here restricts commercial use of the code, though the Gaia2 dataset is hosted separately on Hugging Face and carries its own terms, which you should read before redistributing it. That is a factual pointer, not legal advice.

The repository is not archived, and the last push was on 2026-08-26, which is recent enough to call the project maintained. There are no retrieved releases, so upgrades are not something you schedule against a changelog. You track main.

Upgrade cost is dominated by the exact pins. Bumping litellm or datasets means resolving against a dependency set that the project has fixed deliberately, and the GUI adds a Node toolchain on top. The practical approach is to treat ARE as its own environment, pinned to a commit, and to re-run the validation split at a fixed limit after any bump so you can see whether scores moved because of the library or because of the model.

## Conclusion

Adopt ARE if you need to measure how an agent behaves when the environment changes under it, and if you are willing to work with a research platform whose scenarios are authored as Python rather than as data files. Skip it if you need a stable evaluation harness with a frozen task set, or if you cannot supply model credentials through LiteLLM. Before committing, run the validation split at limit 1 and open the GUI to confirm the scenario DAG renders the way your task expects, then check whether the 800 Gaia2 scenarios cover your domain at all.

## FAQ

### What is a meta agent?

The project does not define the term in the README. It describes ARE as a platform for evaluating AI agents in dynamic scenarios, and the package name is meta-agents-research-environments, where meta refers to Meta's research organisation rather than to a meta-agent architecture.

### What is the difference between an agent and an environment?

The documentation lists agents, environments, apps, events and scenarios as the core concepts. An agent is the model-driven loop that calls tools, while the environment is the simulated state, made of apps and the events that change them during a run.

### What are the different types of agent environments in AI?

The README does not enumerate a taxonomy of environment types. It draws one distinction that matters here: unlike static benchmarks, ARE environments evolve while the agent works, with new information integrated during the run.

## Sources

- [facebookresearch/meta-agents-research-environments on GitHub](https://github.com/facebookresearch/meta-agents-research-environments)
- [Issues](https://github.com/facebookresearch/meta-agents-research-environments/issues)
- [License: MIT](https://github.com/facebookresearch/meta-agents-research-environments/blob/main/LICENSE)
- [Project website](https://facebookresearch.github.io/meta-agents-research-environments/)
- [README](https://github.com/facebookresearch/meta-agents-research-environments/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/facebookresearch-meta-agents-research-environments
