# AReaL: asynchronous RL infrastructure for training agents, not a word game

> AReaL is an Apache-2.0 Python framework from Tsinghua IIIS and Ant Group for asynchronous reinforcement learning on agent applications. Its v2.0 microservice split is the real story, and it is also where the adoption cost lives.

**areal-project/AReaL** — The RL Bridge for LLM-based Agent Applications. Made Simple & Flexible.

- Repository: https://github.com/areal-project/AReaL
- Website: https://areal-ai.io
- Stars: 5,802 · Forks: 613
- Language: Python
- License: Apache-2.0
- Published: 2026-09-09 · Updated: 2026-09-09 · Language: en
- Canonical page: https://hysenlabs.com/projects/areal-project-areal

## The problem AReaL addresses: connecting agent runtimes to RL training

Most agent stacks are built as black boxes. A developer wires a prompt, some tools, and a loop, then wants the policy behind it to improve from experience. The usual answer is to rewrite the agent inside a training harness so the trainer can see every token, reward and action. AReaL takes the opposite position. The README describes replacing the base_url and api_key of an existing agent with AReaL's RL service, with no code changes, and cites an OpenClaw example in examples/openclaw/ as a complete walkthrough. The intended audience is stated plainly: researchers and engineers training large-scale reasoning and agentic models, originally the Tsinghua IIIS and Ant Group teams behind the project. That is a narrower group than the phrase AI agents suggests. If your agent already runs against an OpenAI-compatible endpoint and you have a reward signal, AReaL is aimed at you. If you are building a first prototype with no evaluation loop, the framework's assumptions will not match your stage.

## Why AReaL 2.0 split training, inference, agent and weight-update into services

The July 2026 release notes call AReaL 2.0 a major architectural milestone and describe a refactor into a microservice architecture with independent training, inference, agent and weight-update services. Those map onto directories in the repository: areal/v2/training_service/, areal/v2/inference_service/, areal/v2/agent_service/ and areal/v2/weight_update/. The design follows from the asynchronous paradigm the README leads with. In a synchronous loop, generation and gradient updates alternate, so GPUs idle during rollout. Decoupling the services lets inference keep producing trajectories while training consumes them, and the weight-update service is the component that pushes fresh weights back to the inference side. That separation is also the source of the framework's complexity: four services mean four failure domains, and a stalled agent service can starve the trainer without any single process reporting an error. The README does not document rollback or recovery behaviour for that case, so plan to instrument the boundaries yourself. The earlier AReaL-lite line, announced July 2025, keeps an algorithm-first API with what the notes describe as 80% fewer lines of code and 90% of the performance and core functionality, which is the saner entry point if you do not yet need the microservice topology.

## Installing AReaL with uv and running a first math example

The package targets Linux only. pyproject.toml declares requires-python >=3.11,<3.13, an NVIDIA CUDA 12 environment, and a torch dependency pinned to >=2.9.1,<2.11 except on x86_64 macOS, where it falls back to an older torch. The build backend is uv_build, and the repository ships uv.lock and a separate pyproject.vllm.toml plus uv.vllm.lock for the vLLM variant. Create an environment and install the package from the repository root:

```bash
uv venv --python 3.11
uv sync
```

For a container instead of a host install, the Dockerfile builds an sglang variant by default and a vLLM variant behind a build argument. The file's own usage comment gives both commands:

```bash
docker build -t areal-runtime:dev-sglang .
docker build --build-arg VARIANT=vllm -t areal-runtime:dev-vllm .
```

The Dockerfile starts from lmsysorg/sglang:v0.5.10.post1-runtime, installs cmake, ccache, libibverbs-dev and librdmacm-dev, and sets TORCH_CUDA_ARCH_LIST to "8.0 8.9 9.0 9.0a". That last value is worth reading before you build: the compiled C++ extensions cover those compute capabilities, so older GPUs will not be served by the prebuilt path. For a first real run, the README points at the math examples and names two configs added in June 2026, examples/math/gsm8k_kpop.yaml and examples/math/gsm8k_icepop.yaml. Both are YAML training configs for GSM8K. KPop is described as bidirectional binary KL divergence token masking, selected with rejection_sampling.metric=binary_kl, while IcePop uses importance-ratio-based token masking. Start with the quickstart guide linked from the README rather than editing a config blind.

## Where AReaL is the wrong tool

Pre-alpha is not a formality here. The classifier in pyproject.toml reads Development Status :: 2 - Pre-Alpha even though the version string is 2.1.0, which tells you the API surface is still moving between minor releases. The v2.0 release itself is the evidence: a microservice refactor of this size breaks imports, config keys and launch scripts, and the release notes present it as an architectural milestone rather than a compatibility-preserving change. There is no macOS or Windows support in the classifiers, and the x86_64 macOS torch pin exists only to keep the dependency resolver from failing, not because the training stack runs there. AReaL is also a poor fit when your reward is expensive or subjective. Asynchronous training assumes you can generate many trajectories cheaply; if every reward requires a human review or a slow external service, the throughput advantage of the async loop disappears and you are carrying four services for nothing. Finally, if you only need supervised fine-tuning or preference optimization on static data, this is the wrong layer. AReaL is built around an online loop with a live inference endpoint, and the README's framing of online RL for black-box agent applications makes that assumption explicit.

## AReaL compared with single-process RL libraries and agent frameworks

The closest comparison is a single-process RL library that runs generation and training in one Python job. That approach is easier to debug: one stack trace, one log, one place to set a breakpoint. AReaL trades that away for throughput on multi-GPU clusters, and the trade is deliberate. The second comparison is an agent framework such as the ones AReaL already integrates with, including OpenAI Agents, LangChain, CAMEL and the OpenHands, Claude and Anthropic SDKs listed as dependencies. Those frameworks give you the agent loop and no training. AReaL gives you the training loop and treats the agent as an external service reached over HTTP. The difference in practice: with an agent framework you own the trajectory format and hand it to a trainer; with AReaL you keep your agent code untouched and let the RL service observe it through the endpoint. The scaffolding integration announced in April 2026, built on NVIDIA TensorRT-LLM's scaffolding module, is the middle path, separating agent execution, reward calculation and trajectory acquisition so those pieces can be reused. If your agent logic is already modular, that example is worth reading before you commit to the base_url approach.

## Maintenance, licensing and what an upgrade actually costs

The repository is not archived, and the last push was on 2026-09-09, about a week before this writing. Release cadence over 2026 has been roughly every six to eight weeks: v1.0.4 in May, v2.0.0 in July, v2.1.0 in August. That pace is the upgrade cost. Every one of those releases in the v2 line can move service boundaries, and the microservice layout means an upgrade is not a single pip command but a coordinated redeploy of training, inference, agent and weight-update components plus the config files that wire them. Pin your version and read the release notes before moving. The licence is Apache-2.0, declared both in the LICENSE file and in pyproject.toml, which permits commercial use and modification with the usual attribution and notice requirements. Note that Apache-2.0 covers AReaL's own code; the base image is lmsysorg/sglang, and several integrations pull in third-party SDKs under their own terms. That is a factual boundary, not legal advice, and a legal review of the full dependency tree is a separate exercise. Ascend NPU support lives on a separate ascend branch rather than main, so NPU users track a different line of development.

## Conclusion

AReaL fits teams that already run multi-GPU Linux clusters and want to point an existing black-box agent at an RL service by changing base_url. It does not fit laptop prototyping, macOS work, or anyone who needs a stable API: pyproject.toml classifies the package as Development Status 2 - Pre-Alpha, and the v2.0 release refactored the code into separate training, inference, agent and weight-update services. Before committing, read the v2 technical report, confirm which service boundary your rollout logic must cross, and check whether the AReaL-lite path or the full microservice path matches your team size.

## FAQ

### What is AReaL?

AReaL is a reinforcement learning infrastructure project for bridging foundation model training with agent-based applications. It is written in Python, licensed Apache-2.0, and built around a fully asynchronous RL training paradigm for reasoning and agentic models.

### How do I install AReaL?

The repository uses uv as its build backend and ships a uv.lock file, with requires-python set to >=3.11,<3.13 and a Linux plus NVIDIA CUDA 12 target. A Dockerfile is also provided that supports sglang and vLLM variants through the VARIANT build argument.

### Can I use AReaL with an existing agent without changing its code?

The README states that you can do online RL training for black-box agent applications by simply replacing the base_url, and cites an OpenClaw example that swaps the base_url and api_key for AReaL's RL service with no code changes.

### Does AReaL run on macOS or Windows?

No. The classifiers in pyproject.toml list only Operating System :: POSIX :: Linux, and the environment classifier names NVIDIA CUDA 12. The x86_64 macOS torch pin exists so dependency resolution succeeds, not as a supported training target.

## Sources

- [areal-project/AReaL on GitHub](https://github.com/areal-project/AReaL)
- [License: Apache-2.0](https://github.com/areal-project/AReaL/blob/main/LICENSE)
- [Project website](https://areal-ai.io)
- [README](https://github.com/areal-project/AReaL/blob/main/README.md)
- [Releases](https://github.com/areal-project/AReaL/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/areal-project-areal
