OpenPipe ART: GRPO training for multi-step agents
Agent Reinforcement Trainer: train multi-step agents for real-world tasks using GRPO. Give your agents on-the-job training. Reinforcement learning for Qwen3.6, GPT-OSS, Llama, and more!
At a glance
- What is it?
- OpenPipe ART is a Python framework that puts GRPO reinforcement learning behind an ergonomic harness so an existing agent loop can be trained on its own traces. It is aimed at engineers who already have a working agent and a reward signal, not at teams looking for a managed training service.
- Who is it for?
- Adopt OpenPipe ART if you already have a working agent loop, a programmatic reward, and either GPUs or a W&B API key, and you want the GRPO machinery handled for you. Do not adopt it if your task has no measurable reward, or if you cannot accept a Python 3.12 floor and a dependency set that pins torch 2.11.0, transformers 5.2.0 and unsloth 2026.3.3.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 1 day ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What OpenPipe ART actually trains
Most agent code is a loop: the model picks a tool, the tool returns something, the loop continues until a stop condition. That loop is usually written once and then frozen, because fine-tuning a model on multi-step behaviour requires turning each rollout into a training signal, and the plumbing for that is tedious. OpenPipe ART exists to remove that plumbing. The README describes it as "an open-source RL framework that improves agent reliability by allowing LLMs to learn from experience" and as "an ergonomic harness for integrating GRPO into any python application".
The audience is narrow on purpose. You need three things before ART is useful: an agent that already runs end to end, a reward you can compute in code, and a base model you are allowed to train. The example notebooks make the pattern concrete. There is an email-search agent (ART-E) built on RULER, a 2048 player, a Tic Tac Toe player, Codenames, Temporal Clue, an MCP server task, and a LangGraph variant. Each pairs a task with a reward, and the reward is what the model optimises against. If you cannot write that reward, ART has nothing to optimise.
The framing in the README is "on-the-job training" for agents. Read that literally. This is not a general fine-tuning library and it is not a prompt-optimisation tool. It is a trainer for policies that act over several steps.
The GRPO mechanism and the two backends
ART uses GRPO, group relative policy optimisation. The distinguishing feature of GRPO is that it does not train a separate value network. Instead it samples a group of completions for the same prompt, scores each one with the reward function, and uses the spread within that group as the baseline. A completion that scores above its group average is reinforced; one below is suppressed. That removes the critic model that PPO needs, which is why GRPO fits agent rollouts where a value estimate over long trajectories would be hard to learn.
The repository shows two ways to run that training. The first is a local backend, which the .env.example implies when it says the WANDB_API_KEY is worth setting "if you're going to ssh into a new machine and use the local backend". The second is serverless. The README introduces "W&B Training (Serverless RL)" and gives a code sample in which a TrainableModel is registered against a ServerlessBackend, with the model declared as Qwen/Qwen3.6-27B and the backend constructed from a W&B API key. The same snippet is presented as the alternative to hours of GPU setup, and the README's own bullet points claim 40% lower cost, 28% faster training, and scaling to 2000+ concurrent requests. Those are the project's marketing figures, not measurements I can confirm.
The dependency layout supports the split. The base install is light: aiohttp, anthropic, openai, pydantic, litellm, polars, numpy, scipy and a few others. The heavy pieces sit behind optional extras. The backend extra pulls in peft, bitsandbytes, unsloth, torchao, trl, wandb, duckdb and pyarrow, with torch pinned to 2.11.0+cu128 on Linux and Windows and plain 2.11.0 on macOS. There is also a distributed extra with torchmonarch and transformers, and a separate distributed-cu130 variant for CUDA 13. The repository also carries vllm_runtime/ and megatron_runtime/ directories at the top level, which indicates inference and training runtimes are shipped as part of the tree rather than assumed to be installed separately.
Installing openpipe-art and running a first training job
The project is on PyPI as openpipe-art, and pyproject.toml declares requires-python >=3.12, so a 3.12 or newer interpreter is the floor. The README links a Colab notebook for a first hands-on run, and the .env.example is the file to copy for local configuration. The pyproject.toml file lists the optional extras, and the backend extra is the one that brings torch, unsloth and trl.
The README's serverless example is the shortest complete path to a training run. It constructs a TrainableModel with a project name, a run name and a base model, builds a ServerlessBackend from a W&B API key, and registers the model against it.
from art.serverless.backend import ServerlessBackend
model = art.TrainableModel(
project="voice-agent",
name="agent-001",
run_name="agent-001",
base_model="Qwen/Qwen3.6-27B"
)
backend = ServerlessBackend(
api_key="your_wandb_api_key"
)
model.register(backend)After registration the README says every checkpoint is available through W&B Inference. What you should see is a registered model handle; the actual training loop is defined by your environment and reward, which is the part the README leaves to the notebooks. If you would rather not start from scratch, the README links Colab notebooks for 2048, ART-E and the others, and those are the fastest way to see a full run.
The .env.example is the configuration template. It lists WANDB_API_KEY first and marks it as the one to set for the local backend, and it notes that HF_TOKEN is optional for most models but necessary for training gated models like Llama 3.1.
# I recommend setting your API key here if you're going to ssh into a new machine and use the local backend
WANDB_API_KEY=YOUR_WANDB_API_KEY
# HuggingFace Token (optional for most models, necessary for training gated models like Llama 3.1)
HF_TOKEN=YOUR_HUGGINGFACE_TOKENThe README does not give a pip install line, so the package name comes from pyproject.toml, where the project is declared as name = "openpipe-art".
Where ART is the wrong tool
The strongest limitation is the reward. GRPO needs a scalar that can be computed for a whole trajectory, and it needs enough variance within a sampled group for the relative comparison to mean anything. If every rollout in a group scores identically, the advantage is zero and nothing is learned. Tasks with sparse binary outcomes, where almost every attempt fails, give you exactly that problem. The notebooks pick games and retrieval tasks because those have dense, cheap, deterministic rewards. A customer-support agent graded by a human reviewer does not.
The second constraint is the dependency surface. The backend extra pins torch to 2.11.0 with a CUDA 12.8 build, transformers to 5.2.0, trl to 0.20.0, unsloth to 2026.3.3 and accelerate to 1.7.0. If your environment already has a different torch or transformers, you are looking at a separate virtualenv at best. The distributed extra is stricter still, requiring transformers >=5.2.0,<=5.12.1 and a torchmonarch dependency. There is no documented path for running the backend extra against an existing CUDA 11 or ROCm stack.
Third, Python 3.12 is a hard floor. Teams on 3.10 or 3.11 cannot install it without an interpreter upgrade. Fourth, the README does not document rollback, checkpoint resumption or how to recover a run that dies mid-training. The .env.example mentions S3 configuration "for log and model backups" and lists BACKUP_BUCKET, but the README does not explain the restore procedure. Treat that as an open question to resolve from the docs before you depend on it.
ART against TRL and the rest of the stack
The obvious alternative is TRL, and the relationship is closer than it looks: trl==0.20.0 is itself a dependency of the backend extra. TRL is a general post-training library. It gives you GRPOTrainer, SFTTrainer, DPO and the rest, and it expects you to supply a dataset of prompts and a reward function that operates on model outputs. It is model-agnostic and task-agnostic, and it is the layer ART builds on for the actual optimisation step.
The difference is what sits above the trainer. TRL has no concept of an agent loop. If your task is multi-step, you write the rollout yourself, collect the completions, attach rewards, and feed the result in the format GRPOTrainer expects. ART's contribution is that harness: the TrainableModel abstraction, the backend split between local and serverless, and the example notebooks that show a full agent task wired to a reward. You are trading control over the training loop for a shorter path from a working agent to a training run.
There is also the serverless route, which is not really a library comparison. W&B Training manages the inference and training infrastructure, and the README presents it as the way to avoid GPU management. If you already run your own cluster, the local backend plus the distributed extra is the comparable option, and the trade-off is operational work against a dependency on the W&B service and API key.
Maintenance, licensing and upgrade cost
The repository is not archived and the last push was on 2026-09-10, so it is being worked on. The release cadence is uneven: v0.5.16 landed on 2026-03-04, v0.5.17 on 2026-03-13, and v0.5.19 on 2026-08-14, with pyproject.toml already carrying version 0.5.20. That is a five-month gap between v0.5.17 and v0.5.19, which is worth knowing if you plan to track main.
The version numbers still sit in the 0.5.x range, which usually signals that interfaces can move. The pinned dependencies make that concrete: torch, transformers, unsloth and trl are all exact pins or narrow ranges, so a minor ART release can force a rebuild of your environment. Budget for that. The extras also mean you can keep the base install in one environment and the training stack in another, which softens the upgrade cost if your agent code only touches the base package.
The licence is Apache-2.0, which permits commercial use and modification and includes an explicit patent grant. The repository also ships a THIRD-PARTY-NOTICES file and a licenses/ directory, and the backend extra pulls in unsloth, which has its own licence terms. If you plan to redistribute a trained model or a bundled environment, read those files rather than assuming the Apache-2.0 header covers everything in the tree. This is a description of the licence files present, not legal advice.
Editorial conclusion
Adopt OpenPipe ART if you already have a working agent loop, a programmatic reward, and either GPUs or a W&B API key, and you want the GRPO machinery handled for you. Do not adopt it if your task has no measurable reward, or if you cannot accept a Python 3.12 floor and a dependency set that pins torch 2.11.0, transformers 5.2.0 and unsloth 2026.3.3. Verify first that your reward function can score a full trajectory rather than a single turn, and check the docs at art.openpipe.ai for the current backend options before you commit to one.
Frequently asked questions
Is OpenPipe ART free or paid?
The library itself is open source under Apache-2.0, so installing openpipe-art from PyPI costs nothing. The README also documents a serverless option, W&B Training (Serverless RL), which is a managed service built around a W&B API key and is presented with its own cost claims.
How do I install OpenPipe ART?
Install the openpipe-art package from PyPI, which requires Python 3.12 or newer. For local training you also need the backend extra, which brings in torch, unsloth, trl and the rest of the training stack.
Which models can OpenPipe ART train?
The description names Qwen3.6, GPT-OSS and Llama among others, and the notebooks train Qwen 3.6 27B, Qwen 2.5 7B and Qwen 2.5 3B. The .env.example notes that HF_TOKEN is necessary for gated models such as Llama 3.1.
Does OpenPipe ART require a GPU?
The local backend does, since the backend extra installs CUDA 12.8 torch wheels on Linux and Windows. The serverless path is presented as the alternative that avoids managing that infrastructure yourself.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/openpipe-art)