Model or dataset
agentscope-ai/Trinity-RFT avatar
agentscope-ai/Trinity-RFT

Trinity-RFT review: decoupled explorer, trainer and buffer for LLM reinforcement fine-tuning

Trinity-RFT is a general-purpose, flexible and scalable framework designed for reinforcement fine-tuning (RFT) of large language models (LLM).

703 stars83 forksPythonApache-2.0

At a glance

What is it?
Trinity-RFT splits reinforcement fine-tuning into three coordinated components and ships dozens of runnable examples. It is a research framework first: the payoff is modularity, the cost is an alpha-stage Python stack with heavy dependencies.
Who is it for?
Adopt Trinity-RFT if you are an RL researcher or agent developer who needs to swap algorithms and data operators without rewriting the training loop, and you already have GPU capacity or a Tinker backend account for the no-GPU path.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem Trinity-RFT solves, and who it is actually for

Most reinforcement fine-tuning codebases hard-wire three things together: the loop that generates experience, the loop that updates weights, and the code that shuttles data between them. Change the algorithm and you touch the sampler. Change the sampler and you touch the trainer. Trinity-RFT's answer is to name the three parts and let them run in coordination. The README describes the split directly: the Explorer generates experience data via agent-environment interaction, the Trainer updates model weights by minimizing losses on the data, and the Buffer pipelines data processing throughout the RFT lifecycle.

The README addresses three audiences, and the distinction matters when you decide whether to spend a day on setup. Agent application developers are pointed at a workflow tutorial for training LLM-powered agents in specific domains. Reinforcement learning researchers are pointed at plug-and-play modules intended for non-invasive customization of new RL algorithms. Data engineers are pointed at operators for cleaning, augmentation and human-in-the-loop pipelines. If your job is none of those three, the framework is probably an expensive way to run a supervised fine-tune.

How the Explorer, Trainer and Buffer divide the work

The architecture is a producer-consumer arrangement with a shared data layer rather than a monolithic training script. The Explorer talks to an environment and emits experience. The Buffer accepts that experience and applies processing steps, which is where dataset cleaning, augmentation and human-in-the-loop review live as operators. The Trainer consumes whatever the Buffer releases and minimizes a loss to update weights. Because the Buffer sits in the middle, the two compute-heavy sides do not need to agree on a data format at the call site; they agree on what the Buffer stores.

That indirection is also where the framework's complexity lives. The dependency list in pyproject.toml includes ray[default]>=2.50.0 for distributed execution, tensordict for the data containers, sqlalchemy with aiosqlite and psycopg2-binary for persistence, and networkx. The optional data extra pulls in py-data-juicer, and the comments in pyproject.toml note that PostgreSQL users need asyncpg while MySQL users need aiomysql, neither installed by default. In other words, the Buffer is not an in-memory list. It is a database-backed pipeline, and you inherit the operational surface of that choice.

The release history shows this arrangement being tuned rather than replaced. v0.6.0 added SGLang support and what the release notes call optimized fully async weights synchronization and scheduling to reduce bubbles, plus improved MoE training stability and an upgrade to verl 0.8.0. Earlier, v0.5.0 added a colocate mode for single-GPU scenarios and trainer driven weight synchronization. Those are all synchronization concerns between the Explorer and Trainer sides, which is consistent with the three-component design.

Installing Trinity-RFT and running a first example

The package is published on PyPI as trinity-rft, and pyproject.toml requires Python >=3.10,<3.13. The project also defines a console entry point, so installation gives you a trinity command. The packaging metadata in pyproject.toml declares the extras, and the README points readers at the tutorial site for the step-by-step install, so check that page for the command that matches your backend choice.

The optional extras are where you choose an inference engine. The vllm extra is pinned to a narrow range and sglang to an exact version, so the two are not interchangeable in one environment:

toml
vllm = [
    "vllm>=0.22.0,<=0.23.0",
]
sglang = [
    "sglang==0.5.13",
]

The agent extra installs AgentScope with its tuner component, and the data extra installs py-data-juicer for pipeline work:

toml
agent = [
    "agentscope[tuner]>=1.0.19,<2.0.0"
]
data = [
    "py-data-juicer>=1.4.3",
]

For a first real run, the repository ships examples/grpo_gsm8k/, which is the smallest recognizable target: GRPO on GSM8K. The README links a workflow tutorial at agentscope-ai.github.io/Trinity-RFT for the step-by-step version, and the examples directory carries the configuration for each run. Read the example's config before launching, because the parallelism settings and the inference backend are set there, not on the command line. Expect the first run to spend most of its time on engine startup and weight synchronization rather than on the optimizer step.

Where Trinity-RFT gets in your way

The classifier in pyproject.toml reads Development Status :: 3 - Alpha. That is the project's own label, and it is consistent with the release cadence: v0.5.1 in February 2026, v0.5.2 in April, v0.6.0 in June, each with bug fixes and optimizations listed alongside new features. Alpha software that moves this fast will rename config keys and change defaults between minor versions. Budget time for reading release notes before upgrading, and pin your installed version in whatever environment you build.

The dependency pins are the second friction point. vllm is constrained to >=0.22.0,<=0.23.0 and sglang to exactly 0.5.13. If another part of your stack needs a different vLLM, you cannot satisfy both in one environment. The base install also pulls transformers>=5.12.1 and datasets>=4.0.0, which are large and move quickly.

The third limitation is hardware. The framework is built around distributed execution through Ray and verl, and the README's news entries describe GPU-oriented work: fully async weight synchronization, MoE training stability, single-GPU colocate mode. The Tinker backend, added in v0.4.0, is explicitly described as being for users without GPUs, which tells you the default path assumes you have them. If your only machine is a laptop, the no-GPU route exists but it is a backend choice, not the mainline configuration the examples are written around.

How this differs from a monolithic RLHF stack

The obvious comparison is with frameworks that treat RLHF as a single training script with a reward model bolted on. In that design, the generation step, the reward computation and the policy update share one process and one data structure, and extending it usually means editing that script. Trinity-RFT's difference is the Buffer as a first-class component with its own persistence layer (sqlalchemy over SQLite by default, PostgreSQL or MySQL if you install the drivers). Data processing becomes operators applied in the Buffer rather than functions called inline.

The practical consequence is that a data engineer can change cleaning and augmentation without touching the trainer, and an RL researcher can add an algorithm as a module without touching the sampler. The cost is that you now run a database alongside your training job. For a single researcher reproducing a GRPO result on GSM8K, that is overhead with no benefit. For a team where three people with three different specialisms work on the same pipeline, it is the reason the framework exists. The examples directory shows the breadth this buys: grpo_gsm8k, dpo_human_in_the_loop, grpo_email_search, agentscope_websearch, and research code tied to published papers such as examples/mix_chord and examples/learn_to_ask.

Maintenance, licensing and what upgrading costs

The repository is not archived, and the last push was on 2026-09-09, which is recent relative to the v0.6.0 release on 2026-06-26. Work is ongoing. The upgrade cost is real, though. Because the framework tracks verl (v0.7.1 or newer in the dependency list, with v0.8.0 arriving in the v0.6.0 release notes) and pins inference engines to narrow ranges, an upgrade is rarely a single pip command. You are upgrading a coordinated set: verl, vLLM or SGLang, transformers, and Trinity-RFT itself. The release notes are the place to check what moved.

Licensing is Apache-2.0, which permits commercial use and modification provided you keep the licence and notices intact. That is a permissive choice, and it matters here because the framework is designed to be extended: you can copy an operator or an algorithm module into your own codebase without a copyleft obligation. Note that the dependencies carry their own licences, and a training stack that includes vLLM, Ray, verl and AgentScope is a stack you should review as a whole. Nothing in the repository suggests the project offers indemnification, and this is not legal advice.

Editorial conclusion

Adopt Trinity-RFT if you are an RL researcher or agent developer who needs to swap algorithms and data operators without rewriting the training loop, and you already have GPU capacity or a Tinker backend account for the no-GPU path. Do not adopt it if you want a stable, versioned API for production fine-tuning: pyproject.toml classifies the project as Development Status 3 - Alpha, and the dependency pins (vllm>=0.22.0,<=0.23.0, sglang==0.5.13) mean an unrelated upgrade can break the environment. Before committing, verify that your Python is in the >=3.10,<3.13 range, that you can install the optional extras you actually need (vllm, sglang, agent, data, openjudge), and that the example closest to your task runs end to end on your hardware.

Frequently asked questions

What is Trinity-RFT and what does it do?

It is a general-purpose framework for reinforcement fine-tuning of large language models, published as the trinity-rft package under Apache-2.0. It decouples the work into an Explorer that generates experience data, a Trainer that updates weights, and a Buffer that pipelines data processing.

How do I install Trinity-RFT and can I use it without a GPU?

The package is published on PyPI as trinity-rft and requires Python >=3.10,<3.13; the README points to the tutorial site for install steps, and the optional extras select the inference engine. For machines without GPUs, the README's release notes state that v0.4.0 added a Tinker backend for users without GPUs.

What is in the Trinity-RFT tutorial and examples?

The README links a tutorial site at agentscope-ai.github.io/Trinity-RFT with separate guides for developing a workflow, an algorithm and an operator. The repository's examples directory contains runnable configurations such as examples/grpo_gsm8k and examples/dpo_human_in_the_loop.

Official sources

  1. agentscope-ai/Trinity-RFT on GitHub
  2. License: Apache-2.0
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/agentscope-ai-trinity-rft.svg)](https://hysenlabs.com/projects/agentscope-ai-trinity-rft)