MARTI keeps multi-agent interaction central and policy training distributed, and its wheel tag is hand-built
[ICLR 2026] A Framework for LLM-based Multi-Agent Reinforced Training and Inference
At a glance
- What is it?
- A Tsinghua framework, accepted at ICLR 2026, for training multi-agent LLM systems with reinforcement learning, now extended with tree search for code generation. The design principle is one sentence and the module list is three. The packaging is where the sharp edges are: the wheel tag is assembled by a custom build class, the platform tag is hardcoded to a legacy manylinux value, and a nightly build silently renames the distribution.
- Who is it for?
- Use it if you are already running multi-agent reinforcement learning on a cluster and want heterogeneous models in one graph, since that is the capability the module split is built around and it is not something a single-agent trainer gives you. Expect to do your own environment work: the install is a requirements file and a pointer to separate setup instructions for three large dependencies, and the framework ships no tagged release, so pin a commit.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 46 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 5, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Interactions central, policy training distributed, three modules
The organising principle is stated once and everything else follows from it: all agent interaction and reward allocation happen centrally, while policy training is distributed across individual agents. That split is why there are three modules rather than a pipeline. The Multi-Agent World owns the interactions, Centralized Rewarding computes the reward and the credit assignment, and a Single Agent Trainer handles the optimisation for one agent at a time. Reward shaping and credit assignment are built in rather than left to you, which is the part that distinguishes this from wrapping a single-agent trainer in a loop. Workflows are graph-based, and the named patterns are debate, chain-of-agents, and mixture-of-agents. Several RL algorithms are supported, including PPO, GRPO, REINFORCE++, and TTRL, and inference and training live in the same framework rather than in two codebases.
GSPO and TIS are the two techniques carrying version two
The second version adds tree search-augmented reinforcement learning, released under the name MARS squared, aimed at complex reasoning tasks such as code generation. Two specific techniques do the heavy lifting and both target a known failure of long-horizon RL. GSPO replaces token-level policy optimisation with sequence-level, which the documentation argues suits complex reasoning better than the token-level objective in PPO. TIS, truncated importance sampling, corrects the distribution shift that appears in long sequence generation and, more concretely, corrects vLLM sampling bias during rollout. Two smaller mechanisms sit alongside them: dynamic data filtering and an overlong buffer for token penalties. The stated target is sequences up to 32K tokens, which is the regime where both corrections matter. The implementation sits on top of the OpenRLHF infrastructure, itself built on a single-agent RL framework alongside verl.
Heterogeneous training puts two models in one graph with separate filters
A single agent graph can hold different models, and version two leans on that. The stated capability is training different models simultaneously, with independent roles, independent training strategies, and dynamic sample filtering applied per agent, and the worked example pairs an 8B Qwen model with a different 8B model. That is a meaningfully different thing from a homogeneous ensemble, because each participant can be tuned for its own role in the conversation while the shared interaction graph keeps them coordinated. It also explains why the framework needs centralized rewarding with per-agent filtering rather than one global batch. A caveat sits in the documentation rather than in the code: integration with third-party multi-agent frameworks, specifically AutoGen and CAMEL, is supported but marked experimental, so the graph-based workflows are the path to build on.
The wheel tag is assembled by hand and the platform tag is hardcoded
The build configuration is where this repository stops being a research artifact and starts being something you can trip over. A custom wheel build class overrides option finalisation to force the wheel to be marked as not pure Python, which is correct for a package with compiled parts. It then builds the compatibility tag by hand rather than letting the build backend derive it. The Python tag is the interpreter major and minor, and the ABI tag is set to the exact same string, so a CPython 3.11 build is tagged with an ABI tag that names the interpreter version rather than an ABI. The platform tag is hardcoded to the first-generation manylinux tag on Linux, and on every other system it is the lowercased operating system name, which is not a platform tag format the ecosystem recognises.
A nightly build renames the distribution and stamps the date
The version does not live in the setup file. It is read from a version file at the top of the tree, which is the right call for a project with no tagged releases. An environment variable switches the build into nightly mode, and that switch changes two things rather than one. The version string gains a development suffix built from the current date in year-month-day form, so two nightly builds on the same day are the same version. More significantly, the distribution name itself changes from the normal name to a nightly variant, which is what stops a nightly install from silently overwriting a release install in the same environment. Both are quiet behaviours triggered by a single variable, and neither is mentioned in the installation instructions, which cover only cloning, installing requirements, and then following separate setup instructions for three heavy dependencies.
The codebase shipped in May 2025 and the headline feature in February 2026
The news timeline records five milestones and they are not evenly spread. The framework code was released on 2025-05-27. Async tool use and an async workflow for multi-agent reinforcement learning arrived on 2025-08-05. In October 2025 two external works were reported as built on the framework, one of them accepted at EMNLP 2025. Acceptance at ICLR 2026 came on 2026-01-25, and the version two release with its technical report followed on 2026-02-10. The last commit was 2026-08-20. With no GitHub releases at any point, there is no tag that corresponds to the version two announcement, so the only way to get that code is to take the default branch, which has moved six weeks past its last push. There is no statement of supported Python or CUDA versions anywhere in the documentation.
Requirements: thirty-two packages, five of them pinned exactly
The requirements file lists about thirty packages and pins five to exact versions: a deep learning library at 0.18.0, an attention implementation at 2.8.3, the Ray default extra at 2.48.0, the tensor library at 2.9, and the transformers release at 4.57.0. Four more carry floors rather than pins, including a serving engine at a post-release build above 0.8.5. Installation is the clone and the requirements file:
git clone https://github.com/TsinghuaC3I/MARTI.git
cd MARTI
pip install -r requirements.txtAfter that you are sent to separate setup instructions for three of its largest dependencies. The five exact pins are the ones that decide whether an install works, since an attention extension built from source against a specific tensor release is the usual failure point on a fresh machine. Everything else floats, including the optimiser libraries and the experiment tracker, so two installs a month apart can resolve to different versions of the same package.
A top-level directory named assert, and seven example folders
The tree is small and mostly conventional: a package directory, examples, docs, data, a requirements file, a setup script, a version file, and a licence. One entry is worth naming because it will stop you if you try to import it, since assert is a Python keyword and a directory with that name cannot be imported as a module even though nothing prevents it existing on disk. The examples directory is the more useful part of the layout, split into seven folders that map onto the framework's capabilities rather than onto its modules: the tree search example, multi-agent, multi-turn tool use, plain Python, the framework the external review work used, a schema example, and single-agent. That last one is the useful precedent: it shows how the same training entry points are called when there is only one participant.
Editorial conclusion
Use it if you are already running multi-agent reinforcement learning on a cluster and want heterogeneous models in one graph, since that is the capability the module split is built around and it is not something a single-agent trainer gives you. Expect to do your own environment work: the install is a requirements file and a pointer to separate setup instructions for three large dependencies, and the framework ships no tagged release, so pin a commit. If you intend to publish a wheel rather than run from a checkout, read the build configuration first, because the tag it produces is not one a package index will accept.
Frequently asked questions
What is the MARTI framework?
An open-source framework for training LLM-based multi-agent systems with reinforcement learning, accepted at ICLR 2026. It combines centralized multi-agent interaction with distributed policy training across three modules: Multi-Agent World, Centralized Rewarding, and a Single Agent Trainer, with graph-based workflows such as debate, chain-of-agents, and mixture-of-agents.
Which RL algorithms does MARTI support?
PPO, GRPO, REINFORCE++, and TTRL, with inference and training in one framework. It builds on OpenRLHF and verl, and supports the vLLM v1 engine along with a hybrid engine for faster training.
What does MARTI-v2 add?
Tree search-augmented reinforcement learning under the name MARS squared, aimed at complex reasoning such as code generation, with adaptive node expansion and refinement. It adds a sequence-level GSPO loss instead of token-level PPO, truncated importance sampling to correct distribution shift and vLLM sampling bias during rollout, dynamic data filtering, and an overlong buffer, targeting sequences up to 32K tokens.
How do I install MARTI?
Clone the repository, change into it, and install the requirements file. You then have to follow separate setup instructions for OpenRLHF, Ray, and vLLM. The requirements list about thirty packages with five pinned exactly, including deepspeed at 0.18.0, flash-attn at 2.8.3, Ray at 2.48.0, torch at 2.9, and transformers at 4.57.0.
Does MARTI work with AutoGen and CAMEL?
Integration with both is supported but marked experimental in the documentation, so the built-in graph-based workflows are the supported path. There is also a single-agent example in the repository showing the same entry points with one participant.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/tsinghuac3i-marti)