Model or dataset
PRIME-RL/TTRL avatar
PRIME-RL/TTRL

TTRL supervises a model with majority voting and then beats majority voting

[NeurIPS 2025] TTRL: Test-Time Reinforcement Learning

1,126 stars81 forksPythonMIT

At a glance

What is it?
Test-time reinforcement learning that turns an unlabeled test set into training data by using majority voting as the reward signal. The implementation lives in a vendored verl fork inside the repository, and the headline number is a 211 percent pass@1 gain on AIME 2024 with no labels at all.
Who is it for?
TTRL is a paper-with-code release worth reading for the idea rather than installing as infrastructure, because what it demonstrates is narrow and sharp: majority voting is not just an inference-time trick, it is a usable reward. That claim is what makes the result interesting and also what bounds it, since the same procedure inherits the biases of whatever the majority agrees on.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 174 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 6, 2026, and from our analysis. They are not legal advice.

Editorial analysis

Majority voting is repurposed as a training reward

The problem is stated as reinforcement learning on data without explicit labels for reasoning tasks in large language models, and the core difficulty is named precisely: estimating reward during inference when you have no ground truth. The answer is that a common practice from test-time scaling, majority voting, turns out to yield rewards effective enough to drive RL training. That reframing is the whole contribution. Majority voting normally computes an answer by sampling and taking the most frequent result; here the agreement pattern across samples becomes a scalar signal attached to individual attempts, which is exactly the shape an RL objective needs. No label is read at any point. The practical consequence is that a benchmark's test split, which is public and unlabelled, becomes a training set, and the paper is accepted to NeurIPS 2025.

The 211 percent gain comes with a caveat about maj@n

The headline result is that TTRL boosts the pass@1 of Qwen-2.5-Math-7B by approximately 211 percent on AIME 2024 using only unlabeled test data. The second claim is the more interesting one and is stated with its own bound: although TTRL is supervised only by the maj@n metric, it consistently surpasses that upper limit of the initial model, and approaches the performance of models trained directly on the test data with ground-truth labels. So the procedure is not simply distilling its own vote. What it approaches is a supervised-on-the-test-set model, which is the right yardstick to compare against. The epigraph the project opens with is a line attributed to David Silver and Richard S. Sutton about the era of experience, which is the frame the whole approach sits in.

The verl fork lives inside the repository, not beside it

The top-level layout explains a lot about how this was built and how you should treat it. There are only four entries: `LICENSE`, `README.md`, a `figs/` directory and a single `verl/` directory. So the framework is not a dependency to be resolved; it is vendored, and the setup instructions walk you into it:

bash
git clone https://github.com/PRIME-RL/TTRL.git

cd TTRL/verl

conda create -n ttrl python==3.10
conda activate ttrl
bash scripts/install_ttrl_deps.sh
pip install -e .

You clone one repository, change directory into the fork, install its dependency script and then install the fork in editable mode. That is a different maintenance posture from pip-installing a library, and it explains the release history: the GitHub releases attached to this project are verl at v2.0.0 and OpenRLHF at v1.0.0, upstream forks rather than TTRL releases. There is no TTRL version to pin.

Implementation is described as a reward function change

The reproduction section is short, which tells you how small the change is claimed to be. One script reproduces the AIME 2024 numbers:

bash
bash examples/ttrl/Qwen2.5/aime.sh

And the approach to implementing it is stated as being achieved rapidly by simply modifying the reward function, with a pseudo-code snippet sitting inside a collapsed block that renders as an empty section on the page. So the design claim is that TTRL is a reward swap rather than a new training loop, and that is consistent with the later news entry about verl: as of 2025-08-17, after bumping into verl v0.4.1, TTRL can be enabled by setting a single flag, `+ttrl.enable=True`. If that is true in the version you have, the fork may not be the only way to run it, since upstream gained the hook that TTRL needed.

Three preview runs gave 43.3, 43.3 and 46.7

The reproducibility note is unusually specific and worth taking at face value. The authors ran three independent runs using the preview version of the code, and report that two reached a pass@1 greedy of 43.3 while one reached 46.7, with the full logs on Weights and Biases. The spread is 3.4 points across three runs of the same configuration, which is a useful calibration for how much of the headline improvement you should attribute to the method rather than to seed variance. The compute behind it is stated too: all experiments were conducted on eight NVIDIA A100 GPUs with 80GB of memory each. That is a modest cluster by current standards, which is the most encouraging fact on the page for anyone wondering whether this is reproducible outside a large lab.

Data conversion and multiple benchmarks go through verl tooling

Three notes under the getting-started section define the surrounding workflow. A preprocessing script in the verl tree converts data from JSON format to Parquet for training with verl, so the expected input format is a specific one rather than anything you like. Scripts for running TTRL across multiple models and various benchmarks live in the `verl/examples/ttrl` directory, which tells you the scope of what has been exercised is broader than the single AIME run quoted above. And for anything below that layer the README points at the verl documentation rather than duplicating it, which is the correct boundary for a project whose contribution is a reward function rather than a framework. The paper is on arXiv at 2504.16084, and the citation is a BibTeX entry listing Zuo Yuxin first among the authors.

The follow-up work moved to a branch and a second conference

The news section shows where the project went after the paper. The most recent entry, dated 2026-03-10, describes an investigation into the mechanisms and applications of Unsupervised RLVR, abbreviated URLVR, with the finding that it is particularly well suited for test-time training and for quantifying model priors. The code for that line lives on a branch called `urlvr-dev` rather than on main, and the paper is accepted to ICLR 2026. Two smaller entries explain how the project became usable in the first place: the paper and code were updated on 2025-05-23 to be based on verl, and the code and experimental logs were released on 2025-04-24 after the project was presented on 2025-04-23. The last push to the repository is dated 2026-04-15, so main has been quiet since around when the URLVR branch work was published.

Editorial conclusion

TTRL is a paper-with-code release worth reading for the idea rather than installing as infrastructure, because what it demonstrates is narrow and sharp: majority voting is not just an inference-time trick, it is a usable reward. That claim is what makes the result interesting and also what bounds it, since the same procedure inherits the biases of whatever the majority agrees on. Three practical notes. The repository does not depend on verl, it contains it: the top level is a LICENSE, a README, a figures directory and a single `verl/` tree, so you are cloning a fork and installing from inside it with `pip install -e .`. The GitHub releases belong to that fork and its history, not to TTRL, so there is no TTRL version to pin. And the last push was on 2026-04-15, with a newer line of work on an `urlvr-dev` branch that followed the main paper into an ICLR 2026 acceptance.

Frequently asked questions

What is TTRL and what problem does it solve?

TTRL is test-time reinforcement learning, accepted to NeurIPS 2025, for training on data without explicit labels. The core difficulty is estimating reward at inference time without ground truth, and the answer is that majority voting, a common test-time scaling technique, produces rewards effective enough to drive RL training.

How does TTRL train on data with no labels?

The agreement pattern across sampled solutions, which is what majority voting computes, is used as a scalar reward attached to individual attempts. Nothing reads a label at any point, so a public test split such as AIME 2024 becomes a training set.

How do I install and run TTRL?

Clone the repository, change into the `TTRL/verl` directory, create a conda environment on Python 3.10, run `bash scripts/install_ttrl_deps.sh`, then `pip install -e .`. Reproduction on AIME 2024 is a single script, `bash examples/ttrl/Qwen2.5/aime.sh`. All experiments ran on eight NVIDIA A100 80GB GPUs.

Does TTRL depend on the verl project?

It contains it. The repository's top level holds only a LICENSE, the README, a figures directory and a `verl/` directory, so the framework is vendored rather than installed as a dependency. The GitHub releases attached to the project are verl v2.0.0 and OpenRLHF v1.0.0, not TTRL versions.

What results does TTRL report?

TTRL boosts the pass@1 of Qwen-2.5-Math-7B by approximately 211 percent on AIME 2024 using only unlabeled test data, and although supervised only by the maj@n metric it surpasses that upper limit and approaches models trained on the test set with ground-truth labels. Three independent runs of the preview code gave pass@1 greedy of 43.3, 43.3 and 46.7.

What is URLVR in relation to TTRL?

URLVR is the follow-up line of work, investigated as unsupervised RLVR and reported as particularly well suited to test-time training and to quantifying model priors. Its code is on a branch called `urlvr-dev` rather than on the main branch, and its paper was accepted to ICLR 2026.

Official sources

  1. License: MIT
  2. PRIME-RL/TTRL on GitHub
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/prime-rl-ttrl.svg)](https://hysenlabs.com/projects/prime-rl-ttrl)