Model or dataset
llm-as-a-verifier/llm-as-a-verifier avatar
llm-as-a-verifier/llm-as-a-verifier

LLM-as-a-Verifier: fine-grained reward for agent trajectories

LLM-as-a-Verifier is a general-purpose framework that provides fine-grained feedback for any agent without requiring additional training. It achieves SOTA performance across coding, robotics, and medical agentic benchmarks.

3,291 stars267 forksPythonMIT

At a glance

What is it?
A Python framework that scores agent rollouts by reading the logprob distribution of score tokens, then ranks candidates with a pivot tournament. It is a verifier you call from your own harness, not a model you train.
Who is it for?
Adopt it if you already generate several rollouts per task and want a ranking step that costs O(Nk) verifier calls instead of a full round robin, and if you can point it at a backend that returns token logprobs. Do not adopt it if your only verifier credential is a plain Gemini API key, since the .env.example states Vertex AI is required for logprobs, or if you need a published release feed for every upgrade.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 42 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What LLM-as-a-Verifier is for, and who should reach for it

The problem is selection. An agent that samples five trajectories per task gives you five answers and no reliable way to say which one is best. Pass@1 measures whether any of them works; it does not tell you which. LLM-as-a-Verifier is a Python package whose job is to rank that pool and return an index, a score per candidate, and a full ordering.

The README frames the idea in three steps: use fine-grained scoring granularity, take the expectation over the full logprob distribution of the LLM score tokens, and scale repeated evaluation and criteria decomposition. The practical consequence is that the verifier does not emit a single integer grade. It reads the probability mass the model places on each score token and turns that into a continuous reward in [0, 1].

The audience is narrow and specific. You need a task where multiple candidate answers exist, a way to describe what a good answer looks like as named criteria, and an API key for a model that exposes token-level logprobs. Researchers reproducing the benchmark tables in the README are the clearest fit. Engineers running best-of-N over a coding agent are the second. Anyone who wants a turnkey reward model trained on their own data is in the wrong repository: the README states the framework provides feedback without requiring additional training.

How the fine-grained reward and the pivot tournament fit together

Two mechanisms do the work, and they are separable.

The first is the reward. `select` is built on a pairwise reward model, and the package exposes `compare` for the raw case: given a problem, two candidates and a criteria dict, it returns two fine-grained rewards in [0, 1]. The README example prints 0.99994 and 0 for a pair where one answer is clearly correct. That asymmetry is the point. A coarse grader that returns 1 or 0 loses the ordering information you need when both candidates are wrong but one is closer.

The second is the ranking procedure. Ranking N candidates by comparing every pair costs O(N squared) verifier calls. The README states that `select` instead runs a Probabilistic Pivot Tournament that ranks all N trajectories using O(Nk) pairwise verifications, where k is the number of pivots. The `pivots` parameter is the cost dial, and the README is explicit that more pivots means more comparisons and higher accuracy. That is a genuine trade-off rather than a hidden default, and it is the parameter most likely to matter in production.

Criteria are not decoration. They are passed as a dict of named questions, and each one is evaluated separately. In the trajectory example the two criteria are "Root cause" and "Verification". Decomposing the judgement this way is what lets the same framework cover coding, robotics and medical benchmarks, since only the criteria text changes.

Installing llm-verifier and running a first selection

The package installs from PyPI. A clone install is also documented for tracking the repository directly.

bash
pip install llm-verifier
bash
pip install -e .

Before any call goes out, the verifier backend needs a credential. The `.env.example` file says to copy it to `.env` and fill in one of the keys, and it states that whichever backend you use must expose token-level logprobs, because the fine-grained reward reads the score-token logprob distribution. `DEEPSEEK_API_KEY` is the default backend. `VERTEX_API_KEY` is the alternative, and the same file notes that Vertex AI is required because the plain Gemini API does not expose token logprobs. A third option is any OpenAI-compatible server that returns logprobs, set through `OPENAI_BASE_URL` and `OPENAI_API_KEY`.

The smallest useful call is best-of-N over a short list of candidates. The README gives this example:

python
import llm_verifier

problem = "Write a function that reverses a string."
candidates = [
    "def rev(s): return s[::-1]", "def rev(s): return s", "def rev(s): return ''.join(sorted(s))",
]

result = llm_verifier.select(
    problem=problem,
    candidates=candidates,
    criteria={"Correctness": "Does the code actually reverse the string?"},
)
print(result.index)   # index of the best candidate: 0
print(result.scores)  # candidate scores: [0.73104, 0.38446, 0.38449]

What you should see is `0` for the index and a score list where the first entry is clearly ahead of the other two. Note that the second and third candidates score almost identically, 0.38446 against 0.38449. Both are wrong, and the framework does not pretend to separate them by much.

The same reward can score progress rather than final answers. `track` takes a list of step descriptions and a list of checkpoint steps, and returns a score after each one. In the README example the scores climb from 0.00106 after reading the problem statement to 0.99978 after the fix is tested. That shape is what makes the output usable for progress tracking rather than only final selection.

The logprob requirement is the constraint that decides your setup

This is the limitation that will stop most first attempts, and it is easy to miss because it lives in a comment inside `.env.example` rather than in the installation section.

The fine-grained reward is computed from the distribution over score tokens. If your provider does not return token-level logprobs, the mechanism has nothing to read. The `.env.example` comment is blunt about the consequence for one popular route: Vertex AI only, because the plain Gemini API does not expose token logprobs. If you have a Gemini API key and not a Vertex one, the default path is closed to you.

The escape hatch is a local server. The README shows `vllm serve Qwen/Qwen3.5-9B` with `OPENAI_BASE_URL=http://localhost:8000/v1`, and `pyproject.toml` defines an optional extra for that route: `pip install "llm-verifier[vllm]"`. The same file states that vLLM 0.19 or later is needed for the structured-outputs choice constraint used by the score-tag prefill. That version floor is the kind of detail that turns into an afternoon of debugging if you install an older vLLM and wonder why scores look wrong.

Two more boundaries are worth stating plainly. The package is classified as Development Status 4 - Beta in `pyproject.toml`, so the API is not presented as frozen. And there are no retrieved releases, so the CHANGELOG.md file is the only upgrade record you have; there is no release feed to subscribe to.

Where it sits next to an LLM-as-a-judge setup

The obvious alternative is the thing most teams already have: an LLM-as-a-judge prompt that asks a model to grade an answer, usually on a small integer scale, and returns that grade as text.

The difference is in what gets read back. A judge returns a discrete label. LLM-as-a-Verifier returns the expectation over the logprob distribution of the score tokens, which is why the README can print 0.73104 and 0.38446 rather than 1 and 0. Those intermediate values are what let `track` produce a smooth progress curve across agent steps, and what let the tournament order candidates that are all partially wrong.

The second difference is the cost model. A judge setup that ranks N candidates with a full round robin needs O(N squared) calls. The probabilistic pivot tournament in this framework is documented at O(Nk) pairwise verifications, with `pivots` as the knob. For a pool of five trajectories the difference is modest; it grows with N.

What a judge setup buys you instead is provider freedom. It works with any model that can follow a grading prompt, including ones with no logprob access at all. If your only credential is a plain Gemini API key, a judge prompt will run and this framework will not. That is the trade: continuous rewards and a cheaper ranking procedure, in exchange for a hard dependency on token-level logprobs.

Reproducing the benchmark tables and adapting the verifier to your own task

The repository ships agent trajectories under `data/`, and the README states that each benchmark has its own runner. Running `python scripts/run.py` with no argument lists the available benchmarks.

bash
python scripts/run.py terminal_bench
python scripts/run.py swe_bench
python scripts/run.py medagentbench

The tournament defaults can be overridden on the command line. The README gives this example, which sets two pivots, eight evaluations, a fixed seed and a worker pool of 50:

bash
python scripts/run.py swe_bench --pivots 2 --n-evaluations 8 --seed 0 --max-workers 50

The self-verification runs are separate scripts. `scripts/run_bo3.py` and `scripts/run_bo5.py` reproduce the Terminal-Bench 2.1 best-of-3 and best-of-5 configurations, and the README states that scoring only needs `DEEPSEEK_API_KEY` in `.env` because the trajectories are already in `data/terminal_bench_2.1_trajs/`. That is the cheapest way to confirm your credentials and your logprob path work before you point the framework at your own data.

Adapting it to a new task is documented as three steps. Copy your trajectories into `data/task_name_trajs/`, then follow the remaining steps in the README and in `add_new_benchmark.md`. Benchmarks are defined in `llm_verifier/benchmarks.py`, which the README identifies as the place to add or tweak one. The README also points to a separate Claude Code plugin repository, TurboAgent, for generating criteria and writing a runner.

Maintenance, licence and what an upgrade actually costs you

The repository is not archived, and the last push was on 2026-08-20. That is roughly a month of quiet, which is normal for a research codebase between releases rather than a sign of abandonment, but it does mean there is no stream of small fixes to lean on.

The version in `pyproject.toml` is 0.2.0, and the README documents what changed in that release: a prefix-cache optimization reported at roughly 3.4 times fewer uncached input tokens on trajectory-heavy benchmarks, the Terminal-Bench 2.1 self-verification benchmark, a `deepseek-v4-flash` verifier backend, and token accounting through `llm_verifier.token_usage()`. The prefix-cache change is the one with real operational weight. Trajectory verification sends long, largely repeated context to the verifier, so uncached input tokens dominate the bill. If you are on 0.1.x, that is the reason to move.

The upgrade cost is bounded by the dependency set, which is small: `google-genai`, `openai` and `tqdm`, with `vllm>=0.19` as an optional extra. The licence is MIT, declared in `pyproject.toml` with a `LICENSE` file at the repository root. MIT is permissive, but it covers the framework code only. The verifier models you point it at carry their own terms, and the trajectories under `data/` may carry terms of their own. Check those separately; this is a description of what the repository declares, not legal advice.

There is no retrieved release history, so CHANGELOG.md is your only upgrade record. Read it before bumping the pin.

Editorial conclusion

Adopt it if you already generate several rollouts per task and want a ranking step that costs O(Nk) verifier calls instead of a full round robin, and if you can point it at a backend that returns token logprobs. Do not adopt it if your only verifier credential is a plain Gemini API key, since the .env.example states Vertex AI is required for logprobs, or if you need a published release feed for every upgrade. Before wiring it into anything, run python scripts/run.py terminal_bench on the shipped trajectories and confirm the scores you get back match the numbers in the README table.

Frequently asked questions

What is LLM verification?

In this framework, verification means scoring candidate answers or agent trajectories against named criteria and returning a fine-grained reward in [0, 1] rather than a single pass or fail label. The reward is computed by taking the expectation over the logprob distribution of the model's score tokens.

What is the purpose of a verifier?

The README states the resulting fine-grained feedback can be used for test-time scaling, progress tracking and reinforcement learning. Concretely, `select` uses it to rank a pool of candidates and return the best index, and `track` uses it to score an agent's progress after each step.

What does LLM stand for?

The repository does not expand the acronym. Its README uses LLM in the project name and in the phrase LLM-as-a-Verifier, and the package itself is named llm-verifier, but no definition of the abbreviation appears in the README, the `.env.example` file or `pyproject.toml`.

What are the four types of LLM?

The repository does not describe a taxonomy of LLM types. Its README covers a verification framework that calls a verifier model and reads score-token logprobs, and it names specific backends such as `deepseek-v4-flash`, `gemini-2.5-flash` and `Qwen/Qwen3.5-9B` without classifying model families.

Official sources

  1. Issues
  2. License: MIT
  3. llm-as-a-verifier/llm-as-a-verifier on GitHub
  4. Project website
  5. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/llm-as-a-verifier-llm-as-a-verifier.svg)](https://hysenlabs.com/projects/llm-as-a-verifier-llm-as-a-verifier)