# RankLLM: listwise reranking with open and hosted LLMs

> RankLLM is a Python package from the castorini group for reproducible information retrieval research with pointwise, pairwise and listwise rerankers. It wraps retrieval, reranking, evaluation and serving behind one CLI, and its main constraint is that the useful parts are optional extras you install one at a time.

**castorini/rank_llm** — RankLLM is a Python toolkit for reproducible information retrieval research using rerankers, with a focus on listwise reranking.

- Repository: https://github.com/castorini/rank_llm
- Website: http://rankllm.ai
- Stars: 660 · Forks: 101
- Language: Python
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/castorini-rank-llm

## The measurement problem RankLLM was built to solve

Reranking research has a reproducibility problem that has nothing to do with the models. A paper reports nDCG after reranking the top 100 BM25 documents on DL19, but the retrieval run, the prompt template, the candidate ordering and the evaluation script all live in different repositories. Two groups can implement the same listwise prompt and get numbers that differ for reasons nobody wrote down.

RankLLM targets that gap. It is a Python package that bundles retrievers, rerankers and evaluation into one pipeline, so the same command produces a reranked run and a score. The intended user is an IR researcher or a graduate student who wants to compare RankZephyr against RankVicuna against a hosted model on the same collection, and who needs the comparison to be repeatable. The README describes the project as a suite of rerankers with a focus on open source LLMs compatible with vLLM, plus RankGPT and RankGemini variants for the proprietary listwise case.

It is not a library you drop into a search service and forget. The design assumes you have a collection, a retrieval method and a trec_eval-style metric in mind. If you only want to reorder ten strings inside a web request, the machinery here is heavier than the task.

## How the reranking pipeline is wired

The rerankers come in three families, and the distinction matters for cost. Pointwise models such as MonoT5 score one query-document pair at a time. Pairwise models such as DuoT5 compare two documents. Listwise models take the whole candidate list in a single prompt and emit a permutation, which is where the project's focus sits.

The listwise path is the interesting one. A retrieval method supplies candidates, the reranker builds a prompt from the query and those candidates, and the model returns an ordering. The README notes that the project supports reranking with first-token logits only, an efficiency measure that avoids generating a full ranking string. Prompts are not hardcoded: since the 2025.07.23 release, custom prompt templates can be supplied as YAML files, and the release notes point to the Reasonrank integration as an example of plugging in your own prompt and language model.

Backends are selected by what you installed. Local Hugging Face and PyTorch rerankers come from the local extra, vLLM from the vllm extra, and hosted OpenAI, OpenRouter and Gemini rerankers from the cloud extras. The same Reranker entry point is reused across them, so switching from a local 7B model to a hosted API is mostly an installation and credential change rather than a code rewrite. The repository also ships a ResponseAnalyzer for looking at what the model actually produced, which is useful when a reranking score moves and you want to know whether the permutation changed or the prompt did.

## Installing RankLLM with uv and running a first rerank

The README states that uv is the canonical contributor workflow, with conda and pip kept as fallbacks. Start by installing uv and creating a repository-local environment on Python 3.12, which pyproject.toml requires.

```bash
curl -LsSf https://astral.sh/uv/install.sh | sh
export PATH="$HOME/.local/bin:$PATH"
git clone https://github.com/castorini/rank_llm.git
cd rank_llm
uv python install 3.12
uv venv --python 3.12
source .venv/bin/activate
uv sync --group dev
```

If you would rather not clone, the published package installs from PyPI. Note that the base install is deliberately thin: requirements.txt lists dacite, ftfy, huggingface-hub, pyyaml, requests and tqdm, and nothing else. Torch, vLLM, Pyserini and the provider SDKs all arrive through extras.

```bash
uv venv --python 3.12
source .venv/bin/activate
uv pip install rank-llm
uv pip install "rank-llm[vllm]"
```

With the environment ready, the packaged rank-llm command is the documented entry point for a full run. The README gives this example, which retrieves with BM25, reranks the top 100 candidates for the dl20 dataset with RankZephyr, and leaves results behind for inspection.

```bash
rank-llm rerank --model-path castorini/rank_zephyr_7b_v1_full --dataset dl20 \
  --retrieval-method bm25 --top-k-candidates 100
rank-llm view demo_outputs/rerank_results.jsonl
rank-llm evaluate --model-name castorini/rank_zephyr_7b_v1_full
```

The view subcommand reads the JSONL that the rerank step wrote, so you can confirm the permutation before spending time on evaluation. The evaluate subcommand scores the run. If you need the retrieval and evaluation path, install the pyserini extra instead, and remember that the README states Java 21 is required while JDK 11 is not supported. For an HTTP endpoint rather than a batch run, the documented command is rank-llm serve http --model-path castorini/rank_zephyr_7b_v1_full --port 8082, and rank-llm serve mcp --transport stdio starts the MCP server.

## Where the listwise approach breaks down

Listwise reranking sends the query and its candidate documents to a language model in one prompt. That is the source of its quality advantage and also its main operational limit. The README does not document a sliding window or a sharded listwise pass for candidate sets larger than a single prompt can hold, and the quick-start example fixes the candidate pool at 100 with --top-k-candidates. If your retrieval stage returns thousands of documents, the listwise path is not the tool for the whole set; you would use it as a second stage over a shortlist produced by something cheaper.

Cost and latency follow from the same design. A hosted reranker makes one API call per query with all candidates in the prompt, so token spend scales with candidate count, not with the number of documents in the collection. A local vLLM deployment moves that cost to GPU memory instead. Neither is a drop-in replacement for a cross-encoder that scores pairs in milliseconds.

There is also a dependency trap. The extras are not additive by default. The mcp extra pulls in openai, pyserini and vllm together, and the server extra aggregates api and mcp, so a small experiment can end up installing the full retrieval and inference stack. The README's feature matrix exists precisely because installing everything is the wrong default. Choose the extra that matches the workflow you are actually running.

## RankLLM against a cross-encoder reranker such as FlashRank

The obvious alternative for many teams is a small cross-encoder, the category that FlashRank packages. The difference is architectural rather than a matter of quality claims. A cross-encoder scores each query-document pair independently and sorts by score. RankLLM's listwise mode conditions the ordering of every candidate on the others in the same prompt, which is why it can move a document up because a better one was removed, and why the cost of one reranking pass grows with the candidate list rather than with a fixed per-pair budget.

That conditioning is the reason to choose RankLLM for research. It is also the reason not to choose it for a latency-sensitive endpoint. A cross-encoder runs on CPU, needs no API key, and returns a score you can threshold. RankLLM's listwise rerankers need a GPU or a hosted provider, and the output is a permutation you then evaluate with trec_eval-style metrics. The two tools answer different questions: one asks how relevant is this document, the other asks how these documents should be ordered relative to each other.

RankLLM does cover the cross-encoder case too, through its pointwise models such as MonoT5 and MonoELECTRA from the local extra. So the honest framing is not RankLLM versus cross-encoders but RankLLM as a harness that lets you run both and compare them on the same collection with the same evaluation code.

## Licence, maintenance and what an upgrade costs

RankLLM is Apache-2.0, and pyproject.toml declares license = { file = "LICENSE" }. That is a permissive licence, but it covers the toolkit, not the models you point it at. The README lists RankGPT and RankGemini as proprietary listwise rerankers, and the local models such as RankZephyr and RankVicuna are downloaded from Hugging Face under their own terms. Anyone planning to ship a reranked result commercially should check the model card separately from the repository licence. This is a description of the files, not legal advice.

The repository is not archived, and the last push was on 2026-09-07, so the project is being updated. The version string in pyproject.toml is 0.25.7, matching the release notes entry for the OpenRouter addition. Upgrades are not free, though. The optional dependency pins are tight: transformers>=5 in the local extra, vllm>=0.19.1 in the vllm extra, and pyserini==2.4.0 in the pyserini extra. A pinned Pyserini means the retrieval and evaluation path moves only when the project moves it, which is good for reproducibility and awkward if you need a newer Pyserini for an unrelated reason.

The other upgrade cost is interface churn. The README records that the rank-llm CLI arrived on 2026.03.26 and that the scripts under src/rank_llm/scripts/ now act as compatibility wrappers over it. If you built tooling against those scripts, the wrappers keep it working, but the documented surface is the CLI. Training is a further step: the training extra exists to keep finetuning dependencies out of base installs, and the README notes that flash-attn must be installed separately with --no-build-isolation for optimized training workflows.

## Conclusion

Adopt RankLLM if you need to compare listwise rerankers on standard TREC collections and want the retrieval, reranking and trec_eval steps in one reproducible command. Skip it if you need a low-latency production reranker, since every listwise pass sends the query and its candidate documents to an LLM. Before committing, verify two things yourself: that your Python is 3.12 or newer, which pyproject.toml requires, and that the pinned pyserini==2.4.0 in the pyserini extra resolves against the Java 21 install you plan to use.

## FAQ

### What is LLM ranking in the context of RankLLM?

It is the use of a language model to reorder a list of retrieved documents for a query. RankLLM implements this in three families: pointwise models that score one document at a time, pairwise models that compare two, and listwise models that take the whole candidate list in one prompt.

### How do I install RankLLM?

The README states that uv is the canonical workflow: install uv, clone the repository, run uv python install 3.12 and uv venv --python 3.12, then uv sync --group dev. The published package can also be installed with uv pip install rank-llm, and conda with pip install -e . remains a fallback.

### Which Python version does RankLLM require?

pyproject.toml sets requires-python to >= 3.12, and the installation instructions use uv python install 3.12. Older interpreters are not supported by the declared metadata.

### Does RankLLM need Java?

Only for the retrieval and evaluation path. The README states that Java 21 is required if you plan to use workflows via rank-llm[pyserini], and that JDK 11 is not supported.

### Can RankLLM serve a reranking endpoint?

Yes. The README documents rank-llm serve http --model-path castorini/rank_zephyr_7b_v1_full --port 8082 for an HTTP server and rank-llm serve mcp --transport stdio for the MCP server. Both depend on the api and mcp extras respectively.

## Sources

- [castorini/rank_llm on GitHub](https://github.com/castorini/rank_llm)
- [Issues](https://github.com/castorini/rank_llm/issues)
- [License: Apache-2.0](https://github.com/castorini/rank_llm/blob/main/LICENSE)
- [Project website](http://rankllm.ai)
- [README](https://github.com/castorini/rank_llm/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/castorini-rank-llm
