# FlashRank: CPU-only reranking for retrieval pipelines

> FlashRank is a Python library that adds cross-encoder and LLM reranking to an existing search or RAG pipeline without Torch or Transformers. It is small and runs on CPU, but its default model is a trade-off, not a free win.

**PrithivirajDamodaran/FlashRank** — Lite & Super-fast re-ranking for your search & retrieval pipelines.  Supports SoTA Listwise and Pairwise reranking based on LLMs and  cross-encoders and more.  Created by Prithivi Da, open for PRs & Collaborations.

- Repository: https://github.com/PrithivirajDamodaran/FlashRank
- Stars: 1,008 · Forks: 72
- Language: Python
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/prithivirajdamodaran-flashrank

## The reranking gap FlashRank targets

Vector search and lexical search both return candidates ranked by a cheap similarity function. That function never sees the query and the passage together, so the top of the list is often wrong. A reranker scores each (query, passage) pair jointly and reorders the list before it reaches an LLM or a user. FlashRank is a Python library for exactly that final step, positioned by its README as the last leg of a larger retrieval pipeline.

The audience is narrow and specific. You have a working retriever, you want better ordering, and you do not want to add a GPU or a Torch dependency to do it. The README states the library needs no Torch or Transformers and runs on CPU, and setup.py confirms the runtime dependencies are tokenizers, onnxruntime, numpy, requests and tqdm. That dependency list is the whole pitch: a reranker you can drop into an existing service without changing the deployment shape.

## How FlashRank reranks: ONNX models and two reranker families

FlashRank supports two mechanisms, and the distinction matters more than the model list.

Pairwise and pointwise rerankers are cross-encoder based, with a maximum of 512 tokens. These are the default path. The model runs through onnxruntime, which is why the install is small. The README describes model choice as a size and precision trade-off: ms-marco-TinyBERT-L-2-v2 is the default at roughly 4MB, ms-marco-MiniLM-L-12-v2 is called the best cross-encoder reranker at roughly 34MB, rank-T5-flan is the best non cross-encoder reranker at roughly 110MB, and ms-marco-MultiBERT-L-12 covers 100+ languages at roughly 150MB.

Listwise LLM rerankers are the second family, with a maximum of 8192 tokens, and they require the listwise extra. rank_zephyr_7b_v1_full is a 4-bit quantised GGUF model of roughly 4GB. The README carries an explicit constraint here: the current integration supports a maximum of 20 passages in one pass, and sliding window logic is not yet added. That is a real ceiling, not a footnote.

Data flow is straightforward. You construct a Ranker with a model name and a max_length, build a RerankRequest with a query and a list of passages, and get back a reordered list. Metadata is optional, and the README suggests using your database ids from the retrieval stage as the id field.

## Installing FlashRank and running a first rerank

The README gives two install paths. The plain install covers the cross-encoder rerankers, which is what most people want first.

```bash
pip install flashrank
```

If you need the LLM based listwise rerankers, install the extra. This pulls llama-cpp-python, pinned in setup.py at 0.2.76.

```bash
pip install flashrank[listwise]
```

A first rerank starts by constructing the ranker. The README example sets max_length explicitly, and the default model is the ~4MB TinyBERT.

```python
from flashrank import Ranker, RerankRequest

ranker = Ranker(max_length=128)
```

You then wrap the query and candidate passages in a RerankRequest. The README notes that metadata is optional and that the id can be your retrieval-stage database id or a plain numeric index. The call returns the passages in the new order, and you feed that order to your LLM or your UI.

```python
query = "How to speedup LLMs?"
```

One configuration point deserves attention before you benchmark anything. The README states that max_length should be large enough to hold your longest passage plus the query, and that setting a longer value than needed, such as 512 for short passages, will negatively affect response time. The README suggests using a token estimator like Openai tiktoken to size it. If your longest passage is 100 tokens and your query is 16, the README's own arithmetic points to 128.

## Where FlashRank is the wrong tool

The default model is a 4MB cross-encoder. That is the reason the library is cheap to run and the reason its ranking precision is not the best available in its own model table. The README itself labels ms-marco-MiniLM-L-12-v2 as the best cross-encoder reranker and rank-T5-flan as the best non cross-encoder reranker, which means the default is a speed-first choice. If ranking precision is your primary constraint, you are meant to change the model, and the cost is roughly 8x to 27x the download size.

The multilingual story is more awkward. ms-marco-MultiBERT-L-12 supports 100+ languages, but the README's own comment says not to use it for English. There is also a separate miniReranker_arabic_v1 model described as the only dedicated Arabic reranker. So a genuinely multilingual pipeline is not a single-model configuration here.

The listwise path has the 20-passage ceiling noted above, and the README says sliding window support is yet to be added. If your retrieval stage returns more than 20 candidates and you want listwise reranking over all of them, this integration does not do it today.

Finally, the README states that detailed benchmarking is TBD. The project publishes a time chart image for the example code with the default model, but there is no published benchmark table in the README. Any latency claim you make about your own workload has to come from your own measurement.

## FlashRank versus a hosted reranking API

The natural alternative is a hosted reranking API, and the README's own framing invites the comparison. Hosted rerankers run a larger cross-encoder on the provider's hardware. You send the query and passages over the network and get scores back. You do not manage model files, you do not size max_length, and you get whatever model the provider runs.

The difference is where the cost lands. FlashRank runs in your process on CPU, so there is no per-call network hop and no per-invocation billing, and the README explicitly frames serverless deployments like Lambda as the case where lowest cost per invocation and shorter cold starts matter. A hosted API moves that cost to a per-request price and adds a network round trip, but it also removes the model download, the onnxruntime dependency and the max_length tuning from your plate. FlashRank also keeps passage text inside your infrastructure, which a hosted API does not.

There is no free option here. FlashRank trades peak ranking precision for control over latency, footprint and data path. If your reranking volume is low and your latency budget is generous, the hosted route is less work. If you are already running Python on CPU and want the reranker to look like any other local function call, FlashRank is the shape you want.

## Maintenance, licence and upgrade cost

The repository is not archived, and the last push was on 2026-07-11. The most recent release listed is 0.2.9 from 2024-11-29, described as minor fixes. Before that, 0.2.4 from 2024-04-30 was the stable release with LLM based listwise rerankers, and 0.1.64 from 2023-12-23 added Rank T5 support. The pattern is a small number of releases with substantive model additions separated by longer quiet periods, so plan for the model list to change slowly.

Upgrade cost is dominated by model files, not code. Changing model_name changes the download size, and for the listwise path it changes the runtime shape, since the listwise extra pins llama-cpp-python at 0.2.76. The README lists InRanker as a roadmap model, which means the supported set is expected to grow.

The licence is Apache-2.0, declared in setup.py and shown in the README badge. Apache-2.0 permits commercial use and modification with the usual notice and attribution conditions, but the models FlashRank loads are separate artifacts with their own licences. The README links a model card for each one. Check the model card for the model you actually deploy, because the library licence does not cover the weights. This is a description of what the repository states, not legal advice.

## Conclusion

Adopt FlashRank if you already have a retrieval stage and want reranking that runs on CPU with no Torch or Transformers dependency, especially for serverless or latency-sensitive deployments where the ~4MB default model matters. Do not adopt it if you need multilingual ranking out of the box, since the README explicitly says not to use ms-marco-MultiBERT-L-12 for English, or if you need to rerank more than 20 passages per pass with rank_zephyr. Before committing, verify your own latency at your real max_length setting and check that your chosen model actually improves ranking precision on your data, because the README states detailed benchmarking is still TBD.

## FAQ

### What is FlashRank?

FlashRank is a Python library that adds reranking to an existing search or retrieval pipeline. It supports cross-encoder based pairwise and pointwise rerankers and LLM based listwise rerankers, and the README states it needs no Torch or Transformers and runs on CPU.

### Why do we rerank in RAG?

The README positions reranking as the final leg of a larger retrieval pipeline, applied to search results before they are fed into an LLM. A reranker rescores query and passage pairs so the best candidates reach the model first.

### How do I install FlashRank?

The README gives two commands: pip install flashrank for the lightweight pairwise rerankers, and pip install flashrank[listwise] if you need the LLM based listwise rerankers. The listwise extra depends on llama-cpp-python 0.2.76 per setup.py.

### How do I choose the max_length setting in FlashRank?

The README says max_length should be large enough for your longest passage plus the query, and that setting it longer than needed, such as 512 for short passages, will negatively affect response time. It suggests estimating token density with a library like Openai tiktoken.

### Does FlashRank support multilingual reranking?

The README lists ms-marco-MultiBERT-L-12 as supporting 100+ languages but adds a note not to use it for English, and lists miniReranker_arabic_v1 as the only dedicated Arabic reranker. There is no single model covering all languages well.

## Sources

- [Issues](https://github.com/PrithivirajDamodaran/FlashRank/issues)
- [License: Apache-2.0](https://github.com/PrithivirajDamodaran/FlashRank/blob/main/LICENSE)
- [PrithivirajDamodaran/FlashRank on GitHub](https://github.com/PrithivirajDamodaran/FlashRank)
- [README](https://github.com/PrithivirajDamodaran/FlashRank/blob/main/README.md)
- [Releases](https://github.com/PrithivirajDamodaran/FlashRank/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/prithivirajdamodaran-flashrank
