castorini/pyserini: first-stage retrieval research where the pip install pulls in PyTorch
Pyserini is a Python toolkit for reproducible information retrieval research with sparse and dense representations.
At a glance
- What is it?
- Pyserini is an Apache-2.0 Python toolkit for reproducible information retrieval experiments, bridging Lucene through Anserini and dense search through Faiss. Its dependency list is the story worth reading: torch, transformers and an ONNX runtime are required to run a lexical BM25 search, a Java runtime is required at all, and Faiss is deliberately left out for a reason the project explains.
- Who is it for?
- Pyserini is the right tool if you are running retrieval experiments you need to be able to repeat, and you accept that the environment is large and that some results depend on prebuilt indexes rather than on your own indexing code.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 2 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 28, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The install command and the three heavyweight things it drags in
Installation is one line:
pip install pyseriniThe documentation then warns you about what that line costs, and the warning is unusually blunt. A pip installation automatically pulls in major dependencies including PyTorch, the Transformers library and the ONNX Runtime. For a package whose most basic use is running BM25 over a Lucene index, that is an extraordinary amount of machinery, and it is worth asking why before complaining about it.
The answer is scope. Pyserini is not a search library, it is an experiment harness that covers four families of retrieval model through one interface: traditional lexical models such as BM25, learned sparse models such as uniCOIL and SPLADE, learned dense models such as DPR, Contriever and BGE, and hybrid models that fuse the dense and sparse results. Dense retrieval is a neural embedding workload, so a deep learning framework and an inference runtime are not incidental extras, they are the path through which half the supported models run. Splitting them into an extra would have been defensible. The project chose to make the default install complete instead, which trades a smaller environment for an environment that works without further thought.
The project also flags the failure mode that follows from that choice. pip sometimes fails to resolve the right versions of these dependencies, and the advice is to install them by hand first, possibly with a version manager, then install Pyserini. That is the cost of a large pinned dependency graph, and it is the first thing to expect.
The package metadata confirms the scale. The declared Python floor is 3.12 or newer, and the dependency list includes a data loading library, a web framework, a plotting library, a numerical stack, an inference runtime, an OpenAI client, a data processing library, an image library, a JVM bridge, a tokenizer library, a machine learning toolkit, a scientific computing library, a sentence piece library, a token counting library, a deep learning framework, a progress bar, a model library, and an ASGI server.
Java 21 is a hard requirement, and pyjnius is why
The single most surprising line in the installation section is that Pyserini is built on Python 3.12 and Java 21, with the Java requirement attributed to the dependency on Anserini. A Python information retrieval toolkit requiring a specific major Java runtime is not a typo.
Sparse retrieval in Pyserini is not implemented in Python. It is delegated to Anserini, the group's own toolkit, which is itself built on Lucene, and the badge in the README pins the Lucene version the pairing targets at 10.5.0. The bridge between the two is a dependency named `pyjnius`, which is a Python interface to the Java virtual machine built on the Java Native Interface. That single dependency is the reason the project cannot be pure Python, and it is the reason a Java runtime version is part of the installation contract rather than an optional extra.
The consequences ripple outward in ways the documentation is upfront about. You cannot install Pyserini into an environment with no JVM, which rules out some slim container images. You inherit the startup cost of a JVM next to whatever your Python program was already doing, so a process that runs both pays two runtime initialisations. And you inherit the JVM's failure modes, including the fact that a search call that goes wrong surfaces a Java stack trace rather than a Python one.
The upside is that the lexical half of the toolkit is not a reimplementation. You are driving Lucene, through a toolkit the same group maintains specifically to expose it to researchers, which is why the project can claim it handles traditional lexical models, learned sparse models and the analyzer and query builder APIs on equal footing with the dense side. Reproducing a paper's BM25 baseline means running the same Lucene, not an approximation of it.
The dependency list also carries a second, older web framework alongside the current one, which is discussed in its own section below.
Faiss is left out on purpose, and indexes are the product
Dense retrieval runs through Faiss, and Faiss is not in the dependency list. The project gives the reason plainly: there is a proliferation of variants, with separate CPU and GPU packages among them, so it is easier for the user to install the right variant themselves. That is a defensible call and a common one for packages that wrap a native library, but it means the advertised one line install does not give you working dense search, and a reader who follows the quick start and then asks a dense model for results will hit a missing import rather than a documented error.
Multimodal search, described as image search, is handled the other way. It lives in an optional extra, requested by appending the optional marker to the package name, and that extra pulls in a single additional dependency, a package that provides the multimodal feature layer for Pyserini. Keeping multimodal out of the default is easier to defend than keeping deep learning out, since the same argument applies but the weight is much smaller.
The other thing that sets this package apart is that the indexes come with it. The project describes itself as self-contained, shipping queries, relevance judgments, prebuilt indexes and evaluation scripts for many commonly used test collections, and its stated purpose is reproducible first-stage retrieval in a multi-stage ranking architecture. That last phrase is the one to sit with. Pyserini is deliberately the first stage, not the whole pipeline, and the packaged collections are the standard ones used in the field: MS MARCO, NaturalQuestions, BEIR, with a specific guide for the second version of the MS MARCO document corpus used in recent TREC retrieval-augmented-generation tracks.
So the deliverable is not an index you build, it is a run you can reproduce. Indexing your own corpus is supported through four documented paths, and the choice among them depends on the model family, since a BM25 index, a sparse vector index and a dense vector index are genuinely different objects built by genuinely different code.
The package metadata carries the evidence for that emphasis. Alongside the Python source, the distribution includes an OpenAPI description for the REST server, two JSON files mapping local topic and relevance judgment aliases, and a Java logging properties file packaged as Python data, which is the small detail that tells you how tightly the two runtimes are joined. The declared floors are equally revealing:
requires-python = ">=3.12"
dependencies = [
"fastapi>=0.124",
"numpy>=2",
"onnxruntime>=1.23",
"torch>=2.9",
"transformers>=5,<6",
]Two web frameworks, a REST API and an MCP server in the core install
The dependency list contains both a classic Python web framework and a newer one with an ASGI server beside it, an OpenAI client, and a Model Context Protocol server library. That combination is a design statement rather than an accident, and it is the fastest way to see where this project is heading.
The README announces two interfaces directly: a REST API, documented on its own page, and an MCP server, documented on its own page. Neither is behind an extra. Both arrive with the base install.
The OpenAI client in that list is the piece that needs an explanation, and the repository does not give one directly. The plausible reading, given the stated purpose of first-stage retrieval, is that the package is being set up to be driven by, or to drive, model calls somewhere in an experiment loop, and the client is there so that a user does not have to add it. What is not in evidence is a documented feature built on it, so treat the dependency as a signal of direction rather than a feature to plan around.
The presence of two web frameworks is easier to read. One of them is almost certainly the older interface to the same server and the other the current one, and shipping both means the framework migration has not been finished. For a user this is harmless, since you call an HTTP endpoint either way. For someone maintaining a fork it is a signal to check which one the REST documentation actually targets before assuming a framework is load bearing.
The MCP server is the more interesting of the two for a 2026 reader. Wrapping a retrieval tool in MCP means an assistant can search a prebuilt index, fetch documents and run an evaluation without anyone writing glue code, and doing that as part of the core package rather than a community project is a deliberate bet that experiment harnesses will be operated conversationally. The interfaces directory at the repository root suggests there is more integration surface than the two documented endpoints, and an agents directory at the top level suggests the project is also being made approachable to tools that read the repository itself.
What reproducibility actually costs in this design
The word appears throughout the documentation, and it is worth being precise about what it does and does not guarantee here, because the meaning is narrower than the marketing usually is.
What the project guarantees is that the code path is fixed. When you run BM25 through Pyserini you are running Lucene 10.5.0 through Anserini at a specific version, with a specific analyzer, driven by a specific scoring implementation. Published papers that used Pyserini can therefore be re-run rather than re-implemented, and a result on a standard collection can be compared against a number produced by someone else on a different machine, which is the actual research requirement.
What it does not guarantee is that your environment matches. The dependency floors in the package metadata are aggressive and forward looking: a web framework at or above 0.124, a numerical stack at 2 or newer, an inference runtime at or above 1.23, an OpenAI client at or above 2.12, a deep learning framework at or above 2.9, a tokenizer library at or above 0.12, and the model library constrained to the 5 series with an explicit upper bound below 6. Bounds like that prevent silent major upgrades, which is exactly right for a reproducibility tool. They also mean that pinning the Pyserini version alone does not pin the environment, and that a lock file is your responsibility.
The prebuilt indexes are the other half of the guarantee and the other half of the risk. They remove index building from the experiment, which is where most irreproducibility lives, and they remove a large download and a lot of tuning from your first day. They also mean your results depend on an artefact hosted outside your machine. If a prebuilt index is rebuilt with a different tokenizer or a different Lucene behaviour, your numbers move and your code does not change, and nothing in the version number of Pyserini would tell you. Recording which index you used alongside which Pyserini version is not paranoia; it is the minimum record needed to interpret your own result later.
Editorial conclusion
Pyserini is the right tool if you are running retrieval experiments you need to be able to repeat, and you accept that the environment is large and that some results depend on prebuilt indexes rather than on your own indexing code. It is a poor fit for production serving or for a slim application, since the mandatory dependency set includes a deep learning framework, a model hub library and an inference runtime that a lexical search never touches, and the indexing steps differ by model family. Verify first that your Java version satisfies the Anserini requirement and that you install the Faiss variant matching your hardware, then pin the package version along with the index you used, because a run that cannot be reproduced against the same prebuilt index is not the reproducibility the project is selling.
Frequently asked questions
Why does installing Pyserini require Java?
Sparse retrieval is delegated to Anserini, which is built on Lucene, and the two are joined through a JVM bridge dependency. That is why the project states a Java 21 requirement, and why an environment without a JVM cannot run Pyserini at all.
Why is Faiss not included as a Pyserini dependency?
Because there are many Faiss variants, including separate CPU and GPU packages, so the project expects you to install the one matching your own hardware. Dense retrieval will not work until you do, and a missing import rather than a documented error is what you will see first.
What does a basic pip install of Pyserini pull in?
Beyond the retrieval libraries it automatically pulls major dependencies including PyTorch, the Transformers library and the ONNX Runtime, because the toolkit covers dense and learned sparse models as well as lexical ones. The project also warns that pip can fail to resolve the right versions of these, in which case installing them by hand first is the suggested route.
Can I search with Pyserini through a REST API or an assistant?
Yes. The project documents both a REST API and an MCP server, and neither is behind an optional extra, so both arrive with the base install. The package data includes an OpenAPI description for the REST interface.
How do I make sure a Pyserini result is reproducible?
Record the Pyserini version and the specific prebuilt index you used, and lock the environment yourself, since the package's dependency bounds prevent major upgrades but do not pin exact versions. Prebuilt indexes are the other variable, since an index rebuilt differently moves your numbers without changing your code.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/castorini-pyserini)