Library / SDK
xhluca/bm25s avatar
xhluca/bm25s

bm25s: a pure-Python BM25 retriever built on Numpy and Numba

Fast BM25 search in Python, powered by Numpy and Numba

1,796 stars107 forksPythonMIT

At a glance

What is it?
bm25s implements Okapi BM25 in Python with eager sparse scoring, so query time is mostly matrix lookups. Here is how to install it, index a corpus, persist the index, and where its design stops being the right choice.
Who is it for?
Adopt bm25s if you need lexical retrieval inside a Python process, want no JVM or PyTorch dependency, and can accept that the index is a sparse matrix you manage yourself. Do not adopt it if you need incremental updates, distributed sharding, or a query DSL, because the README documents none of those.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 12 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The gap bm25s fills between rank-bm25 and Elasticsearch

BM25 is the ranking function behind most keyword search, and the README says it is a core component of services like Elasticsearch. The problem for Python developers is that the two obvious ways to get it are awkward. Running Elasticsearch means a JVM service, a separate process, and a client library between your code and the scores. Using rank-bm25 means a Python implementation that recomputes scores at query time, which the project's own comparison chart measures as the baseline for its speedup claims.

bm25s targets the middle: lexical search that lives in the same process as your application, with no Java and no PyTorch. The README states the only hard requirement is Numpy, with optional lightweight dependencies for stemming and for Numba compilation. That makes it a fit for retrieval-augmented generation pipelines, offline evaluation scripts, and small to mid-sized search features where standing up a cluster is disproportionate. It is not a service. It is a library that hands you arrays.

Eager sparse scoring: why query time is cheap

The design decision that separates bm25s from most Python BM25 code is when the work happens. The README describes it as storing "eagerly computed scores for all document tokens" in sparse matrices. Instead of walking the posting list and computing each document's score when a query arrives, the library precomputes per-token contributions at index time and stores them sparsely. A query then becomes a lookup and accumulation step over the rows corresponding to the query's tokens.

That trade is explicit: you pay memory and index build time up front to make retrieval fast. The README claims performance improvements over popular libraries by orders of magnitude, measured as throughput in queries per second on BEIR datasets in a single-threaded setting, and the comparison graphic includes Elasticsearch alongside rank-bm25. Those numbers come from the project's own benchmark repository, not from independent measurement, and the README does not state the corpus sizes behind each bar.

The tokenization step is separate from indexing, which is a deliberate split. bm25s.tokenize() converts text into token ids, optionally applying stopwords and a stemmer, and retriever.index() consumes those ids. Retrieval is position-based: document ID i maps to corpus[i], so the corpus list you pass around must stay in the same order and length as the indexed documents. The README calls this out directly, and it is the kind of constraint that quietly breaks a pipeline when someone sorts a list in between.

Installing bm25s with pip and running a first retrieval

The README gives a plain pip install, plus an extras group it labels highly recommended because it pulls in JSON loading, a progress bar, stemming, and JIT compilation. The full extra adds everything else.

bash
pip install bm25s
pip install "bm25s[core]"

The quickstart is short enough to paste into a file and run. Note that the stemmer is optional and imported separately, and that tokenize() is called twice: once for the corpus with stopwords, once for the query without them.

python
import bm25s
import Stemmer

corpus = [
    "a cat is a feline and likes to purr",
    "a dog is the human's best friend and loves to play",
    "a bird is a beautiful animal that can fly",
    "a fish is a creature that lives in water and swims",
]

stemmer = Stemmer.Stemmer("english")
corpus_tokens = bm25s.tokenize(corpus, stopwords="en", stemmer=stemmer)

retriever = bm25s.BM25()
retriever.index(corpus_tokens)

After index() returns, the retriever holds the sparse structures. Query it with tokenized input, and expect two arrays of shape (n_queries, k): document ids and scores. The README notes that passing corpus=corpus to retrieve() returns the documents themselves rather than ids.

python
query_tokens = bm25s.tokenize("does the fish purr like a cat?", stemmer=stemmer)
results, scores = retriever.retrieve(query_tokens, k=2)

for i in range(results.shape[1]):
    doc, score = results[0, i], scores[0, i]
    print(f"Rank {i+1} (score: {score:.2f}): {doc}")

Persistence is two calls. Saving with the corpus writes it alongside the index, and loading with load_corpus=True restores both; the README says to set load_corpus=False when you do not need the documents back.

python
retriever.save("animal_index_bm25", corpus=corpus)
reloaded_retriever = bm25s.BM25.load("animal_index_bm25", load_corpus=True)

One detail worth knowing before you design a schema: corpus entries can be dictionaries, and dictionaries have no required keys. The README shows a metadata_corpus with id, title, and text fields, where you tokenize only the text field but retrieve the whole dictionary. String entries are serialized to corpus.jsonl as {"id": i, "text": doc}, while dictionaries, lists, and tuples are written as provided, and all of them must be JSON-serializable.

What the index format costs you: updates, memory, and rollback

The eager sparse design has a direct consequence the README does not soften: the index is a snapshot. There is no documented API for adding or deleting a single document after index() runs, and no documented way to merge two indexes. If your corpus changes, the path the documentation supports is re-tokenizing and re-indexing. For a static corpus that is fine. For a feed that ingests documents hourly, it means rebuilding the sparse matrices on every cycle, and the cost of that scales with the corpus, not with the delta.

The second constraint is memory. Precomputed scores for every document token are stored, not computed on demand, so the resident footprint is larger than a posting-list implementation would need. The README does not publish a memory table, and the benchmark chart is about throughput. Before adopting this for a large corpus, measure the saved index directory size and the process RSS after load on your own data. That is the number that decides whether the trade works for you.

Third, the documentation is silent on rollback and versioning. The README does not document an index format version, a compatibility guarantee between releases, or what happens when a directory saved by one version is loaded by another. The release history shows 0.3.9, 0.3.10, and 0.3.11 landing in May, July, and August 2026, so the library is moving. If you persist indexes as build artifacts, pin the bm25s version alongside them and treat a version bump as a reason to rebuild rather than to trust the old directory.

Finally, this is lexical search only. There is no vector index, no hybrid scoring, and no query parser for boolean or phrase operators documented in the README. If your queries need stemming plus synonym expansion plus field boosting, you are assembling that yourself around tokenize().

bm25s versus rank-bm25 and Elasticsearch

The most useful comparison is with rank-bm25, because both are Python libraries for the same job and the difference is architectural rather than cosmetic. rank-bm25 computes scores at query time from term frequencies and document lengths. bm25s precomputes sparse score contributions at index time and looks them up at query time. The project's comparison chart uses rank-bm25 as the baseline and plots bm25s and Elasticsearch as multiples of it, measured in queries per second on BEIR datasets, single-threaded. If your workload is many queries against a fixed corpus, that ordering favors bm25s. If your workload is a few queries against a corpus that changes constantly, the precomputation is wasted work and the rebuild dominates.

Against Elasticsearch the difference is operational. Elasticsearch gives you sharding, replication, incremental indexing, analyzers, and a query DSL, in exchange for a JVM service to run and monitor. bm25s gives you an in-process object and arrays, in exchange for managing all of the above yourself or doing without. Neither is a downgrade of the other; they answer different questions. The README's own framing, that BM25 is a core component of services like Elasticsearch, is honest about the fact that bm25s is the component, not the service.

There is also a Numba path. The README marks version 0.2.0 as introducing a numba backend and links a release discussion claiming roughly 2x speedup for larger datasets. The repository ships examples/index_and_retrieve_with_numba.py and examples/retrieve_with_numba_advanced.py, which is where to look for the actual call pattern, since the README's quickstart does not use it.

Maintenance, licence, and what a release bump implies

The repository is not archived, and the last push was on 2026-09-10, with 0.3.11 released on 2026-08-25. The project is MIT licensed, which permits commercial use, modification, and redistribution provided the copyright notice and permission notice are included. That is a permissive licence with no copyleft obligation on your application code. It says nothing about the licences of the optional dependencies you add through the core or full extras, and the README does not enumerate them, so check what pip resolves if your organization audits transitive licences.

Upgrade cost is dominated by the index, not the code. If you persist indexes to disk, a version bump means re-tokenizing and re-indexing, because the README documents no format compatibility contract. Keeping the bm25s version pinned in the same artifact as the saved index, and rebuilding both together, is the cheapest way to avoid a load failure in production. The setup.py resolves versions from environment variables such as BM25S_VERSION and RELEASE_TAG and from git tags, which matters only if you build from source rather than from PyPI.

Editorial conclusion

Adopt bm25s if you need lexical retrieval inside a Python process, want no JVM or PyTorch dependency, and can accept that the index is a sparse matrix you manage yourself. Do not adopt it if you need incremental updates, distributed sharding, or a query DSL, because the README documents none of those. Before committing, verify the Numba path on your own corpus by running examples/index_and_retrieve_with_numba.py, and check the memory footprint of the saved index directory on your real document count.

Frequently asked questions

What are the disadvantages of BM25?

The README does not discuss BM25's ranking weaknesses. What it does show is a library-level constraint: bm25s precomputes scores for all document tokens, so the index is a snapshot with no documented incremental update path, and the memory footprint is larger than a query-time scorer's.

What does BM25 stand for?

The README identifies the algorithm as Okapi BM25, and the repository topics list okapi-bm25 alongside bm25-l and bm25-plus. The README does not expand the acronym further.

Is BM25 still relevant today?

The README states BM25 is a widely used ranking function for text retrieval and a core component of services like Elasticsearch, and it lists rag and retrieval among the repository topics. That is the project's position, not an independent assessment.

How is BM25 different from TF-IDF?

The README does not compare BM25 to TF-IDF. It describes bm25s as an implementation of BM25 that ranks documents against a query, and shows tokenization, indexing, and retrieval but no scoring formula breakdown.

Official sources

  1. License: MIT
  2. Project website
  3. README
  4. Releases
  5. xhluca/bm25s on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/xhluca-bm25s.svg)](https://hysenlabs.com/projects/xhluca-bm25s)