pg_bestmatch.rs makes BM25 vectors and leaves the index to somebody else
Generate BM25 sparse vector inside PostgreSQL
At a glance
- What is it?
- A PostgreSQL extension that turns BM25 statistics into sparse vectors for pgvecto.rs or pgvector, with an explicit warning that HNSW does not handle them well. Building it needs cargo-pgrx at an alpha version.
- Who is it for?
- pg_bestmatch.rs fits a team that wants BM25 lexical retrieval to live inside the same PostgreSQL transaction as the rest of its RAG pipeline, and that already runs pgvecto.rs or pgvector and is willing to index with ivfflat rather than HNSW. It does not fit a team that wants a self-contained search extension, since it generates vectors and nothing else.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Probably not. The repository last received commits 23 months ago, on November 5, 2024.
- What is it written in?
- Mainly Rust, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
HNSW does not handle these vectors, so the index is not this extension's job
The warning near the top of the page is the most important sentence in it, and it is flagged as IMPORTANT.
Based on the project's initial tests, HNSW indexing does not support the sparse vectors generated by BM25 very well, because the high sparsity prevents effective navigation within the graph. That is a structural mismatch rather than a tuning problem: BM25 vectors are mostly zeros, and a graph built to navigate between nearby dense points does not get good edges out of a vector that is zero almost everywhere.
The consequence shows up in the workflow. The optional index step in the usage example uses ivfflat for pgvector, not HNSW, and the extension itself does not implement index-based search at all. The page says so directly: it provides methods for generating sparse vectors, and index-based search is something pgvecto.rs or pgvector handles.
So the division of labour is: this extension produces the vectors, you choose the vector extension and the index type, and you accept that the obvious index choice does not work. That is a smaller commitment than a search extension, and it is also the whole of what this one is.
Installation is two statements, and the second one is easy to forget
Installing the extension into an existing database is:
CREATE EXTENSION pg_bestmatch;
SET search_path TO public, bm_catalog;Two statements, and the second one does real work. The extension keeps its statistics catalog in a schema called `bm_catalog`, and putting that schema on the search path is what makes the `bm25_` functions resolvable without qualifying every call.
This is the kind of step that gets lost. A deployment that runs `CREATE EXTENSION` in a migration and then relies on the connection's default search path will fail at the first `SELECT bm25_create(...)` with an unrecognized function, and the error will point at the SQL rather than at the missing `SET`. If you take one operational note from this page, make it that the `SET search_path` belongs in the same place as the `CREATE EXTENSION`, in every environment including development.
Note also what the page does not tell you. There is no statement about upgrading, and no mention of a schema version or a migration path for the catalog.
A third argument picks svector or sparsevec
The API is small, and the mechanism that trips people up is a single optional argument.
The functions are `tokenize`, `bm25_create`, `bm25_refresh`, `bm25_drop`, `bm25_document_to_svector` and `bm25_query_to_svector`. Scoring is a dot product between the query sparse vector and the document sparse vector, which means the distance operator belongs to whichever vector extension you installed rather than to this one.
Creating statistics takes the BM25 parameters directly:
SELECT bm25_create('documents', 'passage', 'documents_passage_bm25', 0.75, 1.2);The reference section documents `b` with a default of 0.75 and `k` with a default of 1.2, so the two positional values here are the degree of length normalisation and the term saturation parameter.
Then the conversion, and here is the flag:
ALTER TABLE documents ADD COLUMN embedding svector; -- for pgvecto.rs users
ALTER TABLE documents ADD COLUMN embedding sparsevec; -- for pgvector users
UPDATE documents SET embedding = bm25_document_to_svector('documents_passage_bm25', passage)::svector; -- for pgvecto.rs users
UPDATE documents SET embedding = bm25_document_to_svector('documents_passage_bm25', passage, 'pgvector')::sparsevec; -- for pgvector usersTwo arguments give an `svector` for pgvecto.rs. Adding the string `'pgvector'` as a third argument gives a `sparsevec` instead. The same pattern applies to `bm25_query_to_svector`.
Statistics live in a materialized view, so refresh is a step you own
The statistics behind BM25 are not computed on demand. Creating them with `bm25_create` creates a materialized view to record the stats, which means they are a snapshot of the document set at the moment you ran the statement.
That makes the lifecycle three functions rather than one. `bm25_create` builds the statistics. `bm25_refresh` updates them to reflect any changes in the underlying data. `bm25_drop` deletes them for a given table and column.
So an ingestion pipeline has a step people forget with lexical scoring: after loading or updating documents, the statistics are stale until you refresh them. A search over stale BM25 statistics still returns results, and they will look plausible, which is worse than an error. If your corpus updates incrementally, the refresh belongs in the same job as the load, not in a separate manual step.
One detail worth noting: the tokenize function is exposed for inspection, with `SELECT tokenize('i have an apple')` returning the token set. If your vocabulary looks wrong in the results, that call is the fastest way to see whether the problem is in the statistics or in how the text is being split.
The tokenizer is fixed to bert-base-uncased, and the code carries others
The last bullet in the How does it work section names the tokenizer and then flags its own limit. Currently it uses a Hugging Face tokenizer with the `bert-base-uncased` vocabulary set to tokenize words, and more tokenizer configuration might be supported in the future.
Read literally, that is an English uncased vocabulary with that model's token boundaries, chosen for you. For an English corpus that is a reasonable default. For a corpus with heavy domain vocabulary, or for any language whose morphology does not survive BERT's wordpiece splits, it is a constraint you did not choose.
The dependency list is more encouraging and more confusing at the same time. `Cargo.toml` pulls in `tokenizers` with its http and onig features, `tiktoken-rs`, and both `jieba-rs` and `tiniestsegmenter`, which are Chinese and Japanese word segmenters respectively. So the code carries several tokenizers while the page documents one.
That gap is worth resolving before you commit: the machinery for other tokenizers appears to exist in the build, but nothing on the page says how to select one. Read `tokenizer/` and `src/` for that rather than assuming the documented vocabulary is the only option.
Building needs cargo-pgrx at an alpha version and a matching pgrx pin
The build section starts from a plain requirement list: PostgreSQL, Rust and Cargo on your system. Then three steps, and the first pins an alpha.
cargo install cargo-pgrx --version v0.12.0-alpha.1cargo pgrx init --pg16=$(which pg_config) # assuming that you have PostgreSQL 16 installedcargo pgrx install --release # if you want to install it on your machine
cargo pgrx package # if you want to package `pg_bestmatch`The `--pg16` argument in the init step is a template for whichever major version you have. The manifest supports a range rather than a single one: features named `pg12` through `pg17` each enable the matching `pgrx` feature, and `pgrx` itself is pinned exactly at `=0.12.7` with `default-features = false`.
Two details sit oddly together. The build instructions name a cargo-pgrx alpha while the library dependency is an exact pin on a 0.12.7 release, so the CLI and the library are versioned separately and you have to match them yourself. And `Cargo.toml` still says `version = "0.0.0"` while the newest published tag is v0.0.1, so the manifest does not carry the release version.
The recall query uses the pgvecto.rs operator, and 0.77 is your correctness check
The workflow ends with a top-1 recall computation, and it doubles as the acceptance test:
SELECT sum((array[answer_pids] = array(SELECT pid FROM documents WHERE queries.dataset = documents.dataset ORDER BY queries.embedding <#> documents.embedding LIMIT 1))::int) FROM queries;The `<#>` operator is pgvecto.rs's, the `LIMIT 1` inside a correlated subquery gives one document per query, and the outer sum compares against the expected answer ids. The page then states the number: top-1 recall of BM25 on this dataset is 0.77, and if you reproduce that result your operations are correct.
That is the most useful sentence on the page for anyone evaluating the extension, because it turns a vague claim into a check you can fail.
One gap to be aware of. The index step shows both extensions, with `vectors` and `svector_dot_ops` for pgvecto.rs and `ivfflat` with `sparsevec_ip_ops` for pgvector, but the recall query only shows the pgvecto.rs operator. If you are on pgvector, the operator in that query has to change and the page does not say to what.
The dataset loader is two `wget` commands pulling the LoCoV1 Documents and Queries parquet files from Hugging Face, plus a psycopg2 block that registers adapters for numpy float64, int64, float32, int32 and numpy arrays. Those adapters exist because psycopg2 does not know numpy types natively, and they are easy to omit until a query returns an unhelpful type error.
pg_search puts the engine outside Postgres; this one stays inside
The page names the alternative directly and describes the trade without puffing either side.
`pg_bestmatch.rs` only provides methods for generating sparse vectors and does not support index-based search, which is left to pgvecto.rs or pgvector. `pg_search` performs BM25 retrieval through the external `tantivy` engine, which may have limitations when combined with transactions, filters or JOIN operations. And because `pg_bestmatch.rs` is entirely native to Postgres, it offers full compatibility with those operations inside Postgres.
That is the whole argument, and it is about where the work happens rather than about retrieval quality. A tantivy-backed extension has an engine outside the database, so a transaction that inserts and then searches, a filtered query, or a join has to cross a boundary. A native extension has no boundary: the same connection, the same transaction, the same planner-visible operations.
The price is that you now own two moving parts instead of one. You pick a vector extension, you pick an index type that is not the one you would otherwise use, and you maintain the statistics refresh. If you already run pgvecto.rs or pgvector, that second dependency is not a new thing. If you do not, adopting this to get BM25 means adding a vector extension as well.
Editorial conclusion
pg_bestmatch.rs fits a team that wants BM25 lexical retrieval to live inside the same PostgreSQL transaction as the rest of its RAG pipeline, and that already runs pgvecto.rs or pgvector and is willing to index with ivfflat rather than HNSW. It does not fit a team that wants a self-contained search extension, since it generates vectors and nothing else. Verify four things before building on it: that adding `bm_catalog` to your search path is actually in your deployment scripts, since the two-statement install will otherwise fail at the first function call; which tokenizer you need, because the documented one is bert-base-uncased; that `bm25_refresh` runs after every ingestion; and that you have a recall baseline to compare against, since the page gives 0.77 top-1 on the LoCo benchmark as its own correctness check. Licence is Apache-2.0, the last push was on 2024-11-05 and the newest tag is v0.0.1.
Frequently asked questions
Why does pg_bestmatch.rs not work well with HNSW indexing?
The page states that HNSW does not support the sparse vectors generated by BM25 very well, because the high sparsity prevents effective navigation within the graph. Its optional index example uses ivfflat for pgvector instead, and the extension does no indexing of its own.
What does installing pg_bestmatch.rs into PostgreSQL require?
Two statements: CREATE EXTENSION pg_bestmatch, then SET search_path TO public, bm_catalog so the extension's statistics catalog is on the path. Building from source needs PostgreSQL, Rust and Cargo, plus cargo-pgrx installed at v0.12.0-alpha.1.
How does pg_bestmatch.rs choose between svector and sparsevec?
By a third argument. bm25_document_to_svector with two arguments emits an svector for pgvecto.rs, while passing the string 'pgvector' emits a sparsevec. bm25_query_to_svector follows the same pattern.
What top-1 recall should pg_bestmatch.rs produce on the LoCo benchmark?
0.77, and the page says that reproducing that result means your operations are correct. The number is computed with the recall query shown in the usage section.
Which tokenizer does pg_bestmatch.rs use?
Currently a Hugging Face tokenizer with the bert-base-uncased vocabulary, and the page says more tokenizer configuration might be supported later. The manifest also depends on tiktoken-rs, jieba-rs and tiniestsegmenter, so other tokenizers are present in the build.
How does pg_bestmatch.rs differ from pg_search?
pg_bestmatch.rs only generates sparse vectors and leaves index-based search to pgvecto.rs or pgvector, while pg_search runs BM25 through the external tantivy engine, which may have limitations with transactions, filters or JOIN operations. Being native to Postgres gives full compatibility with those.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/tensorchord-pg-bestmatch-rs)