pg_textsearch: BM25 Ranking Inside Postgres
PostgreSQL extension for BM25 relevance-ranked full-text search. Postgres OSS licensed.
At a glance
- What is it?
- pg_textsearch is a PostgreSQL extension from Timescale that adds a bm25 index access method and a `<@>` scoring operator, so relevance ranking happens in the database instead of in a separate search service. It supports PostgreSQL 17 and 18, and ships as source or prebuilt Linux binaries.
- Who is it for?
- pg_textsearch is worth adopting when your corpus already lives in PostgreSQL 17 or 18, your ranking needs are BM25 with tunable k1 and b, and you want ranking and boolean filtering in the same query planner. Skip it if you need cross-cluster search, faceting, or aggregations, or if you run PostgreSQL 16 or older, since the README states 17 and 18 only.
- Can I use it commercially?
- Yes. PostgreSQL is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 5 days ago.
- What is it written in?
- Mainly C, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 26, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The gap pg_textsearch fills between tsvector and Elasticsearch
PostgreSQL has shipped full-text search for years, but `ts_rank` is not BM25. Teams that want relevance scoring that behaves like Lucene or Elasticsearch usually end up running a second system, syncing documents out of Postgres, and accepting the consistency and operational cost that comes with it. pg_textsearch is aimed at the middle of that range: you keep the text in a normal table, add a bm25 index, and rank with an operator.
The target user is an engineer who already has Postgres 17 or 18 and a text column, and whose ranking requirements stop short of needing a dedicated search cluster. The README describes the query surface as `ORDER BY content <@> 'search terms'`, with BM25 parameters `k1` and `b` configurable per index. That is a deliberately small API. There is no separate query DSL to learn and no document ingestion pipeline, because the table is the index source.
It is not a general search platform. There are no facets, no aggregations, no cross-index federation, and no suggestion API in the documented surface. If those are requirements, this extension is the wrong layer.
How the bm25 access method and the <@> operator work
The extension registers a new index access method named `bm25`, so an index is declared with `CREATE INDEX ... USING bm25(content)`. The index is built over a text column or an immutable text expression, and it is configured with a PostgreSQL text search configuration through the required `text_config` option.
Queries use the `<@>` operator. The README states that `<@>` returns negative BM25 scores so that ascending index scans put the best matches first, which is why the idiomatic form is `ORDER BY content <@> 'query' LIMIT n`. The index is detected from the column, and `to_bm25query('query', 'index_name')` names it explicitly when detection is ambiguous or when the index is partial.
Two details in the architecture are worth flagging. First, scoring can happen outside an index scan: the README describes standalone scoring, which still uses corpus statistics from the selected BM25 index but does not require an ordered index scan. Second, boolean filtering is a separate scan mode. The `@@` operator with `tsquery` filters through a BM25 index, and the README notes that a query combining `WHERE content @@ ...` with `ORDER BY content <@> ...` cannot use one BM25 index scan for both operations. Phrase and weight checks may be rechecked against the table row after the index finds candidates.
The repository layout reflects a storage engine rather than a thin wrapper: `src/memtable/` holds posting, page, stringtable and log files, and `src/segment/` holds dictionary, merge and scan files. The README also lists fast top-k queries with Block-Max WAND and parallel index builds for large tables.
Installing pg_textsearch and running a first ranked query
The README gives two install paths: prebuilt binaries from the Releases page, available for Linux on amd64 and arm64, or a source build. The source build is three commands.
cd /tmp
git clone https://github.com/timescale/pg_textsearch
cd pg_textsearch
make
make install # may need sudoAfter installing, the library must be preloaded. The README says to add it to `shared_preload_libraries` in `postgresql.conf` and restart the server. If that setting already lists other libraries, append rather than replace.
shared_preload_libraries = 'pg_textsearch' # add to existing list if neededThen enable the extension in each database you want to use it in.
CREATE EXTENSION pg_textsearch;A first real use starts with a table and an index. The index requires `text_config`, and the README's example sets it to `english`.
CREATE TABLE documents (
id bigserial PRIMARY KEY,
content text,
category_id integer
);
CREATE INDEX docs_idx ON documents USING bm25(content) WITH (text_config='english');The ranked query is a single `ORDER BY` clause. Lower scores rank first because `<@>` returns negative BM25 values.
SELECT * FROM documents
ORDER BY content <@> 'database system'
LIMIT 5;To confirm the index is actually being used, run `EXPLAIN` on the same query. The README warns that PostgreSQL may prefer standalone scoring with a sequential scan on small tables, and suggests `SET enable_seqscan = off;` to test the index plan. On a small demo table, seeing a sequential scan is expected and not a bug.
Index options, expression indexes and the partial-index naming rule
The documented index options are `text_config` (required, no default), `k1` with a default of 1.2 and a range of 0.1 to 10.0, `b` with a default of 0.75 and a range of 0.0 to 1.0, plus `compaction` and `compaction_schedule`. Setting `k1` and `b` per index means you can tune term-frequency saturation and length normalization differently for a title column and a body column without changing the query.
Expression indexes extend this to JSONB and multi-column search. The expression must evaluate to `text` and use only IMMUTABLE functions, and the query must repeat the same expression in the `ORDER BY` clause. The README gives `(data->>'description')`, `(lower(content))` and a `coalesce(title, '') || ' ' || coalesce(body, '')` concatenation as examples. The IMMUTABLE requirement is the real constraint here: anything that depends on session settings or time cannot be indexed this way.
Partial indexes add a `WHERE` clause, and the README is explicit that partial indexes require explicit index naming through `to_bm25query()`. The implicit `text <@> 'query'` form is not enough, because the planner cannot infer which partial index applies. If you build a category-scoped index, every query against it must name it.
Pre-filtering, post-filtering and the 100,000-result scan cap
Filtering behavior depends on which plan PostgreSQL picks. When it chooses standalone scoring, a separate index can pre-filter rows before scoring, and the README shows a plain btree index on `category_id` doing that job alongside the BM25 ordering.
When PostgreSQL instead chooses the ordered BM25 index scan, other conditions become post-filters applied after BM25 scoring. The README shows `WHERE length(content) > 100` in that role. The mechanism for keeping results correct is that post-filtered scans grow their internal scoring batch until the `LIMIT` is filled, matches are exhausted, or a 100,000-result scan cap is reached.
That cap is the sharpest limitation documented. A post-filter that is selective relative to the corpus, combined with a large `LIMIT`, can exhaust the batch before the limit is satisfied. The README does not describe what happens at the cap beyond stating that it exists, and it does not document rollback or recovery behavior for the extension. If your query pattern depends on a highly selective post-filter returning many rows, test it against your own data rather than assuming the batch growth will cover it.
Boolean filtering, config matching and prepared-statement traps
Boolean filtering uses PostgreSQL's own `tsquery` syntax through `@@`, with `&`, `|`, `!`, phrase operators such as `<->`, prefix matching with `:*`, and weight restrictions all listed as supported. This is a real advantage over inventing a second query language: the operators are the ones Postgres users already know.
The constraint is configuration matching. The README states that the `default_text_search_config` used to parse the left-hand `text` value must match the index configuration, and shows `SET default_text_search_config = 'english';` as the fix. Mismatched configs are the kind of problem that surfaces as wrong results rather than an error.
There is also a prepared-statement caveat specific to boolean queries. If `default_text_search_config` changes after a boolean prepared statement has switched to a generic plan, the README says to `DEALLOCATE` and prepare the statement again. A newly planned query can choose the correct sequential fallback, while the cached plan is rejected to avoid incorrect index results. This is a correctness-first design, but it means a connection pool that caches prepared statements across a config change needs to be aware of the behavior.
Compaction modes and what they cost you
The `compaction` option controls spill-time compaction and takes `inline`, `background`, or `manual`, with `inline` as the default. The `compaction_schedule` option defaults to `pg_textsearch.background_compaction_schedule` and is described as an optional cron schedule captured when the index enters background mode.
This is a genuine operational trade-off rather than a cosmetic setting. Inline compaction keeps maintenance inside the writing path, which is predictable but puts the work in front of inserts. Background compaction moves it out of the way but introduces a schedule to configure and a mode transition to reason about. Manual compaction hands the decision to the operator. The README does not document the cost profile of each mode, so the choice has to be made from your own write patterns.
The repository includes a `benchmarks/` directory and the README links a benchmarks site, but the README itself does not publish numbers. Treat any performance expectation as something to measure on your own hardware and data shape.
pg_textsearch compared with pg_search and with a separate search engine
The closest comparison is with other in-database search options, and the difference is the ranking model. PostgreSQL's built-in full-text search gives you `tsvector`, `tsquery` and `ts_rank`; pg_textsearch gives you a `bm25` access method and BM25 scores with tunable `k1` and `b`. If your ranking complaints are about `ts_rank` producing unintuitive ordering, that is the specific problem this addresses.
Against a separate engine such as Elasticsearch or OpenSearch, the difference is architectural rather than algorithmic. Both compute BM25-style relevance, but the separate engine requires moving documents out of Postgres, keeping two stores consistent, and running a second service. pg_textsearch removes that sync path and keeps ranking inside the transaction boundary of the database you already operate. What you give up is everything a dedicated engine provides beyond ranking: distributed search across clusters, faceting, aggregations, and a query DSL built for those features. The README's documented surface is scoring and boolean filtering, not analytics.
The version constraint matters in this comparison too. pg_textsearch supports PostgreSQL 17 and 18, with PostgreSQL 19 beta supported on a best-effort basis, its CI allowed to fail and no prebuilt binaries published for it yet. If you are on PostgreSQL 16 or older, this extension is not an option and the comparison is moot.
Editorial conclusion
pg_textsearch is worth adopting when your corpus already lives in PostgreSQL 17 or 18, your ranking needs are BM25 with tunable k1 and b, and you want ranking and boolean filtering in the same query planner. Skip it if you need cross-cluster search, faceting, or aggregations, or if you run PostgreSQL 16 or older, since the README states 17 and 18 only. Before committing, verify three things on your own data: that EXPLAIN shows the bm25 index scan rather than standalone scoring with a sequential scan, that your default_text_search_config matches the index text_config, and that a combined WHERE content @@ ... ORDER BY content <@> ... query is acceptable given that it cannot use one BM25 index scan for both operations.
Frequently asked questions
What is pg_textsearch?
It is a PostgreSQL extension that adds a bm25 index access method and a `<@>` operator for BM25 relevance-ranked full-text search. It is licensed under the PostgreSQL licence, written in C, and supports PostgreSQL 17 and 18.
How do I install pg_textsearch?
Download prebuilt binaries for Linux on amd64 or arm64 from the Releases page, or clone the repository and run make followed by make install. After installing, add pg_textsearch to shared_preload_libraries in postgresql.conf, restart the server, and run CREATE EXTENSION pg_textsearch in each database.
Can pg_textsearch combine boolean filtering with BM25 ranking in one index scan?
No. The README states that boolean filtering and BM25 ranking are separate scan modes, and a query combining WHERE content @@ ... with ORDER BY content <@> ... cannot use one BM25 index scan for both operations.
Why does EXPLAIN show a sequential scan instead of using my pg_textsearch index?
The README notes that PostgreSQL may prefer standalone scoring with a sequential scan for small tables. It suggests setting enable_seqscan to off to test whether the index plan is chosen.
Which PostgreSQL versions does pg_textsearch support?
The README states that pg_textsearch supports PostgreSQL 17 and 18. PostgreSQL 19 beta is supported on a best-effort basis, with CI allowed to fail and no prebuilt binaries published for it yet.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/timescale-pg-textsearch)