ParadeDB and pg_search: full-text, vector and aggregate search inside Postgres
One Postgres for your application data, full-text search, vector retrieval, and aggregations. Home of the pgsearch extension.
At a glance
- What is it?
- ParadeDB is a Postgres distribution plus the pg_search extension, built on pgrx, Tantivy and Apache DataFusion. It removes the second search system, and it asks you to accept an AGPL-3.0 extension and a young vector feature in return.
- Who is it for?
- Adopt ParadeDB when your search workload already lives next to relational data and you want BM25 ranking, filters and aggregates in the same transaction without running a separate index cluster. Stay away if you need a mature vector index today, or if AGPL-3.0 obligations are incompatible with how you ship your product.
- Can I use it commercially?
- Yes, with strict conditions. AGPL-3.0 is a network copyleft licence: if people use a modified version over a network, for example as a hosted service, you must offer them its source code under the same licence.
- Is it still maintained?
- Yes. The repository last received commits 4 days ago.
- What is it written in?
- Mainly Rust, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 25, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem ParadeDB removes: a second search system next to Postgres
Most applications that need ranked text search end up running two stores. Postgres holds the rows; a separate engine holds the inverted index; a sync pipeline keeps them consistent. The README frames the pitch directly: "Search without a second system." ParadeDB is a Postgres distribution that ships a custom index for full-text and vector search, BM25 scoring, filtering and aggregations, so application data and search engine sit in one database with, in the project's words, "no second system to deploy and nothing to sync."
The audience is narrow but real. If your queries look like "find documents matching these words, rank them, filter by tenant and date, then count by category," you currently pay for that with a sync job and a second set of credentials. ParadeDB targets the team that would rather write one SQL statement. It is not a general analytics warehouse replacement, and the README does not position it as one.
How pg_search works: pgrx, Tantivy and DataFusion inside the Postgres process
The mechanism is stated in the README's dependency list. pgrx bridges Postgres and Rust, Tantivy powers full-text and vector search, and Apache DataFusion handles OLAP processing. The README says ParadeDB "integrates battle-tested Rust libraries for search and analytics inside Postgres, contributing upstream whenever possible." That is the whole architecture in one sentence: the search engine runs in-process as an extension, not as a sidecar.
The repository layout matches that description. The workspace in Cargo.toml has members benchmarks, dst, macros, pg_search, stressgres, tests and tokenizers, with default-members set to pg_search and tokenizers. So the extension crate and the tokenizer crate are the product; benchmarks, dst (deterministic simulation testing) and stressgres are internal verification harnesses. Columnar storage and aggregates are documented as separate index types from the full-text path, which means the DataFusion side is not the same code path as BM25 scoring.
One detail worth reading carefully: the README's feature list has Vector Search unchecked while Full-Text Search, Hybrid Search, Filtering, Aggregates and JOINs are checked. The docs link exists, but the checkbox is the project's own signal about maturity.
Installing ParadeDB in Docker and running a first BM25 query
The README gives one install command for a fresh Docker container that drops you straight into a psql session. Run it and you should land at a Postgres prompt connected to the ParadeDB container.
curl -fsSL https://paradedb.com/install.sh | shFor deployment, the README points to the hosting options page rather than describing a production topology. If you are building the extension from source instead of using the container, the Makefile is the entry point. It reads the pgrx version out of Cargo.toml and installs exactly that, then initializes pgrx against the PostgreSQL version reported by pg_config.
make install-pgrx
make pgrx-initInstalling into the cluster that pg_config points at is a separate target, and it builds in release mode. That target runs cargo pgrx install --package pg_search --release --pg-config "$(PG_CONFIG)".
make installThe package target is the default goal and produces a packaged extension rather than an installed one. Note that a bare cargo build at the repository root covers only pg_search and tokenizers; the Cargo.toml comment warns that other members pull in dependencies the extension does not need, and that benchmarks builds the bundled DuckDB from C++ source.
Where ParadeDB is the wrong tool
Three limits are visible without running anything. First, vector search is unchecked in the README feature list. If your product is semantic retrieval over embeddings, the project's own documentation status tells you the vector path is the least settled part of the stack, and the docs page is where you would confirm what is actually supported.
Second, the integration surface is broad but shallow in places. Drizzle, Django, SQLAlchemy, Rails and EF Core each have their own repository, and the README lists Prisma as "coming." A broad list of thin adapters is a different thing from one deep ORM integration, and the README does not claim otherwise.
Third, this is an extension, not a fork you can ignore. It installs into a specific PostgreSQL version determined by pg_config, and the Makefile's install-pgrx target pins the pgrx version to whatever Cargo.toml declares. That coupling is fine on a managed platform that already offers ParadeDB, and awkward if you run an unusual Postgres build or depend on extensions that conflict. Nothing in the README describes a compatibility matrix for third-party extensions.
ParadeDB versus Elasticsearch, pgvector and DuckDB
The comparison people actually search for is ParadeDB versus Elasticsearch. The difference is architectural, not a benchmark result: Elasticsearch is a separate service with its own cluster, shards and client, and you synchronize data into it. ParadeDB keeps the index inside Postgres, so a query can join search results against ordinary tables in one statement. The trade is that you inherit Postgres scaling behaviour instead of a system designed from the start for distributed search. The README makes no performance claim against Elasticsearch, and neither should you assume one.
Against pgvector, the split is what each extension is for. pgvector adds a vector type and index to Postgres; ParadeDB's README lists full-text with BM25 scoring, top K, highlighting, tokenizers and token filters, hybrid search, filtering, aggregates with columnar storage and facets, and JOINs. Vector retrieval appears as one unchecked item. If you only need nearest-neighbour lookup, pgvector is the narrower dependency. If you need ranked text search and aggregations, ParadeDB covers ground pgvector does not.
Against DuckDB, the README's own Cargo.toml is the clearest statement: the workspace comment describes DuckDB as a "benchmark comparison engine," and the profile section notes its objects with -g cost gigabytes per target directory. DuckDB is an embedded analytical engine for files and local analysis; ParadeDB is an extension that runs inside a Postgres server serving concurrent application traffic. Same analytical ambition, different deployment shape.
Licence and upgrade cost: AGPL-3.0 and a pinned pgrx version
ParadeDB Community is licensed under the GNU Affero General Public License v3.0, and the workspace Cargo.toml declares license = "AGPL-3.0". The README separates that from ParadeDB Enterprise, pointing to the enterprise deployment docs and a sales address for commercial licensing. Whether AGPL-3.0 obligations reach your product depends on how you distribute or expose the software, and that is a question for your own counsel, not something this article can settle.
Upgrade cost is tied to the pgrx pin. The Makefile derives PGRXV from Cargo.toml and installs that exact cargo-pgrx version, then initializes it against the PostgreSQL version from pg_config. Moving to a new PostgreSQL major version therefore means a new pgrx init and a rebuild, not just a package upgrade. Release cadence looks tight: v0.25.6 landed on 2026-08-27, the same day as the rc.3 and rc.1 candidates, and the last push to the default branch was 2026-08-27. Frequent releases are good for fixes and bad for anyone who wants to pin and forget.
Who should adopt ParadeDB, and what to check first
The fit is a team already on Postgres whose search requirements are text ranking, filters and aggregates, and who is willing to run an extension rather than a separate cluster. The mismatch is a team that needs a proven vector index now, or one whose product cannot absorb AGPL-3.0 terms.
Before adopting, confirm the install path against your actual Postgres version by running the Makefile targets that print the values it will use.
make pg-version
make pgrx-versionThen verify that your queries hit the custom index and not a sequential scan, since the README describes the index but does not promise a planner outcome. Finally, open the vector and hybrid search docs pages and check the current state of the unchecked item against your recall requirements.
Editorial conclusion
Adopt ParadeDB when your search workload already lives next to relational data and you want BM25 ranking, filters and aggregates in the same transaction without running a separate index cluster. Stay away if you need a mature vector index today, or if AGPL-3.0 obligations are incompatible with how you ship your product. Before committing, verify three things on your own hardware: that pg_search installs cleanly against your exact PostgreSQL version through the Makefile targets, that your queries actually use the custom index rather than falling back to sequential scans, and that the vector path meets your recall needs given the unchecked box in the README.
Frequently asked questions
Which is better, ParadeDB or Elasticsearch?
They solve the same problem with different architectures. Elasticsearch is a separate service you synchronize data into; ParadeDB puts the search index inside Postgres so search results can join against ordinary tables in one query. The README makes no performance comparison between the two.
What is ParadeDB?
It is a Postgres distribution that adds a custom index for full-text and vector search, BM25 scoring, filtering and aggregations, so application data and search live in one database. The README describes it as "One Postgres for your application data, full-text search, vector retrieval, and aggregations."
Is ParadeDB open source?
Yes. ParadeDB Community is licensed under the GNU Affero General Public License v3.0, and the workspace Cargo.toml declares the same licence. A separate ParadeDB Enterprise offering exists, with licensing handled through the sales contact in the README.
Is ParadeDB free?
ParadeDB Community is available under AGPL-3.0 at no stated cost. The README directs commercial and enterprise licensing enquiries to [email protected], so paid options exist alongside the community edition.
How does ParadeDB compare with pgvector?
pgvector adds vector storage and search to Postgres. ParadeDB's README lists full-text search with BM25 scoring, top K, highlighting, tokenizers, hybrid search, filtering, aggregates and JOINs as checked features, while Vector Search is unchecked. The overlap is narrower than the names suggest.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/paradedb-paradedb)