Open-source project
timescale/pgvectorscale avatar
timescale/pgvectorscale

pgvectorscale: a StreamingDiskANN index for pgvector, and what its release history tells you

Postgres extension for vector search (DiskANN), complements pgvector for performance and scale. Postgres OSS licensed.

3,136 stars157 forksRustPostgreSQL

At a glance

What is it?
A PostgreSQL extension written in Rust that adds Microsoft's DiskANN algorithm and statistical binary quantization on top of pgvector, with a 2026 security release that matters if you are on an earlier version.
Who is it for?
pgvectorscale is worth evaluating if you have real embedding volume inside an existing PostgreSQL deployment and would rather not operate a separate vector database, because the operational saving is the whole argument.
Can I use it commercially?
Yes. PostgreSQL is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 6 days ago.
What is it written in?
Mainly Rust, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 7, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What it adds on top of pgvector

pgvectorscale is an extension, not a database. It installs into a PostgreSQL server you already run and works on the data pgvector already stores, which is the entire premise. The README describes it as complementing pgvector rather than replacing it, and in practice you create the pgvectorscale extension alongside the vector type from pgvector and never write the extension name in a query again.

Three additions are named. The first is a new index type called StreamingDiskANN, inspired by Microsoft's DiskANN algorithm. The second is statistical binary quantization, described as developed by Timescale researchers and as improving on standard binary quantization, which is a compression technique that shrinks stored vectors so more of them fit in memory. The third is label based filtered vector search, based on Microsoft's filtered DiskANN research, which lets you combine a similarity search with a label filter rather than filtering after retrieval.

The filtered search is the one that changes how you write application queries. Post-filtering a similarity search means fetching the nearest neighbours and then discarding the ones that fail a metadata condition, which can leave you with fewer results than you asked for. Combining the filter inside the index traversal is what fixes that, and it is a real difference in kind rather than a constant factor.

The project is written in Rust using the pgrx framework, which the README justifies by saying it offers the PostgreSQL community a new avenue for contributing to vector support. pgvector itself is C.

Installing the extension with CREATE EXTENSION

Once the extension is available on your server, the install is one statement:

sql
CREATE EXTENSION IF NOT EXISTS vectorscale CASCADE;

The `CASCADE` is doing real work. The README states plainly that it automatically installs pgvector, so a single statement gives you both extensions in the right order.

The README offers three routes to get there. A pre-built TimescaleDB Docker image is the fastest, and a Timescale Cloud service can have the extension enabled for you. The third is building from source, which is the route that reveals the version constraints.

bash
# install rust
curl --proto '=https' --tlsv1.2 -sSf https://sh.rustup.rs | sh

# download pgvectorscale
cd /tmp
git clone --branch <version> https://github.com/timescale/pgvectorscale
cd pgvectorscale/pgvectorscale
# install cargo-pgrx with the same version as pgrx
cargo install --locked cargo-pgrx --version $(cargo metadata --format-version 1 | jq -r '.packages[] | select(.name == "pgrx") | .version')
cargo pgrx init --pg18 pg_config
# build and install pgvectorscale
cargo pgrx install --release

The `cargo pgrx init --pg18` line pins the build to PostgreSQL 18, and the `cargo metadata` invocation before it exists so that the cargo-pgrx version matches the pgrx version the crate depends on. Skipping that step is a common way to end up with a confusing build failure.

There is one platform restriction stated plainly: building on macOS on Intel is not supported, with an open issue tracking it. The listed alternatives are an ARM based Mac, building on Linux, or the pre-built containers.

Creating the table and the DiskANN index

The working example uses the `VECTOR` column type from pgvector, which tells you the two extensions share the same storage format rather than requiring a translation layer.

postgresql
CREATE TABLE IF NOT EXISTS document_embedding  (
    id BIGINT PRIMARY KEY GENERATED BY DEFAULT AS IDENTITY,
    metadata JSONB,
    contents TEXT,
    embedding VECTOR(1536)
)

A JSONB column alongside the vector is the pattern to note. With label based filtered search, that metadata column is what you filter on, so deciding how you will partition your documents is a schema decision made before the index exists.

The index creation names the index access method as `diskann` and the operator class as `vector_cosine_ops`:

postgresql
CREATE INDEX document_embedding_idx ON document_embedding
USING diskann (embedding vector_cosine_ops);

`vector_cosine_ops` is a pgvector operator class name, which is the clearest evidence that the index works directly on pgvector's stored representation. The 1536 dimensions in the example match a common embedding size for a particular generation of text embedding models.

The README's own note on populating the table defers to the pgvector documentation rather than repeating it, which is a reasonable division: how to get vectors into a column is pgvector's question, and how to index them once they are there is pgvectorscale's.

What version 0.9.1 fixed, and why it should change your upgrade decision

The most useful thing in this repository's release notes is the 0.9.1 entry, published on 2026-09-04. It is a security hardening release, and the description is specific enough to be useful rather than vague.

The stated problem: the operator class SQL referenced pgvector's `vector` type and operators unqualified, so a type pre-created in the extension's target schema could capture the binding. The native code trusted the typmod, the persisted metapage dimensions and the datum layout without validation. The release notes state plainly that together these could let attacker-controlled bytes reach native DiskANN code paths and cause a backend crash, memory disclosure or an out of bounds write.

That is a serious class of problem, because the attack requires someone to create a type in your schema, which is the kind of thing that happens in a multi tenant database where tenants create their own objects.

The fixes are equally concrete. Operator classes now bind explicitly to pgvector's extension schema with quoted, schema-qualified type and operator references, and the extension script's search_path is preserved. Vector type, signed typmod, persisted metapage dimensions, and detoasted datum size and dimension are all validated before any slice is constructed. `amvalidate` is implemented for the vector DiskANN operator classes. And direct upgrade scripts to 0.9.1 ship from every prior release.

The upgrade path is stated as `ALTER EXTENSION vectorscale UPDATE TO '0.9.1'`, with verification that existing vector operator classes are bound to pgvector correctly. Prior releases were 0.9.0 in November 2025 and 0.8.0 in July 2025, both a long way from this.

The project is PostgreSQL licensed, which matters for a PostgreSQL extension, and the repository carries both a `LICENSE` file and a `NOTICE` file.

The benchmark claim, and what it does and does not settle

The README leads with a specific performance claim: on a benchmark dataset of 50 million Cohere embeddings with 768 dimensions each, PostgreSQL running pgvector plus pgvectorscale achieves 28x lower p95 latency and 16x higher query throughput compared to Pinecone's storage optimized index for approximate nearest neighbour queries at 99% recall, at 75% less cost when self-hosted on AWS EC2.

Read those numbers for what they are. They are a comparison against a hosted service's index configuration, on a dataset of a particular dimensionality, at a particular recall target, on particular cloud infrastructure. A 768 dimension Cohere embedding is not a 1536 dimension OpenAI embedding, and recall targets are a choice you make rather than a fixed property. The measurement is also the vendor's own, and the README points to a Timescale blog post for methodology rather than publishing a harness you can rerun here.

What is fair to conclude is that the extension exists because pgvector's default index types stop scaling well at this volume, and that DiskANN with quantization is a well understood answer to that problem, borrowed from research Microsoft published rather than invented here. Whether the specific ratios transfer to your data is something you have to determine on your own data.

The repository is also honest about its own build maturity. The Makefile at the root carries a literal question in a comment about erroring out if the target is not PostgreSQL 14, pins `PGRX_VERSION=0.9.8`, and derives the PostgreSQL version by parsing `pg_config --version` with awk. A regression test harness is configured with roles named `superuser`, `tsdbadmin` and `test_role_1`, and a locale is chosen conditionally on Darwin. That is a working test setup rather than a polished one, and it is a fair signal about what to expect from a project whose last release was a security fix.

Editorial conclusion

pgvectorscale is worth evaluating if you have real embedding volume inside an existing PostgreSQL deployment and would rather not operate a separate vector database, because the operational saving is the whole argument. Two things to check before you rely on it: whether your Postgres version is covered by the pgrx build path the README assumes, since the repository's own Makefile still carries a question mark about version 14 support, and whether you are on 0.9.1 or later, because 0.9.1 fixed validation gaps in native DiskANN code that the release notes describe as capable of causing a backend crash or memory disclosure. The last push was on 2026-09-07 and version 0.9.1 shipped on 2026-09-04, so the project is being worked on.

Frequently asked questions

How do I install pgvectorscale in PostgreSQL?

Enable it in your database with CREATE EXTENSION IF NOT EXISTS vectorscale CASCADE, where CASCADE also installs pgvector. To have the extension files on the server, use a pre-built TimescaleDB container, a Timescale Cloud service, or build from source with cargo pgrx install --release.

What index type does pgvectorscale add to pgvector?

A new index type called StreamingDiskANN, created with USING diskann in a CREATE INDEX statement. It is inspired by Microsoft's DiskANN algorithm, and the index operates on pgvector's stored vectors directly, which is why the operator class is named vector_cosine_ops.

What changed in pgvectorscale 0.9.1?

It is a security hardening release. The operator class SQL bound pgvector's vector type and operators unqualified, and native code trusted typmod and datum layout without validation, which together could let attacker-controlled bytes reach native DiskANN code paths and cause a crash, memory disclosure or an out of bounds write. Upgrade with ALTER EXTENSION vectorscale UPDATE TO '0.9.1'.

Is pgvectorscale faster than Pinecone?

The README reports 28x lower p95 latency and 16x higher throughput against Pinecone's storage optimized index on 50 million 768 dimension Cohere embeddings at 99% recall, self-hosted on AWS EC2. Those are the project's own figures for a specific dataset and recall target, with methodology described in a linked Timescale blog post rather than a harness in the repository.

Can I build pgvectorscale on macOS?

Not on Intel Macs: the README states that building on macOS X86 is currently unsupported and links an open issue. The alternatives it lists are an ARM based Mac, building on Linux, or using the pre-built Docker containers.

Official sources

  1. Issues
  2. License: PostgreSQL
  3. README
  4. Releases
  5. timescale/pgvectorscale on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/timescale-pgvectorscale.svg)](https://hysenlabs.com/projects/timescale-pgvectorscale)