# pgvector: Vector Similarity Search Inside PostgreSQL

> pgvector adds vector columns and L2, inner product, cosine, L1, Hamming and Jaccard distance operators to Postgres, with exact or approximate nearest neighbor search. It suits teams that already run Postgres and want embeddings next to relational data, and it is the wrong choice when you need a purpose-built vector service.

**pgvector/pgvector** — Open-source vector similarity search for Postgres

- Repository: https://github.com/pgvector/pgvector
- Stars: 23,193 · Forks: 1,334
- Language: C
- License: NOASSERTION
- Published: 2026-09-21 · Updated: 2026-09-21 · Language: en
- Canonical page: https://hysenlabs.com/projects/pgvector-pgvector

## What problem pgvector solves, and who it is for

Most vector stores are separate systems. You embed your documents, push the vectors into that system, and then keep a second copy of the metadata in Postgres so you can filter by user, tenant or date. pgvector removes that split. The README frames the pitch as "Store your vectors with the rest of your data", and the extension keeps the vectors in ordinary Postgres tables, which means the same transaction that writes a row also writes its embedding.

The audience is Postgres users who already have a database and want nearest neighbor search without adding a second datastore. Because the extension is queried through SQL, any language with a Postgres client can use it. The README lists ACID compliance, point-in-time recovery and JOINs as inherited properties rather than new features, which is the honest framing: pgvector does not reimplement durability or backup, it borrows them.

It is not aimed at people who want a managed vector API with no database administration. If you do not want to run Postgres, this is not the tool.

## How the vector type, operators and indexes fit together

The core is a `vector` column type with a fixed number of dimensions, declared as `vector(3)` in the README examples. Distance is expressed as operators in an ORDER BY clause rather than as a function call: `<->` for L2 distance, `<#>` for negative inner product, `<=>` for cosine distance, `<+>` for L1 distance, and `<~>` and `<%>` for Hamming and Jaccard distance on binary vectors.

One detail in the README is easy to miss and causes real bugs. The `<#>` operator returns the negative inner product, because Postgres only supports ASC order index scans on operators. So a query sorted by `<#>` gives you the right ordering but the wrong sign if you read the value directly. The README's own fix is to multiply by -1, and to compute cosine similarity as `1 - (embedding <=> '[3,1,2]')`.

By default the search is exact, and the README says that provides perfect recall. Adding an HNSW or IVFFlat index switches you to approximate search and trades recall for speed. The README is explicit that "you will see different results for queries after adding an approximate index", which is the single most important behavioural change in the project and the one most likely to surprise a team that adds an index after testing without one.

HNSW builds a multilayer graph. The README states it has better query performance than IVFFlat in speed-recall terms, but slower build times and higher memory use, and unlike IVFFlat it needs no training step, so an index can be created on an empty table. Each distance function needs its own index, and the operator class must match the column type: `vector_l2_ops` for `vector`, `halfvec_l2_ops` for `halfvec`, `sparsevec_l2_ops` for `sparsevec`, and `bit_hamming_ops` or `bit_jaccard_ops` for binary vectors.

## Installing pgvector and running a first nearest neighbor query

The README supports Postgres 13 and later. On Linux and Mac the source build is three commands after cloning the tagged release. The `make install` step may need sudo, because it copies files into the Postgres installation directory.

```bash
cd /tmp
git clone --branch v0.8.6 https://github.com/pgvector/pgvector.git
cd pgvector
make
make install # may need sudo
```

The README also lists Docker, Homebrew, PGXN, APT, Yum, pkg, APK and conda-forge as installation routes, and notes the extension comes preinstalled with Postgres.app and many hosted providers. The repository ships a Dockerfile that builds against `postgres:$PG_MAJOR-$DEBIAN_CODENAME` with `PG_MAJOR` defaulting to 17 and Debian codename bookworm, then removes the build toolchain from the final image. That is the path to take if you want a container rather than a host install.

Once the extension is available, enabling it is per database, not per server. That is a detail worth internalizing before you script anything.

```sql
CREATE EXTENSION vector;
```

Then create a table with a vector column, insert a couple of rows, and query by L2 distance.

```sql
CREATE TABLE items (id bigserial PRIMARY KEY, embedding vector(3));
INSERT INTO items (embedding) VALUES ('[1,2,3]'), ('[4,5,6]');
SELECT * FROM items ORDER BY embedding <-> '[3,1,2]' LIMIT 5;
```

You should get both rows back, ordered by distance from the query vector. At this point the search is exact. If you later add an HNSW index for L2 distance, the query text does not change, but the results can.

```sql
CREATE INDEX ON items USING hnsw (embedding vector_l2_ops);
```

The two HNSW parameters are `m`, the max connections per layer, defaulting to 16, and `ef_construction`, the candidate list size during graph construction, defaulting to 64. The README states a higher `ef_construction` gives better recall at the cost of index build time and insert speed.

## Dimension limits and the quantization trade-off

The index type you choose caps your dimensions. HNSW supports `vector` up to 2,000 dimensions, `halfvec` up to 4,000, `bit` up to 64,000, and `sparsevec` up to 1,000 non-zero elements. If your embedding model outputs more than 2,000 dimensions, you either reduce the dimension, move to half precision, or accept an exact scan with no index. That is a hard boundary, not a tuning knob.

The README points to quantization under its scaling section as the answer for large vector counts. Half-precision, binary and sparse vectors are the mechanisms, and each has its own operator class. The trade-off is not spelled out in the README: quantizing to binary vectors changes the distance function to Hamming or Jaccard, which is a different notion of similarity from cosine distance on the original floats. Teams that switch types to save memory should expect to re-evaluate recall rather than assume the ranking is preserved.

The build itself has a portability wrinkle. The Makefile defaults `OPTFLAGS` to `-march=native` for auto-vectorization, but disables it on Mac ARM, PowerPC and RISC-V64. The Dockerfile passes `make OPTFLAGS=""` explicitly. If you build on one machine and deploy to another with a different CPU, the default flags are the reason a binary can fail; the repository's own container avoids this by turning them off.

## Where pgvector is the wrong tool

The clearest failure mode is operational, not algorithmic. pgvector is a Postgres extension. It does not exist without a Postgres server, and installing it means compiling C against the server's headers or finding a package for your platform. On Windows the README requires C++ support in Visual Studio and an `x64 Native Tools Command Prompt for VS` run as administrator, then `nmake /F Makefile.win` and `nmake /F Makefile.win install`. If your environment cannot build extensions and your host does not ship one, you are stuck.

The second limitation is recall. Exact search gives perfect recall but scans, and the README's own wording is that an approximate index trades recall for speed. There is no documented recall target, no built-in recall measurement, and no rollback procedure described in the README for removing an index that degraded quality. You can drop the index, but the README does not present that as a supported operational workflow.

The third is scope. pgvector stores and searches vectors. It does not generate embeddings, does not chunk documents, and does not manage a pipeline. If you expected a retrieval framework, this is one layer below it. And if your workload is mostly vector search at a scale where Postgres itself becomes the bottleneck, a dedicated vector system may fit better, at the cost of running two datastores.

## pgvector compared with Qdrant, ChromaDB and OpenSearch

The comparison people search for most is pgvector against Qdrant. The difference is architectural. Qdrant is a standalone vector database: you run a separate service, and your application talks to it over its own API. pgvector is an extension inside Postgres, queried with SQL, and it inherits transactions, point-in-time recovery and JOINs from the database. With pgvector you can filter by a relational column and search by distance in one statement. With a separate service you either duplicate metadata into the vector store or fetch candidate IDs and join them back in application code.

ChromaDB is closer to an embedded library for prototyping, with its own persistence model. The distinction is the same: pgvector has no separate process, so there is nothing to deploy beyond the extension, but there is also no separate process to scale independently of your database.

OpenSearch is a search engine first, with vector search added to an existing indexing and aggregation stack. If you already run OpenSearch for text search, adding vectors there avoids a second system too, but you inherit its cluster management and its own query language. The pgvector answer is narrower: SQL operators, two index types, and whatever Postgres already gives you. For a team already on Postgres, that narrowness is the point. For a team that needs independent scaling of the vector tier, it is the limitation.

## Maintenance, versions and licence

The repository's last push was on 2026-09-10 and it is not archived, so it is being worked on. The README pins the current release as v0.8.6, and the Makefile sets `EXTVERSION = 0.8.6`, so the extension version and the source tag line up. Upgrades go through the SQL migration files in `sql/`, which the Makefile picks up with a wildcard over `sql/*--*--*.sql` and ships as `DATA`, while the base script is generated by copying `sql/vector.sql` to `sql/vector--0.8.6.sql`. In practice that means an upgrade is a rebuild and reinstall followed by `ALTER EXTENSION vector UPDATE`, and you should read `CHANGELOG.md` before doing it, since the README does not document a downgrade path.

The licence is the open question. The repository metadata reports `NOASSERTION` rather than a recognized SPDX identifier, and the README does not state a licence in the text. The Dockerfile copies `LICENSE` into `/usr/share/doc/pgvector`, so the file exists at the repository root, but the metadata does not classify it. If you need a specific licence for compliance, read the `LICENSE` file itself rather than trusting the repository metadata. This is not legal advice, and it is worth a look before you ship it inside a product.

The practical maintenance cost is low if you already run Postgres: one extension to keep in step with your server version. It rises if you build from source on every platform you deploy to, because the `-march=native` default in the Makefile is not portable across CPUs.

## Conclusion

Adopt pgvector if you already run Postgres 13 or later and want embeddings stored beside relational data, with ACID compliance, point-in-time recovery and JOINs from the same database. Do not adopt it if you need a dedicated vector service with its own scaling model, or if you cannot compile an extension against your server. Before committing, verify your Postgres version is 13 or higher, check which installation path your host supports, and measure recall after adding an HNSW or IVFFlat index, since the README states that query results change once an approximate index exists.

## FAQ

### What is pgvector?

pgvector is an open-source extension that adds vector similarity search to Postgres. It provides a vector column type and distance operators for L2, inner product, cosine, L1, Hamming and Jaccard distance, with exact or approximate nearest neighbor search.

### What is the difference between PostgreSQL and pgvector?

PostgreSQL is the database; pgvector is an extension installed into it. The README describes pgvector as storing vectors with the rest of your data, so the vectors live in ordinary Postgres tables and inherit ACID compliance, point-in-time recovery and JOINs.

### Is pgvector any good?

The README states that exact search provides perfect recall, and that adding an HNSW or IVFFlat index trades some recall for speed. Whether that trade-off is acceptable depends on your workload, and the README does not publish recall numbers, so you should measure it on your own data.

### How do I install pgvector in PostgreSQL?

On Linux and Mac, clone the v0.8.6 tag, run make, then make install, which may need sudo. The README also lists Docker, Homebrew, PGXN, APT, Yum, pkg, APK and conda-forge as installation routes, and notes it comes preinstalled with Postgres.app and many hosted providers.

### How do I install pgvector on Windows?

The README requires C++ support in Visual Studio, then an x64 Native Tools Command Prompt for VS run as administrator, followed by nmake /F Makefile.win and nmake /F Makefile.win install. You can also use Docker or conda-forge on Windows.

### How do I use pgvector in PostgreSQL?

Enable the extension once per database with CREATE EXTENSION vector, create a table with a vector column such as vector(3), then query with a distance operator like ORDER BY embedding <-> '[3,1,2]' LIMIT 5. Add an HNSW or IVFFlat index when you want approximate search instead of exact.

## Sources

- [Issues](https://github.com/pgvector/pgvector/issues)
- [pgvector/pgvector on GitHub](https://github.com/pgvector/pgvector)
- [README](https://github.com/pgvector/pgvector/blob/master/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/pgvector-pgvector
