What is Vector database?
A vector database (also called a vector store) stores data as high-dimensional vectors and returns the nearest neighbours to a query vector, usually by approximate search. It is the retrieval layer behind semantic search, recommendation and retrieval-augmented generation.
How a vector database works
A vector database stores each item as a fixed-length list of numbers, the embedding, produced by a model that maps text, images or audio into a point in a high-dimensional space. Similar items land near each other, so retrieval becomes a nearest-neighbour problem: embed the query with the same model, then ask for the k vectors closest to it.
Exact nearest-neighbour search compares the query against every stored vector, which is linear in the collection size and quickly becomes too slow. Production systems therefore build an index. The common families are graph indexes such as HNSW, which walk a layered proximity graph and trade recall for speed through parameters like the number of neighbours per node; inverted-file indexes such as IVF, which cluster vectors and probe only the closest cells; and quantisation, which compresses vectors to fewer bits so more of the index fits in memory. Each carries a recall and latency trade-off that must be measured on the actual data.
Distance is a choice, not a constant. Cosine similarity, dot product and Euclidean distance give different rankings unless the vectors are normalised, and a mismatch between the metric used at index time and at query time silently degrades results. Most engines also support payload filtering: a query can combine a vector similarity condition with a structured predicate on stored metadata. Filtering is the hard part of the design, because a selective filter can leave too few candidates for the graph to traverse well. Implementations differ in whether they pre-filter, post-filter or use a filtered graph traversal, and the README of each project is the place to check which applies.
Finally, the vector is only half the record. A vector store typically persists the original object, its metadata and the vector, so that a result can be returned without a second lookup. Some engines, such as Weaviate, describe storing objects and vectors together as their core model. Others, such as Qdrant, expose payloads alongside vectors and treat payload filtering as the distinguishing feature.
When you need a vector database, and when you do not
You need one when the query is semantic rather than literal. Keyword search fails when the user types a paraphrase that shares no tokens with the document; an embedding search can still find it. You also need one when the collection is too large for a linear scan to meet your latency budget, or when results must be filtered by tenant, permission or category at query time.
You do not need one for a few thousand vectors. A NumPy array and an exact dot product will be faster to build and easier to debug than any service. You do not need one if your retrieval is purely lexical and BM25 already works; adding embeddings adds a model, a dimension and an index without fixing a problem you have. And you do not need a separate service if your data already lives in a database that has added vector search, or if your application is single-process and can use an embedded library.
The decision usually comes down to deployment shape rather than algorithm. A client-server vector database gives you shared indexes, concurrent writers and independent scaling, at the cost of operating a stateful service with backups, upgrades and capacity planning. An embedded vector store runs inside your process, removes the network hop and the operational burden, and gives up shared access. Alibaba's Zvec is described as an embedded C++ vector database with SDKs for Python, Node.js, Go, Rust and Dart, putting dense vectors, sparse vectors, full-text search and structured filters behind one local collection object; the same analysis concludes it fits single-writer applications and is the wrong tool once you need concurrent writers or a shared index across machines. sqlite-vec takes the same embedded route inside SQLite, adding a vec0 virtual table for float, int8 and binary vectors, and is still pre-v1 per its analysis.
Milvus sits at the other end. It is an Apache-2.0 vector database written in Go and C++, with a Standalone Docker mode, an embedded Milvus Lite mode for Python, and a distributed Kubernetes architecture for larger workloads; the question is how much of that machinery you actually need. Qdrant is an Apache-2.0 similarity search engine written in Rust, shipped as a client-server service or an in-process EdgeShard, and its main cost is that you operate a stateful service. Weaviate stores objects and vectors together and exposes vector search, BM25 filtering, RAG and reranking through one interface, with a Docker install and a Python client.
Limits and common pitfalls
Approximate search does not guarantee exact results. Recall depends on index parameters, data distribution and filter selectivity, and a configuration that works on one collection can fail on another. If the application assumes exactness, it will occasionally return a worse neighbour than a brute-force scan would, and no amount of tuning removes the trade-off entirely.
Embeddings drift. Change the model, the version or the preprocessing and the old vectors are no longer comparable with new queries; the index must be rebuilt. This is an operational event, not a configuration flag, and it is the most common source of silent quality loss.
Dimension and memory are coupled. A 1536-dimensional float vector occupies about 6 KB before index overhead, so a hundred million vectors is a serious memory footprint unless quantisation or disk-based storage is used. Quantisation reduces memory and can reduce recall; the two must be measured together.
Deletion and updates are harder than insertion. Graph indexes degrade as vectors are removed, and many engines implement deletes as tombstones plus periodic compaction. The README of a given project is often silent on compaction behaviour, and that silence matters when planning capacity.
Filtering interacts badly with graph traversal when the predicate is selective. A tenant filter that matches one percent of the collection can starve the search, and the engine may fall back to a slower exact path. Testing with production-like filter selectivity is more informative than testing on an unfiltered index.
Finally, a vector store is not a database in the transactional sense unless it says so. TiDB is described as built for agentic workloads with ACID guarantees and native support for transactions, analytics and vector search, which is a different promise from a pure similarity engine. Mixing the two expectations is a common planning error.
How vector search shows up in open-source projects
Some projects use a vector store as their retrieval layer. LangChain4j is an idiomatic Java library for LLM applications on the JVM, offering a unified API over popular LLM providers and vector stores; its analysis notes one API over more than 20 model providers and more than 30 embedding stores, plus agents and RAG, and that the release cadence is the thing to plan around. RocketRide Server is an MIT-licensed AI pipeline runtime with a multithreaded C++ core and Python-extensible nodes, built visually inside VS Code and shipped as portable JSON; its description lists support for more than eight vector databases among its nodes, and its analysis notes the documentation is thin on the operational edges. TiDB approaches the same space from the database side, separating SQL computation from storage and adding a columnar replica for analytics, while its description claims native vector search alongside transactions and analytics.
Other projects deliberately avoid a vector store. Graphify-Labs/graphify turns a codebase, with its docs, SQL schemas, configs and PDFs, into a queryable knowledge graph. Its analysis states that it maps a codebase using deterministic tree-sitter parsing with no vector store, targeting developers using Claude Code, Cursor, Codex and Gemini CLI who want to query code structure instead of grepping. That is the clearest illustration that vector search is one retrieval mechanism among several: when the question is structural, a graph answers it without embeddings.
On the engine side, the projects differ in shape more than in purpose. Milvus offers Standalone Docker, embedded Milvus Lite for Python, and distributed Kubernetes. Qdrant offers client-server or in-process EdgeShard. Weaviate stores objects and vectors together and exposes vector search, BM25 filtering, RAG and reranking through one interface. Zvec is embedded and single-writer. sqlite-vec is a pure C SQLite extension that installs through a language package manager or as a loadable extension, and is still pre-v1. Microsoft's SPTAG combines space partition trees with relative neighborhood graphs and adds distributed serving, but its build pulls in SPDK and RocksDB, and the analysis notes the setup is where it bites. SPTAG's last push was 2025-07-18, so it is not actively maintained.
In practice
A vector database is a retrieval index for embeddings, and the useful question is not whether to use one but which deployment shape fits: embedded for single-process, client-server for shared and concurrent access, and a general database with vector support when transactions and analytics matter alongside similarity. Read the project README for the index type, the filtering behaviour and the update model before committing, and test recall and latency on your own data with production-like filters. If your retrieval is structural rather than semantic, Graphify is a concrete example of solving it without a vector store at all.