Qdrant: a Rust vector database with a real filtering engine
Qdrant - High-performance, massive-scale Vector Database and Vector Search Engine for the next generation of AI. Also available in the cloud https://cloud.qdrant.io/
At a glance
- What is it?
- Qdrant is an Apache-2.0 vector similarity search engine written in Rust, shipped as a client-server service or an in-process EdgeShard. Its distinguishing feature is payload filtering, and its main cost is that you operate a stateful service.
- Who is it for?
- Adopt Qdrant if you need filtered vector search over a service you control, or if you want the same query model embedded in-process through Qdrant Edge. Do not adopt it if you want a single Python process with no server: an embedded library such as Chroma is a shorter path, and pgvector is the better fit when your vectors already live in Postgres.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository received new commits within the last day.
- What is it written in?
- Mainly Rust, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
DEEP OPEN-SOURCE ANALYSIS
What Qdrant solves, and who ends up running it
Embedding models turn text, images and products into vectors, but a vector alone answers only one question: which of these is closest. Real applications ask narrower questions. Show me red shoes under 80 euros that are in stock. Qdrant is built around that second question. The README describes it as "a vector similarity search engine and vector database" that stores points, which are vectors plus an additional payload, and it states that Qdrant is "tailored for extended filtering support". The payload is the part that carries your metadata, and filtering happens inside the search rather than after it.
The project targets teams building semantic matching, faceted search and recommendation systems who are willing to run a stateful service. It ships as a client-server system with REST and gRPC interfaces, and also as Qdrant Edge, a lightweight version that runs inside the application process for edge devices and resource-constrained environments. That second form matters for anyone who wants the query model without a separate deployment. The README does not describe a shared-nothing story for Edge beyond local storage and snapshot restore, so treat it as a different deployment shape, not a drop-in replacement for the server.
Points, payloads and why filtering is the interesting part
The data model is a collection of points. Each point has an id, one or more named vectors, and a payload of arbitrary JSON. The README states that Qdrant supports dense vectors for semantic similarity, sparse vectors for full-text search, and multi-vector search, which is why hybrid retrieval is possible inside one engine instead of being stitched together from two systems. Sparse vectors let you keep lexical matching; dense vectors handle paraphrase. A single query can combine them.
The mechanism that separates Qdrant from a plain nearest-neighbour index is that the payload filter is an input to the search, not a post-processing step. The project's own framing is that filtering is extended, and the README points at faceted search as a use case. That is a design commitment with a cost: payload indexes have to be declared and maintained, and a filter on an unindexed field will not behave like a filter on an indexed one. The README does not document the internal execution order of filtered search, so anyone who needs to reason about latency under selective filters should read the documentation rather than assume.
Rust is the implementation language, and the README ties it to behaviour under load: Qdrant is "written in Rust, which makes it fast and reliable even under high load", with a link to published benchmarks. Treat that as the project's claim, not a measured result.
Installing Qdrant locally with Docker and running a first search
The README's client-server quick start is a single container command. It publishes port 6333, which is the REST port used by the Python client below.
docker run -p 6333:6333 qdrant/qdrantThe README adds a warning that this starts an insecure deployment without authentication, open to all network interfaces, and links to a security page. For a laptop this is acceptable. For anything reachable from a network it is not, and the README is explicit about that.
With the container running, the Python client connects by URL. The README gives exactly this snippet.
from qdrant_client import QdrantClient
client = QdrantClient(url="http://localhost:6333")If you prefer to keep everything in one process, the README shows Qdrant Edge instead. It creates a shard on disk, then upserts a point with a four-dimensional vector and a payload.
from qdrant_edge import Distance, EdgeConfig, EdgeVectorParams, EdgeShard, Point, UpdateOperation
shard = EdgeShard.create("./shard", EdgeConfig(
vectors={"my-vector": EdgeVectorParams(size=4, distance=Distance.Cosine)}
))
shard.update(UpdateOperation.upsert_points([
Point(id=1, vector={"my-vector": [0.1, 0.2, 0.3, 0.4]}, payload={"color": "red"})
]))Note the vector size of 4 and the Cosine distance in that example: it is a demonstration of the API shape, not a configuration you would ship. Real embeddings are hundreds or thousands of dimensions, and the size must match whatever encoder produced them. The README does not document how to change the vector size of an existing collection, so plan the schema before you load data.
Where Qdrant is the wrong tool
The strongest argument against Qdrant is operational. The default deployment is a server you have to run, secure, back up and upgrade. The README's own quick start produces an unauthenticated service open to all network interfaces, which means the first production task is reading the security guide, not writing queries. If your team has no appetite for a stateful service, that cost is real and recurring.
Qdrant Edge reduces the footprint but does not remove the constraint. It is a separate product with its own Python and Rust entry points, and the README describes data as stored and queried locally, with optional synchronization to a Qdrant server. It does not promise that every server feature exists in the embedded form.
There is also a schema cost. Vector size and distance are set at collection creation. If you plan to swap embedding models, you need a migration path, and the README does not document one. The project ships agent skills that it says cover model migration and quantization decisions, which suggests the project knows this is a common problem, but the README itself does not walk through the procedure.
Finally, if your corpus is small enough to fit in memory and you want no moving parts, a library that runs inside your application is simpler. Qdrant is aimed at the point where filtering, scale and concurrency stop being incidental.
Qdrant compared with Chroma and Milvus
The two comparisons people search for most are against Chroma and Milvus, and the difference is mostly about where the engine lives.
Chroma is the embedded-first option. You add it to a Python project and it runs there. Qdrant's equivalent path is Qdrant Edge, which the README describes as running inside the application process with local storage and snapshot restore. The distinction is that Qdrant's primary product is still the client-server service, and Edge is presented as the constrained-environment variant. If your mental model is a library, Chroma matches it more directly.
Milvus is the closer architectural relative: a separate vector database service. The README does not compare the two, so any claim about which scales further would be guesswork. What can be said from the repository is that Qdrant is written in Rust and exposes both REST with an OpenAPI 3.0 specification and a gRPC interface, which the README recommends for "faster, production-tier searches". If your stack already speaks gRPC and you care about the wire format, that is a concrete difference to evaluate. The README's feature list is where the honest comparison starts, not a benchmark table.
Licence, releases and the cost of staying current
Qdrant is Apache-2.0, and the Cargo.toml confirms the package licence field. Apache-2.0 is permissive: you can use it commercially, modify it and redistribute it, subject to the usual conditions around notices and patent language. This is not legal advice, and if you embed Qdrant in a product you should have your own counsel review the notice requirements.
The release cadence is visible in the tags. v1.19.0 landed on 2026-08-05, v1.18.3 on 2026-07-17 and v1.18.2 on 2026-06-04. The last push to the default branch was on 2026-08-05, so the repository is not archived and the project is not dormant. Patch releases between minor versions are normal here, which is good for bug fixes and means you should expect to pin versions rather than float on latest.
Upgrade cost depends on how you deploy. A container image is a tag change plus a restart. A cluster is a rolling upgrade with a compatibility window you have to check per release. The README does not document rollback or downgrade procedures, so that is a question for the release notes and the documentation site before you upgrade anything carrying production traffic.
The build itself is not trivial. The Dockerfile uses cargo-chef for layer caching, installs clang, lld, cmake and protobuf-compiler, and pins a Rust toolchain image. Cargo.toml sets rust-version to 1.97 and edition to 2024. If you build from source rather than pulling the published image, budget for a long first compile.
Editorial conclusion
Adopt Qdrant if you need filtered vector search over a service you control, or if you want the same query model embedded in-process through Qdrant Edge. Do not adopt it if you want a single Python process with no server: an embedded library such as Chroma is a shorter path, and pgvector is the better fit when your vectors already live in Postgres. Before anything else, verify the security defaults, because the documented quick-start container has no authentication and is open to all network interfaces.
Frequently asked questions
What is Qdrant used for?
It stores vectors together with a JSON payload and searches them by similarity. The README names semantic matching, faceted search and recommendation systems as the intended uses.
Is Qdrant DB free?
The repository is licensed Apache-2.0, so the software itself is free to use under that licence. The README also mentions Qdrant Cloud as a fully managed option that includes a free tier.
Is Qdrant a RAG?
No. Qdrant is a vector search engine and vector database, not a retrieval-augmented generation pipeline. It is the retrieval layer such a pipeline would query.
What is Qdrant in Python?
It is the official Python client library, qdrant_client. The README shows constructing a QdrantClient with a URL pointing at the server, for example http://localhost:6333.
How to install Qdrant locally?
The README's quick start runs the container with docker run -p 6333:6333 qdrant/qdrant. The README warns that this starts an insecure deployment without authentication and is open to all network interfaces.
How to use Qdrant vector database in Python?
Install the qdrant_client package, then create a QdrantClient pointing at the server URL as shown in the README. From there you create a collection and upsert points with vectors and payloads.
Official sources
Where this project is recommended
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/qdrant-qdrant)
Community notes