milvus vs qdrant: distributed scale versus filtering-first simplicity
Milvus splits storage and compute into separate services so it can grow to billions of vectors; Qdrant keeps a single Rust server process and leans on payload filtering. They solve adjacent problems, and for many teams the real question is which one fits the workload you actually have.
At a glance
| Project | milvus-io/milvus | qdrant/qdrant |
|---|---|---|
| Licence | Apache-2.0Permissive: commercial use allowed | Apache-2.0Permissive: commercial use allowed |
| Maintenance | Commits in the last dayLast push September 29, 2026 | Commits in the last dayLast push September 29, 2026 |
| Language | Go | Rust |
| GitHub stars | 46,282 | 34,881 |
| Read more | Our analysisGitHub | Our analysisGitHub |
Which one to choose
Choose milvus if you are already running Kubernetes, expect your collection to reach hundreds of millions or billions of vectors, need to scale reads and writes independently, or want a pip-installable Lite mode for local development that shares the same client API as the server.
Choose qdrant if your queries combine similarity with precise metadata constraints, you want one container to start and a single Rust binary to reason about, or you need to embed search in a process through Qdrant Edge without running a cluster.
Two architectures that disagree about where complexity belongs
Milvus, written in Go and C++, is described in its README as a fully distributed and K8s-native architecture. The design separates compute from storage: query nodes serve reads, data nodes absorb writes, and coordinators manage the whole set. That separation is the reason the README claims Milvus can independently scale query nodes for read-heavy traffic and data nodes for write-heavy traffic. It also supports replicas, which load the same data segments onto multiple query nodes for fault tolerance and throughput. The cost is that a production deployment is a collection of stateless microservices rather than one process. Milvus does offer a Standalone mode for a single machine and a Lite mode installed with pip, so the distributed shape is not mandatory at every stage.
Qdrant takes the opposite stance. It is a single Rust service that exposes REST and gRPC, started with one docker run command in the README. There is no separate storage tier to operate. Filtering is treated as a first-class part of the query engine rather than a layer bolted on top: the README says Qdrant is tailored for extended filtering support and lists faceted search as a target use case. It also stores points as vectors plus a payload, so metadata travels with the vector. Qdrant Edge runs the same engine inside the application process, which the README contrasts with the client-server mode. The trade-off is that horizontal scale is something you add through sharding and replication rather than something the base architecture hands you.
Getting each one running on a laptop
Milvus has the gentler first five minutes for Python users. The README shows pip install -U pymilvus, then MilvusClient("milvus_demo.db") to create a local file-backed database with no server at all. Installing pymilvus[milvus-lite] pulls the Lite variant. The same client class then accepts a uri and token to talk to a self-hosted server or Zilliz Cloud, so code written against the local file can move to a remote endpoint by changing constructor arguments. That continuity is the strongest practical argument for Milvus in a prototype that is expected to grow. The README also links a Docker standalone install document, which is the middle rung between Lite and a full cluster.
Qdrant's quickstart is a container. The README gives docker run -p 6333:6333 qdrant/qdrant, then QdrantClient(url="http://localhost:6333") from Python. The README explicitly warns that this command starts an insecure deployment without authentication, open to all network interfaces, and points to a security guide before production. That warning matters: the one-line start is a development convenience, not a deployment recipe. Qdrant Edge is the embedded path, initialized through EdgeShard.create with a local directory and a vector configuration, and the README says it can synchronize with a Qdrant server. Official clients exist for Go, Rust, JavaScript, Python, .NET and Java, with community clients for Kotlin and PHP. Neither README documents a rollback procedure for a failed upgrade.
Filtering, index choice and query behaviour
Milvus separates the system from the core vector search engine, which the README says lets it support all major vector index types, naming HNSW, IVF, FLAT, SCANN and DiskANN, along with quantization-based indexes. The README also states that Milvus implements hardware acceleration for CPU and GPU. That breadth is the point: you pick an index to match the recall, latency and memory profile of a specific collection. It also means the index decision is yours to get wrong, and the earlier analysis of Milvus advises verifying your exact index type and GPU support against current documentation before committing. The README describes vector search combined with metadata filtering and hybrid search, and stores integers, strings and JSON alongside vectors.
Qdrant's pitch is narrower and more specific. It pairs dense, sparse and multi-vector search with payload filtering, according to the earlier analysis, and the README frames extended filtering as the differentiator. Where Milvus asks which ANN index fits, Qdrant asks how your filter patterns interact with the query planner. The earlier analysis warns to test how payload filter patterns affect query latency, which is the right instinct: a filter that is selective early in the plan behaves very differently from one applied after candidate retrieval. Qdrant also ships agent skills, per the README, aimed at helping a coding assistant choose quantization, sharding, tenant isolation and hybrid search settings. That is a documentation delivery choice, not a search feature, but it lowers the cost of the configuration decisions.
Operations, scaling and failure modes
Milvus is the heavier operational commitment and the README is upfront about why. Stateless microservices on Kubernetes allow quick recovery from failure, and coordinators have a high-availability document linked from the README. Replicas improve fault tolerance and throughput. The same properties mean you are operating a distributed system: multiple node types, a coordinator layer, and the storage and message infrastructure that a Milvus cluster depends on. The earlier analysis says plainly to confirm that your deployment environment can handle the operational overhead of the distributed mode. If your team has no Kubernetes practice, Standalone mode or Milvus Lite is the honest starting point, and the migration to distributed is a project rather than a config flag.
Qdrant's operational surface is a single service, which is easier to reason about until you need to scale writes across machines. Sharding and replication are configuration you apply, and the earlier analysis says to verify your sharding needs against the current release notes. Qdrant Edge changes the failure model again: because it runs in-process and stores data locally, the README says it suits low latency and offline functionality, with synchronization back to a server. That is a genuinely different deployment shape, and it is the one case where Qdrant reaches somewhere Milvus Lite does not, since Lite is a local file for development rather than a documented edge synchronization story. Neither README documents rollback, so upgrade planning rests on the changelogs.
Licence, governance and release cadence
Both projects are Apache-2.0, so the licence is not a differentiator. The governance is. Milvus is under the LF AI & Data Foundation, with Zilliz as its major contributor, and the README links a managed Zilliz Cloud with Serverless, Dedicated and BYOC options. That structure gives the project a neutral foundation home and a commercial vendor with a direct interest in the hosted product. For most users this is neutral to positive: the open core is permissively licensed and the hosted option exists if you would rather not operate the cluster. It does mean roadmap priorities are shaped by a company whose revenue comes from the managed service, which is worth knowing when you depend on a feature that mainly benefits self-hosting.
Qdrant is a company-backed project too, with Qdrant Cloud including a free tier per the README. Neither repository is archived, and both show a last push of 2026-09-14, so both are being worked on now. Release cadence differs in shape: Milvus shipped v3.0.0 on 2026-07-29 with v2.6.23 on 2026-08-28, so a major line and a maintenance line are moving in parallel. Qdrant shipped v1.19.0 on 2026-08-05 and v1.18.3 on 2026-07-17, a steadier minor-version rhythm. The earlier analysis of Qdrant notes frequent releases and advises checking the changelog for breaking changes before upgrading. Milvus carries the same risk with the added wrinkle of a 3.0 line landing next to 2.6 patches.
Where each one is the wrong tool
Milvus is the wrong choice when the workload is a small prototype and you have not yet decided whether vector search belongs in the product. The earlier analysis says to avoid it if a small prototype is better served by Milvus Lite or a simpler tool, and that is the correct read. Standing up the distributed mode to serve a few hundred thousand vectors buys you operational work and no search quality. The second failure mode is index and hardware assumptions: the README advertises GPU acceleration and a long index list, but the earlier analysis warns to verify your exact index type, GPU support and multi-tenancy isolation level against current documentation. Multi-tenancy isolation in particular is a property you should confirm rather than assume, because it determines whether one tenant's load can affect another's.
Qdrant is the wrong choice when you need a fully embedded database with no server process, unless you accept Qdrant Edge, which the earlier analysis notes still runs in-process but with a smaller footprint. If your mental model is an embedded library with the full feature set of the server, Qdrant Edge is a different product with its own constraints. The second limitation is scale shape: Qdrant's filtering strength does not remove the need to plan sharding, and the earlier analysis advises verifying sharding needs against release notes and testing filter latency. If your query mix is dominated by unfiltered brute-force similarity at very large scale, the filtering advantage is mostly unused while the scaling work remains.
Choosing for a concrete workload
For a RAG service over a few million documents with heavy metadata constraints (tenant, date range, document type), Qdrant is the cleaner fit. The filter engine is the product's stated focus, one container covers the whole deployment, and the Python client connects with a URL. You avoid operating a cluster for a scale that does not need one.
For a recommendation or image search system expected to reach hundreds of millions of vectors with bursty read traffic, Milvus is the better shape. Independent scaling of query and data nodes and replicas for throughput are architectural properties, not tuning flags, and the earlier analysis frames Milvus as the choice for horizontally scalable search with real-time streaming updates. If you already run Kubernetes, the marginal operational cost is lower than it looks.
For an application that must work offline on a device and sync later, Qdrant Edge is the option with a documented story in the README. Milvus Lite is a local file for development, not an edge synchronization design. For an air-gapped proof of concept that may become a cluster, Milvus Lite to Standalone to distributed is a path where the client code survives the transition, which is a real advantage when the prototype outgrows its first home.
Bottom line
Pick Milvus when scale and independent read/write scaling are the requirement and you can absorb distributed operations; pick Qdrant when filtering precision and a single-service deployment matter more. Before committing, verify on the Milvus side your index type, GPU support and multi-tenancy isolation against current docs, and on the Qdrant side your sharding plan and how your filter patterns affect latency. Neither README documents rollback, so read the changelogs before upgrading.