Milvus: A Distributed Vector Database for AI Similarity Search
Milvus is a cloud-native vector database in Go and C++ for scalable ANN search over billions of vectors, with CPU/GPU acceleration and real-time streaming updates.
At a glance
- What is it?
- Milvus is an open-source, Apache 2.0-licensed vector database written in Go and C++, built to store and search high-dimensional vectors at scale. It targets AI engineers who need approximate nearest neighbor search over billions of vectors, with support for metadata filtering, hybrid search, and GPU indexing.
- Who is it for?
- Milvus suits teams building AI applications that need similarity search over large vector collections, particularly when GPU acceleration, multi-tenancy, or a K8s-native deployment model are requirements. Milvus Lite is a reasonable starting point for early prototyping.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository received new commits within the last day.
- What is it written in?
- Mainly Go, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
DEEP OPEN-SOURCE ANALYSIS
What Milvus Solves and Who Needs It
Traditional relational and document databases do not support efficient similarity search over high-dimensional vectors. Searching for the nearest neighbors of a 768-dimensional text embedding in a table of 100 million rows requires either brute-force comparison of every row (too slow) or a specialized index structure. Milvus provides that index structure, along with the storage, ingestion pipeline, and query layer to operate it at scale.
The README describes its use case as powering AI applications that work with unstructured data such as text, images, and multi-modal information. This includes retrieval-augmented generation (RAG) pipelines, image similarity search, recommendation engines that rely on learned embeddings, and anomaly detection systems. Milvus stores vectors alongside scalar data types including integers, strings, and JSON objects, so metadata filtering can be combined with vector search in a single query.
Architecture: Separated Compute and Storage on Kubernetes
Milvus uses a fully distributed architecture that separates the query (compute) layer from the storage layer. The README states that this allows read-heavy and write-heavy workloads to scale independently: query nodes handle search, and data nodes handle ingestion. The components are stateless microservices deployed on Kubernetes, which the README says enables quick recovery from failure.
The system uses etcd for metadata coordination, MinIO or S3-compatible object storage for vector data, and Pulsar or Kafka for the streaming message queue. These dependencies are visible in the docker-compose.yml, which defines services for etcd, minio, pulsar, and an azurite emulator for Azure Blob Storage. For production Kubernetes deployments, the standard path is through Helm charts documented at milvus.io.
Milvus also ships a Standalone mode for single-machine deployment and Milvus Lite, a lightweight version that runs as a local file-backed database from a Python process with no server required.
Getting Started: Milvus Lite for Local Prototyping
The README shows three entry points depending on scale requirements. For a local prototype without running any server, install the pymilvus package with the milvus-lite extra:
pip install -U pymilvusThen instantiate a client pointed at a local file:
from pymilvus import MilvusClient
client = MilvusClient("milvus_demo.db")This creates a file-backed vector database in the current directory. To create a collection with 768-dimensional vectors and insert data:
client.create_collection(
collection_name="demo_collection",
dimension=768,
)Once data is inserted, a vector search looks like this:
res = client.search(
collection_name="demo_collection",
data=query_vectors,
limit=2,
output_fields=["vector", "text", "subject"],
)The same MilvusClient interface works against a self-hosted Milvus server by passing the server URI and authentication token. This means code written against Milvus Lite can be moved to a production server without changing the application logic.
Index Types and GPU Acceleration
The README lists the supported vector index types: HNSW (hierarchical navigable small world graphs), IVF (inverted file), FLAT (brute-force exact search), SCANN, and DiskANN. DiskANN stores the index on disk rather than in memory, which allows searching over datasets larger than available RAM. The README notes that quantization-based variations of IVF are available for reducing memory footprint.
For GPU-accelerated search, Milvus supports NVIDIA's CAGRA index, which is part of the cuVS library. This is relevant for teams with GPU infrastructure who need to reduce search latency at high query rates. The Makefile shows that disk_index is set to OFF on macOS and ON on Linux by default. GPU indexing has the same Linux-only constraint in practice.
The go.mod file shows the project requires Go 1.26.6. The build system is a Makefile that orchestrates Go compilation, CGo linking for the C++ vector search components, and Rust compilation for parts of the storage layer (MILVUS_CARGO_TARGET_ROOT is an environment variable used in the Makefile).
Operational Complexity and When Milvus Is the Wrong Choice
Running Milvus in production means operating a microservices stack with at minimum etcd, MinIO, and a message broker (Pulsar or Kafka). For teams without Kubernetes experience or without the operator bandwidth to manage these dependencies, the operational cost is high relative to simpler alternatives.
Milvus Lite is not a substitute for the full server in production. The README describes it as suitable for quickstart and prototyping. It does not offer the high availability, replication, or horizontal scaling of the distributed deployment.
Multi-tenancy isolation in Milvus is configurable at the database, collection, partition, or partition key level. The README notes that a single cluster can handle from hundreds to millions of tenants. However, fine-grained access control at the row level is not a feature described in the available documentation. Teams with strict row-level security requirements should verify this against the current documentation at milvus.io before committing.
The README also mentions that Zilliz, the primary contributor, offers a fully managed Milvus service on Zilliz Cloud with Serverless, Dedicated, and Bring Your Own Cloud options. This is the path for teams who want Milvus capabilities without managing the infrastructure.
Comparison with Qdrant
Qdrant is a direct alternative for approximate nearest neighbor search. It is written in Rust, runs as a single binary with a REST and gRPC API, and supports its own HNSW implementation with payload filtering. Qdrant is designed to be self-contained and easy to run on a single machine, which makes it a lower-friction starting point for teams that are not on Kubernetes and do not need the distributed scaling Milvus offers.
Milvus has a longer production history and is a graduated project under the LF AI and Data Foundation. Its architecture is explicitly built for Kubernetes-native horizontal scaling and GPU acceleration, which Qdrant does not offer natively. The choice between them generally comes down to scale and infrastructure: teams already on Kubernetes who anticipate growing to billions of vectors have more reason to choose Milvus; teams that want a single-server deployment with simpler operations typically find Qdrant easier to manage.
Maintenance Status and License
The repository is not archived. The last push was on 2026-09-26. Recent releases include v3.0.2 on 2026-09-20, v2.6.24 on 2026-09-16, and the Python client v3.0.0 on 2026-09-18. The simultaneous maintenance of a v2 and v3 release line indicates that the project supports multiple major versions concurrently.
The project is licensed under Apache 2.0, which permits commercial use, modification, and redistribution without requiring that derivative works be open-sourced. There is no copyleft restriction. The project is a graduated project of the LF AI and Data Foundation, as stated in the README.
Editorial conclusion
Milvus suits teams building AI applications that need similarity search over large vector collections, particularly when GPU acceleration, multi-tenancy, or a K8s-native deployment model are requirements. Milvus Lite is a reasonable starting point for early prototyping. It is not the right choice for teams who need a simple single-node setup with minimal operational overhead and have no plans to scale to billions of vectors. Before deploying, check the Makefile note that disk_index is disabled on macOS (Darwin), which means DiskANN is not available outside Linux environments. The Apache 2.0 license permits commercial use without restriction.
Frequently asked questions
What is Milvus used for?
Milvus is used to store and search high-dimensional vectors in AI applications, including retrieval-augmented generation pipelines, image similarity search, recommendation systems based on learned embeddings, and semantic text search. It supports metadata filtering and hybrid search alongside vector queries.
Is Milvus free?
Milvus is open-source and licensed under Apache 2.0, so the self-hosted version is free to use, including in commercial products. Zilliz offers a managed Milvus service on Zilliz Cloud with both a free tier and paid dedicated options, as described in the README.
How do I use Milvus Lite?
Install pymilvus with pip install -U pymilvus, then create a MilvusClient pointing to a local file path such as MilvusClient("milvus_demo.db"). This starts a file-backed vector database without any server process. The README describes Milvus Lite as suitable for quickstart and prototyping, not production deployments.
How do I install Milvus locally?
The simplest local installation is Milvus Lite via pip install -U pymilvus. For a full server installation on a single machine, the milvus.io documentation describes a Docker Compose setup that runs Milvus Standalone, which includes etcd and MinIO as dependencies.
How do I use Milvus with LangChain?
The README does not describe a LangChain integration directly, but LangChain has a community-maintained Milvus vector store integration. The connection parameters would use the same MilvusClient URI and token pattern shown in the Quickstart section of the README.
Official sources
Where this project is recommended
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/milvus-io-milvus)
Community notes