# Milvus Lite gives you a local vector database in one pip install

> A distributed vector database written in Go and C++ whose quickest path runs inside your own Python process, backed by a separate Milvus server that needs etcd, object storage and a message queue. The same codebase ships both, and the client is the only part that behaves the way a library should.

**milvus-io/milvus** — Milvus is a cloud-native vector database in Go and C++ for scalable ANN search over billions of vectors, with CPU/GPU acceleration and real-time streaming updates.

- Repository: https://github.com/milvus-io/milvus
- Website: https://milvus.io
- Stars: 46,282 · Forks: 4,276
- Language: Go
- License: Apache-2.0
- Published: 2026-08-08 · Updated: 2026-08-18 · Language: en
- Canonical page: https://hysenlabs.com/projects/milvus-io-milvus

## Milvus Lite is a file, not a server, and the client call barely changes

The quickest route in the project starts with one install:

```bash
pip install -U pymilvus
```

That gives you pymilvus, the Python SDK. Installing pymilvus[milvus-lite] adds the embedded variant, and the whole local database comes into existence when you pass a filename to the client:

```python
client = MilvusClient("milvus_demo.db")
```

The same class, MilvusClient, connects to a deployed Milvus server or to Zilliz Cloud, where the call takes a uri and a token instead of a path. That symmetry is the design decision worth understanding. Your collection, insert and search code does not change when you move from a laptop to a cluster, which is convenient, and it also hides the difference: one process owns the data locally, the other is a client of a service with replicas, coordinators and object storage behind it. Treat the local file as a development convenience, not as a deployment target.

## Compute and storage are separate layers, which is why the server needs three services

The distributed architecture splits compute from storage, and horizontally scales by adding query nodes for read-heavy traffic and data nodes for write-heavy traffic. Stateless microservices on Kubernetes allow quick recovery from failure, and replicas load data segments onto multiple query nodes. The repo's docker-compose.yml makes the dependency list concrete: the builder service depends_on etcd, minio, pulsar, azurite and gcpnative, and it takes ETCD_ENDPOINTS, MINIO_ADDRESS, PULSAR_ADDRESS and AZURE_STORAGE_CONNECTION_STRING from the environment. Object storage comes from MinIO in development with Azurite standing in for Azure. A reader expecting a single binary gets a service topology instead, and a reader expecting pgvector gets none of this, so the operational surface is the real cost of running the server rather than the client library.

## The index choice is HNSW, IVF, FLAT, SCANN, DiskANN, and the tradeoff is yours

Milvus separates the system from the core vector search engine, which is what allows it to support the major index types: HNSW, IVF, FLAT as brute-force, SCANN and DiskANN, with quantization-based variations such as IVFPQ and mmap as options. Search features are optimised alongside them, including metadata filtering and range search, and hardware acceleration covers GPU indexing with NVIDIA's CAGRA. Two consequences follow. The first is that recall, latency and memory are a decision you make per collection rather than a global setting, and the documentation pushes you to an external benchmark tool on Zilliz's site for the comparison, so no index recommendation ships with the project. The second is that FLAT is the honest baseline, and picking it for a large collection is a decision about brute force that the architecture will not rescue.

## Multi-tenancy is four levels deep, and hot/cold storage changes the bill

Isolation is available at database, collection, partition or partition key level, which the project says lets a single cluster handle hundreds to millions of tenants with access control attached. That is a real capability and also a real footgun, because the level you choose is the level your query planner has to respect: partitioning by key gives you the same table with cheaper isolation, splitting into collections gives you stronger separation and more objects to manage. Hot and cold storage sit alongside it, with frequently accessed data held in memory or on SSDs and the rest elsewhere, so cost depends on a placement decision you configure rather than on data volume alone. Nothing in the project states a default tier, so a deployment inherits whatever the default is until someone reads the storage configuration.

## 2.6 and 3.0 are releasing side by side, so pick a generation deliberately

The release history is the first thing to read. Recent tags include v2.6.25 on 2026-09-29, v3.0.2 on 2026-09-20 and v2.6.24 on 2026-09-16, so a 2.6 patch and a 3.0 patch shipped in the same week on the master branch, and the last push was 2026-09-29. This is a project in the middle of a major version transition, which means the version number on your deployment is a compatibility decision rather than a formality. The repository makes that concrete with a document named milvus20vs1x.md comparing the 2.0 line to 1.x, and another named UPDATE_MILVUS_API.md, both of which exist because the API has changed across generations. Pin a version, read the migration notes for the jump you are making, and do not assume a tutorial written for one generation works on the other.

## disk_index is silently off on macOS because aio is missing there

The Makefile sets disk_index based on the build host's operating system. On Darwin it is OFF, everywhere else ON, with a comment explaining that macOS does not support aio, and a disk_index variable is provided for a manual override. The same file sets build tags for jemalloc and a bytedance_tango tag, and exports a Cargo target root for the Rust workspaces, so a local build is doing more than compiling Go. The consequence is a platform difference in capability that is invisible in the client API. A developer on a Mac who benchmarks a local deployment is not benchmarking the DiskANN path a Linux deployment will take, and the two results will not be comparable. Check which way that variable resolved before you trust a local measurement.

## Governance is outsourced to Zilliz, and the managed tiers are the same vendor

The project sits under the LF AI and Data Foundation and ships with Apache 2.0, with Zilliz named as its major contributor. The same company runs the hosted options the README points at: Zilliz Cloud in Serverless, Dedicated and BYOC flavours, and the repository's own Makefile header still carries a 2019-2020 Zilliz copyright. The licence is unambiguous, which is the part that matters for adoption, but the operational relationship is worth naming. Bug reports and feature requests go to GitHub Issues and Discussions, community help goes to Discord and Slack, and self-hosted deployments are supported. A team that wants a managed vector database and a team that wants an open-source one are looking at the same code, and the support path differs.

## Conclusion

Use Milvus Lite for prototypes, notebooks and tests, where a local file and the pymilvus client are enough, and deploy the standalone or distributed server when replicas, hybrid search or GPU indexing matter. Before you commit, read milvus20vs1x.md to find out which generation you are writing code against, and check the disk_index default, since the Makefile turns it off on macOS because that platform has no aio.

## FAQ

### What is Milvus used for?

It is a vector database for approximate nearest neighbour search over learned representations of unstructured data such as text, images and multi-modal information, stored alongside scalar fields such as integers, strings and JSON. Searches can be combined with metadata filtering or run as hybrid search.

### Is Milvus free to use?

The open-source project is distributed under the Apache 2.0 licence and is under the LF AI and Data Foundation, with Zilliz as its major contributor. Zilliz Cloud is a separate commercial offering with Serverless, Dedicated and BYOC options.

### How do I install Milvus Lite?

Install the client with pip install -U pymilvus, then install pymilvus[milvus-lite] for the embedded version. You create a local database by instantiating the client with a file name, for example client = MilvusClient("milvus_demo.db").

### How do I install Milvus?

The project supports Standalone mode for a single machine deployment, and the README links a standalone Docker install guide. The repository's docker-compose.yml shows the server needing etcd, minio and pulsar alongside it, with Azurite standing in for Azure storage.

### How do I connect to a deployed Milvus server with pymilvus?

Instantiate MilvusClient with a uri pointing at the endpoint of your self-hosted Milvus server or Zilliz Cloud, and a token holding your username and password or your Zilliz Cloud API key. The client class is the same one used for the local Milvus Lite file.

### Which vector index types does Milvus support?

The major index types are HNSW, IVF, FLAT for brute-force search, SCANN and DiskANN, with quantization-based variations such as IVFPQ and mmap. Milvus also supports GPU indexing, including NVIDIA's CAGRA, and optimises search for metadata filtering and range search.

## Sources

- [Official documentation](https://milvus.io)
- [Official README](https://github.com/milvus-io/milvus#readme)
- [Project repository](https://github.com/milvus-io/milvus)
- [Release notes](https://github.com/milvus-io/milvus/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/milvus-io-milvus
