# NVIDIA cuVS: GPU Vector Search and CAGRA Indexes

> cuVS is NVIDIA's library for building approximate nearest neighbor indexes and running clustering on the GPU, exposed through Python, C++, C, Rust, Java and Go bindings. It is a building block for databases and retrieval systems, not an end-user application, and it installs through conda, pip or a tarball.

**NVIDIA/cuvs** — cuVS - a library for vector search and clustering on the GPU

- Repository: https://github.com/NVIDIA/cuvs
- Website: https://docs.nvidia.com/cuvs/
- Stars: 857 · Forks: 238
- Language: Cuda
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/nvidia-cuvs

## What cuVS solves, and who is actually supposed to use it

cuVS is a GPU library for two jobs: approximate nearest neighbor search and clustering. The README frames the motivation around embeddings generated from unstructured data, where you need semantic search over vectors rather than exact keyword matching. That is the retrieval-augmented generation case, recommender systems, image and audio search, and molecular search. The second job is data mining: clustering algorithms, visualization algorithms such as UMAP and t-SNE, k-NN graph construction, sampling and class balancing.

The audience is narrower than the topic list suggests. cuVS is not a vector database you point an application at. The README states it can be used directly or through the databases and other libraries that have integrated it, and that its primary goal is to simplify the use of GPUs for vector similarity search and clustering. So the realistic adopter is someone writing the retrieval layer, or maintaining a database that embeds cuVS. If you want an HTTP endpoint with collections and metadata filters, this is one layer below what you want.

One design point worth noting: the README lists interoperability as a benefit, described as building on GPU and deploying on CPU. That is a real constraint on how you think about indexes. The index is not a black box you can only query where you built it.

## CAGRA and the RAFT layer underneath

The getting started example builds a CAGRA index, which is the algorithm cuVS uses for its headline example across every language binding. The shape of the call is consistent everywhere: you construct index parameters, then build an index from a dataset. In Python that is cagra.IndexParams() followed by cagra.build(index_params, dataset). In C++ it is cagra::index_params and cagra::build(res, index_params, dataset), where res is a raft::device_resources object. The C API follows the same three steps but with explicit handles: cuvsResourcesCreate, cuvsCagraIndexParamsCreate, cuvsCagraIndexCreate, then cuvsCagraBuild, and then a matching destroy call for each handle.

That C API shape tells you something about the intended integration. Every object is created and destroyed manually, and the dataset is passed as a DLManagedTensor, which is the DLPack interchange struct. This is how a database written in C or C++ hands cuVS a tensor it already has in device memory without copying it into a cuVS-specific container. The Rust binding uses the same DLPack types, importing AsDlTensor, AsDlTensorMut and DLTensorView.

cuVS itself sits on top of RAPIDS RAFT, described in the README as a library of high performance machine learning primitives. RAFT supplies the device resources and matrix views that cuVS APIs take as arguments. The README also states that cuVS shoulders the burden of keeping accelerated code current as new NVIDIA architectures and CUDA versions are released. That is the maintenance argument for the library, and it is also the reason the library is versioned against CUDA and RAPIDS releases rather than standing alone.

## Installing cuVS and building a first CAGRA index

The README says pre-built packages are available through conda and pip, or as a tarball from the NVIDIA download page, with different packages per supported language. It points to the Build and Install Guide for the details. The repository also ships a Dockerfile that builds a full image from nvidia/cuda with Miniforge, and its usage comment gives the two commands directly.

```bash
docker build -t cuvs:latest .
docker run --gpus all -it cuvs:latest
```

The Dockerfile takes three build arguments you can override, with defaults of CUDA_VER=12.9.1, PYTHON_VER=3.12 and RAPIDS_VER=25.06. Supported Python versions are listed as 3.11, 3.12 and 3.13. If you need a different CUDA toolkit, the documented form is a --build-arg override, for example CUDA_VER=12.8 and RAPIDS_VER=25.06.

Once the package is installed, the README's Python example is the shortest path to a working index. It loads a dataset, creates default CAGRA index parameters, and builds the index.

```python
from cuvs.neighbors import cagra

dataset = load_data()
index_params = cagra.IndexParams()

index = cagra.build(index_params, dataset)
```

The load_data() call is a placeholder in the README, not a cuVS function, so you supply your own array. After the build returns, you have an index object that the search routines in the same module accept. The README does not show the search call in this snippet; the examples directory in the repository is where the self-contained code examples live, and the README points to it for C++ and C, including drop-in CMake project templates.

## Binary size is a documented installation problem

The README carries a note that is unusual to see stated so plainly: if compiled binary size is a concern, cuVS builds for CUDA 13 are roughly half the size of CUDA 12 builds. The stated cause is improved compression rates in the newer supported CUDA drivers. NVIDIA says it will adopt the newer drivers for CUDA 12 builds in Spring of 2026, which would bring them down to roughly the CUDA 13 size.

Until then, the README offers two options. Build from source, or use the pre-built libcuvs-static conda package to link statically. For anyone shipping a container or a desktop application, this is a real decision point rather than a footnote. A dynamic link to a CUDA 12 build can roughly double the footprint of the cuVS portion of your image compared with CUDA 13, and the note explicitly recommends static linking as the workaround. That is a trade-off, not a free win: static linking removes the shared library dependency but puts the burden of rebuilding on you whenever cuVS or CUDA moves.

There is a second dimension to this. The README's maintenance claim is that cuVS absorbs the churn of new NVIDIA architectures and CUDA versions. That is true for the library, but it does not remove the churn from your build. You still choose a CUDA version, and that choice is visible in your binary size.

## Where cuVS is the wrong tool

cuVS assumes a GPU. The README's whole framing is GPU-accelerated vector search, and the C++ example requires a raft::device_resources object, which is a device-side resource handle. There is no described CPU-only execution path for building or searching an index. The interoperability benefit means you can build on GPU and deploy on CPU, but the build step itself is a GPU operation. If your deployment target has no NVIDIA GPU and you cannot afford one for index construction, this library is not the answer.

Memory is the second boundary. The dataset in every example is a device matrix or a DLPack tensor, which implies the vectors need to be resident on the GPU. The README does not document an out-of-core or host-memory index build path, and it does not describe how to shard an index across multiple devices. If your embedding set is larger than device memory and you have no sharding strategy of your own, you are outside what the documentation covers.

The third boundary is scope. cuVS is a library, not a service. The README lists no persistence format, no replication, no snapshotting and no rollback for indexes. Nothing in the README describes how an index survives a process restart. If you need durability guarantees, you get them from whatever system embeds cuVS, not from cuVS.

## How it compares with building on FAISS or running a vector database

The closest comparison in the README is not named as a competitor, so the honest one is architectural. A CPU-oriented ANN library such as FAISS gives you index types you build and query on the host, with GPU support as an add-on path. cuVS inverts that: the GPU is the primary target, RAFT supplies the device primitives, and the CPU story is described as a deployment target for an index built elsewhere. If your vectors are small enough that a CPU index answers within your latency budget, the GPU dependency buys you nothing and costs you a device.

The other comparison is against vector databases. A database gives you an HTTP or client API, collection management, filtering, and persistence. cuVS gives you build and search calls over tensors you already hold. The README states cuVS can be used through the databases and other libraries that have integrated it, which is the intended relationship: the database owns the operational surface, cuVS owns the accelerated kernel. Choosing between them is really choosing whether you want to operate a service or write one.

A third option the README mentions indirectly is the graph route. It notes that converting dense vectors into nearest neighbor graphs unlocks graph analysis algorithms such as those in GraphBLAS and cuGraph. If your downstream work is graph analysis rather than top-k retrieval, that framing is more relevant than a pure search comparison.

## Licence, releases and the cost of staying current

cuVS is Apache-2.0, and the source files carry SPDX headers naming NVIDIA Corporation and Affiliates. Apache-2.0 is a permissive licence with an explicit patent grant and a requirement to preserve notices, which matters if you link cuVS into a product. The repository also includes a SECURITY.md and a CONTRIBUTING.md, and the README points to the RAPIDS community pages and the GitHub issue tracker for help and feature requests. This is not legal advice; if you are redistributing a binary that statically links libcuvs-static, read the licence text and your own obligations rather than relying on a summary.

On maintenance, the repository is not archived and the last push was on 2026-09-10. Recent releases are v26.08.01 on 2026-08-06, v26.08.00 on 2026-08-05, and v26.06.00 on 2026-06-04. The version scheme tracks RAPIDS releases, which means upgrades arrive on a monthly-ish cadence rather than a slow one, and the Dockerfile's default RAPIDS_VER of 25.06 shows how quickly a pinned version drifts from the current release.

The upgrade cost is the part to budget for. Your build pins a CUDA version and a RAPIDS version, the library's stated job is to track new NVIDIA architectures and CUDA releases, and the binary size note shows that even the CUDA version you pick changes your artifact. There is also a dependencies.yaml and a conda/ directory in the repository, which is where the pinned dependency set lives if you build from source. Expect to rebuild rather than to install once and forget.

## Conclusion

Adopt cuVS if you are building a retrieval or clustering pipeline that already runs on NVIDIA GPUs and you need index construction and search as a library rather than a server. Do not adopt it if your vectors live on CPU-only hardware, if you want a managed vector database with replication and filtering, or if you need a stable ABI you cannot rebuild against. Before committing, verify that a conda or pip package exists for your CUDA version, that your data fits in device memory, and whether the static libcuvs-static package is required to keep your binary size down.

## FAQ

### How do I install cuVS?

The README says pre-built packages are available through conda and pip, or as a tarball, with different packages for the different supported languages. It points to the Build and Install Guide for the full instructions, and the repository ships a Dockerfile that builds an image with CUDA, Miniforge and a conda environment.

### What is NVIDIA cuVS?

cuVS is a library for vector search and clustering on the GPU, containing implementations of approximate nearest neighbor and clustering algorithms. The README states it can be used directly or through databases and libraries that have integrated it, and that its primary goal is to simplify the use of GPUs for vector similarity search and clustering.

### Which programming languages can I use cuVS from?

The README shows the same CAGRA index example in Python, C++, C and Rust, and the repository contains c/, cpp/, go/, java/, python/ and rust/ directories plus matching example folders. The README describes multiple language support as one of the benefits of using cuVS.

### Does cuVS need a GPU to build an index?

The examples pass device-side data: the C++ example uses raft::device_resources and the C API takes a DLPack tensor. The README lists interoperability as building on GPU and deploying on CPU, so the build step itself is described as a GPU operation.

### Why are CUDA 12 cuVS builds larger than CUDA 13 builds?

The README states that cuVS builds for CUDA 13 are roughly half the size of CUDA 12 builds, attributing this to improved compression rates in the newer supported CUDA drivers. It says newer drivers will be adopted for CUDA 12 builds in Spring of 2026, and suggests the libcuvs-static conda package or a source build if binary size matters now.

### What licence is cuVS released under?

cuVS is Apache-2.0, and source files in the repository carry SPDX headers identifying NVIDIA Corporation and Affiliates as the copyright holder.

## Sources

- [License: Apache-2.0](https://github.com/NVIDIA/cuvs/blob/main/LICENSE)
- [NVIDIA/cuvs on GitHub](https://github.com/NVIDIA/cuvs)
- [Project website](https://docs.nvidia.com/cuvs/)
- [README](https://github.com/NVIDIA/cuvs/blob/main/README.md)
- [Releases](https://github.com/NVIDIA/cuvs/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/nvidia-cuvs
