# JVector: a graph-based vector index for Java that can be built larger than memory

> JVector is an Apache-2.0 embedded vector search engine for Java that merges the HNSW hierarchy with Vamana graphs and supports two-pass construction and search. It suits JVM teams that need incremental indexing without a separate service, and it is the wrong tool if you want a managed database.

**datastax/jvector** — JVector: the most advanced embedded vector search engine

- Repository: https://github.com/datastax/jvector
- Stars: 1,751 · Forks: 158
- Language: Java
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/datastax-jvector

## Why JVector exists: exact KNN collapses in high dimensions

The README opens with the problem rather than the product. Exact k-nearest-neighbor search segments space with structures such as quadtrees or k-d trees, and those structures work in two or three dimensions but devolve into linear scans as dimensionality rises. The README calls this one aspect of the curse of dimensionality. Once that happens, an approximate answer in logarithmic time is usually more useful than an exact answer in linear time, which is the tradeoff ANN search accepts.

The project's audience follows from that framing. It is for JVM engineers who have vectors in memory or on disk and want an index inside their own process, not behind an HTTP endpoint. The README contrasts two ANN families: partition-based indexes such as LSH, IVF and SCANN, and graph indexes such as HNSW and DiskANN. Its stated reason for choosing graphs is not raw speed but mutability. Graph indexes can be constructed and updated incrementally, while partitioning approaches only work on static datasets specified up front. That argument is why the README says all the major commercial vector indexes use graph approaches.

## How the JVector graph is layered: HNSW on top, Vamana inside

JVector is explicitly a merge of two lineages. It borrows the hierarchical structure from HNSW and uses Vamana, the algorithm behind DiskANN, within each layer. The result is a multi-layer graph with nonblocking concurrency control, which the README says allows construction to scale linearly with core count.

The memory layout is the part worth understanding before you adopt it. Upper layers are held as an in-memory adjacency list per node, so navigation through them performs no I/O. The bottom layer is an on-disk adjacency list per node. JVector stores extra data inline to support two-pass search: the first pass runs against lossily compressed vectors kept in memory, and the second pass reads a more accurate representation from disk.

The first pass can use product quantization, optionally with anisotropic weighting, or binary quantization, or fused PQ where the codebooks are written inline with the graph adjacency list. The second pass can use full-resolution float32 vectors or NVQ, described as a non-uniform technique for high-accuracy quantization. The README's claim is that this two-pass design lowers memory use and latency while preserving accuracy.

The more unusual capability is that the same two-pass machinery can build the index, not just query it, which the README says allows indexes larger than memory. The stated benefit is staying inside one index with logarithmic search instead of merging results linearly across several smaller indexes. If you have ever split a corpus because it did not fit, that is the design decision being targeted.

## Installing JVector from Maven and running the intro tutorial

The README does not print a Maven coordinate, so the artifact has to come from the repository's own build rather than a copy-paste snippet. The release metadata in package.json names version 4.0.0-rc.9, which matches the newest release listed on the repository, so expect a release candidate rather than a stable 4.x line. The project is a multimodule Maven build intended to produce a multirelease jar usable as a dependency from Java 11 code, with optimized vector providers when it runs on Java 20 or newer with the Vector module enabled.

The repository uses a Git submodule for Google Highway, and the README says to initialize it after cloning:

```bash
git submodule update --init
```

A single-step alternative is given as well, which pulls the submodule during the clone itself:

```bash
git clone --recurse-submodules <repo-url>
```

The native SIMD library, libjvector.so, is built with Meson and Ninja and requires g++ 11 or later. The entry point is a shell script, and the README runs it from its own directory:

```bash
cd jvector-native/src/main/native
bash build_native_lib.sh
```

On a fresh Ubuntu machine the README offers a one-step path that installs g++, Meson and Ninja before building. Note the directory change into src, which differs from the plain build above:

```bash
cd jvector-native/src/main/native/src
bash build_native_lib.sh --auto-install-deps
```

On other distributions the script prints the install commands it needs rather than running them, so you install those yourself. For a first real use, the README points at docs/tutorials, starting with 1-intro-tutorial.md, and at VectorIntro.java under jvector-examples as a simple example. The older docs/legacy/jvector-step-by-step.md guide is still described as useful commentary for advanced users, though new users are told to start with the tutorials.

## Where JVector is the wrong choice, and what the README leaves out

JVector is a library, not a service. There is no server to start, no port to open and no REST API in the README, so if your architecture assumes a standalone vector database that other languages call over the network, this is not that. The homepage field is empty and the README points to no hosted product; distribution runs through the repository and its build.

The build itself is a real constraint. The native path needs Meson, Ninja and g++ 11 or later, plus a Git submodule that must be initialized or the native module will not build. The project is structured to be built with JDK 20 or newer, and the README notes that with JAVA_HOME set to Java 11 through Java 19 certain build features are still available, which implies others are not. The optimized vector providers are conditional on Java 20 or newer with the Vector module enabled, so on an older JVM you are not running the configuration the project is designed around.

Version maturity is the other thing to weigh. The most recent release is 4.0.0-rc.9 from 2026-07-21, and the one before it is labelled a hotfix for a release candidate. Nothing in the README promises API stability across those candidates. The README also does not document rollback, migration between index versions, or what happens to an on-disk index when you upgrade, although the repository does carry an UPGRADING.md file at the top level. That file is the place to look, and its existence is the only evidence available here about upgrade practice. The last push to the repository was on 2026-09-10.

## JVector against Lucene and Faiss: different problems, different shapes

The most common comparison for a Java vector index is Apache Lucene, which also ships graph-based vector search. The difference in approach is scope. Lucene's vector support lives inside a full-text search library with segments, commits, analyzers and a query language; adopting it means adopting that indexing model. JVector is a standalone index with no text layer, so you own storage, persistence and lifecycle yourself. That is less machinery if vectors are all you have, and more work if you also need filtering over text.

Against Faiss, the split is language and mutability. Faiss is a C++ library with Python bindings and is widely used for offline evaluation and static indexes. JVector is Java-first and, per the README, built around incremental construction and updates rather than a dataset fixed in advance. If your pipeline is a Python notebook scoring a frozen corpus, Faiss fits that shape. If your pipeline is a long-running JVM service where vectors arrive continuously, JVector's design is aimed at exactly that case.

There is also an OpenSearch connection implied by the related searches around opensearch-jvector and DataStax. The README does not describe that integration, so treat it as a separate project to evaluate on its own terms rather than as documentation for this library.

## Licence and upgrade cost

JVector is licensed under Apache-2.0, and the repository carries LICENSE.txt and NOTICE.txt at the top level along with a CONTRIBUTIONS.md. Apache-2.0 is a permissive licence with an explicit patent grant, which matters for a search component you embed in a commercial product. The NOTICE file exists, so if you redistribute the library, check what it asks you to carry forward. This is a description of the repository contents, not legal advice; your own counsel decides what your distribution obligations are.

Upgrade cost is harder to pin down from the README. The repository has an UPGRADING.md file at the top level, and CHANGELOG.md is maintained automatically: package.json wires the version script to auto-changelog with the -p flag and then stages CHANGELOG.md. That means the changelog is generated from commit history rather than written by hand, which is useful for seeing what changed and less useful for judging whether a change is safe. With releases still on 4.0.0-rc numbers, budget for reading UPGRADING.md before each bump rather than assuming drop-in compatibility. If you build indexes with the native library, upgrades may also mean rebuilding libjvector.so, since that artifact is produced by the Meson and Ninja build rather than shipped as a managed dependency.

## Conclusion

Adopt JVector if you are building a JVM service or library that needs an in-process graph index which can be updated incrementally, and you are willing to build the native SIMD library or fall back to the pure-Java path. Do not adopt it if you need a managed service, a non-Java client, or a stable 4.x release, since the newest published artifact is a release candidate. Before committing, verify three things: which artifact version you resolve, whether your JVM is Java 20 or newer with the Vector module enabled for the optimized providers, and whether the native library builds on your distribution with g++ 11 or later.

## FAQ

### What is JVector?

JVector is an embedded, graph-based approximate nearest neighbor search engine written in Java and licensed under Apache-2.0. The README describes it as merging the DiskANN and HNSW family trees: it takes the hierarchical structure from HNSW and uses Vamana, the algorithm behind DiskANN, within each layer.

### What is JVector used for?

It performs approximate nearest neighbor search over vectors, trading an exact answer in linear time for an approximate answer in logarithmic time. The README argues graph indexes fit general-purpose use because they can be constructed and updated incrementally, unlike partition-based indexes that need a static dataset defined up front.

### How do I install JVector?

The README does not give a Maven coordinate, so you build from the repository. Initialize the Google Highway submodule with git submodule update --init, then build the native library by running bash build_native_lib.sh from jvector-native/src/main/native, which needs Meson, Ninja and g++ 11 or later.

### Does JVector require a newer Java version?

The project is structured to be built with JDK 20 or newer, and the multirelease jar is intended to be usable as a dependency from Java 11 code. When it runs on a Java 20 or newer JVM with the Vector module enabled, the README states that optimized vector providers are used.

### What is the difference between the two passes in JVector search?

The first pass uses lossily compressed vectors held in memory, either product quantization with optional anisotropic weighting, binary quantization, or fused PQ with codebooks written inline with the adjacency list. The second pass reads a more accurate representation from disk, either full-resolution float32 vectors or NVQ.

## Sources

- [datastax/jvector on GitHub](https://github.com/datastax/jvector)
- [Issues](https://github.com/datastax/jvector/issues)
- [License: Apache-2.0](https://github.com/datastax/jvector/blob/main/LICENSE)
- [README](https://github.com/datastax/jvector/blob/main/README.md)
- [Releases](https://github.com/datastax/jvector/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/datastax-jvector
