# ANN-Benchmarks: benchmarking approximate nearest neighbor libraries in Docker

> ANN-Benchmarks runs approximate nearest neighbor libraries over shared HDF5 datasets inside per-algorithm Docker containers and plots recall against throughput. The README states the project is no longer actively maintained.

**erikbern/ann-benchmarks** — Benchmarks of approximate nearest neighbor libraries in Python

- Repository: https://github.com/erikbern/ann-benchmarks
- Website: http://ann-benchmarks.com
- Stars: 5,735 · Forks: 903
- Language: Python
- License: MIT
- Published: 2026-09-22 · Updated: 2026-09-22 · Language: en
- Canonical page: https://hysenlabs.com/projects/erikbern-ann-benchmarks

## The problem ANN-Benchmarks was built to solve

Approximate nearest neighbor search sits behind recommendations, image similarity and retrieval-augmented generation. The methods are numerous and the claims are hard to compare: an HNSW graph, an inverted file with product quantization, a random projection forest and a locality-sensitive hash table all expose different knobs, and each library's own README tends to report the configuration that flatters it.

ANN-Benchmarks exists to remove that asymmetry. The README describes the goal plainly: doing fast searching of nearest neighbors in high dimensional spaces has few empirical attempts at comparing approaches in an objective way. The project answers that by fixing the dataset, fixing the evaluation metric, and running every algorithm inside its own Docker container so that build dependencies and runtime environments stay separate. The intended reader is someone choosing an index, or an author who wants a number that other people can reproduce.

The README opens with a status note that matters more than any benchmark table: ann-benchmarks is no longer actively maintained, and it points readers toward other benchmarks such as VIBE. The last push to the repository was on 2026-07-10, so the code is not frozen, but the maintainer's own statement is that the project is not being developed. Treat the published results as a historical record and the harness as a tool you can still run.

## How the harness runs an algorithm: datasets, Docker, and the recall-throughput curve

The repository layout makes the data flow visible. Datasets live in HDF5 format and are pre-generated; create_dataset.py is the entry point for producing new ones. Each algorithm gets a Docker container built from a Dockerfile under the algorithms directory, and install.py builds those images. run.py drives the actual measurement, with run_algorithm.py handling a single algorithm, and plot.py turns the recorded results into the recall versus queries-per-second curves the project is known for. create_website.py renders the static site from templates/.

The reason for the container boundary is dependency conflict. FAISS, NMSLIB, ScaNN, Annoy and the rest do not share a compatible set of native libraries, and several of them are not pip-installable in a clean way. Wrapping each in its own image means the harness can add an algorithm without touching the others. The cost is that every algorithm pays a Docker build, and the results depend on the image definition as much as on the library.

The evaluation itself is a sweep. For each algorithm the harness runs a series of parameter settings, records recall against the exact brute-force answer and measures throughput, then plots the frontier. That is why the output is a curve rather than a single number: an ANN index trades accuracy for speed, and a single point on the curve says almost nothing. The README also mentions a test suite on GitHub Actions that verifies function integrity, which is the project's guard against an algorithm silently returning wrong neighbours.

The requirements.txt pins the harness's own Python dependencies, including docker, h5py, numpy, pyyaml, matplotlib, jinja2, scikit-learn and psutil. Those pins are exact versions, which is what you want for a benchmark harness and also what makes it brittle to upgrade.

## Installing ANN-Benchmarks and running your first algorithm

The README does not reproduce a full install walkthrough, but the repository ships install.py, run.py, run_algorithm.py, plot.py, create_dataset.py and create_website.py at the top level, plus a requirements.txt and a results/ directory. What follows describes what each of those scripts is for; the exact flags each one accepts are defined by the scripts themselves, and the README does not list them.

Install the harness's Python dependencies from the pinned file:

```bash
pip install -r requirements.txt
```

Build the Docker images for the algorithms you want to measure. install.py is the repository's install entry point:

```bash
python install.py
```

Drive the measurement with run.py, which is the top-level driver, or with run_algorithm.py when you want a single algorithm:

```bash
python run.py
```

Render the recall against queries-per-second chart from the recorded results with plot.py:

```bash
python plot.py
```

What you should see is a results file under results/ for the run and a chart plotting recall against queries per second for each parameter setting. Datasets are HDF5 files; if the one you want is not present locally, the harness fetches the pre-generated copy, so the first run of a new dataset takes longer than later ones. The README does not document a rollback path for a failed Docker build, so if an algorithm image fails to build, the practical move is to inspect that algorithm's Dockerfile rather than to expect the harness to recover.

## Where the harness gets in the way

The strongest limitation is stated by the project itself. The README says ann-benchmarks is no longer actively maintained and recommends submitting work to different benchmarks. A harness that is not being developed will drift from the libraries it measures: a new index version, a changed default parameter, or a dropped dependency can leave an algorithm's Dockerfile stale without anyone noticing. The test suite helps, but it verifies the harness, not the freshness of each algorithm image.

The second limitation is the dataset. ANN-Benchmarks measures on its own pre-generated HDF5 datasets. Those are useful for cross-library comparison and useless for predicting behaviour on your corpus. Vector dimensionality, cluster structure, filter selectivity and update rate all change which index wins, and none of them are properties of the benchmark datasets. A library that tops the chart on a static 100-dimensional collection can behave very differently under heavy deletes.

The third is the container boundary. Docker isolation makes the comparison fair across libraries, and it also means you are benchmarking the image, not the library as you would deploy it. Memory limits, thread counts and CPU pinning inside the container are set by the Dockerfile. If your production deployment differs, the throughput number is not transferable.

Finally, the harness is single-machine and Python-driven. It is the wrong tool for measuring a distributed vector service under concurrent load, and it is the wrong tool for latency percentiles under a mixed query workload. It answers one question well: given these datasets and this metric, which algorithm gives the better recall at a given speed.

## What to compare it against, and why the difference matters

The README itself names VIBE, the Vector Index Benchmark for Embeddings, as an alternative to submit work to. The difference in approach is the dataset philosophy. ANN-Benchmarks ships pre-generated HDF5 datasets and asks each algorithm to run against them inside a per-algorithm Docker container. VIBE is described in the README as a place to submit work rather than as a drop-in replacement harness, which implies a different contribution model: you bring an index and it gets evaluated in a shared, maintained setting instead of you maintaining a Dockerfile in a repository that is no longer developed.

There is also the option of not using a general benchmark at all. If your question is whether pgvector, Qdrant or Milvus will serve your application, the honest answer comes from replaying a sample of your own queries against a copy of your own data, with your filters applied. ANN-Benchmarks can tell you which family of index is worth trying first; it cannot tell you what your p99 latency will be. The README's own list of evaluated systems spans embedded libraries like Annoy and hnswlib, database extensions like pgvector and Elasticsearch, and standalone services like Milvus and Vespa, and those categories are not interchangeable in production even when they appear on the same chart.

## Maintenance, licence and the cost of keeping a fork alive

The repository is not archived, and the last push was on 2026-07-10, but the README's status section states the project is no longer actively maintained. That combination is common for research harnesses: the code still runs, the results are still published, and nobody is triaging new algorithms. Practically, if you add an algorithm today you own that Dockerfile indefinitely.

The upgrade surface is small and pinned. requirements.txt fixes exact versions of docker, h5py, numpy, matplotlib and the rest, so a fresh environment is reproducible. It also means that moving to a newer numpy or h5py is a deliberate act that may break plotting or dataset loading, and there is no changelog in the repository to tell you which versions were validated together. pyproject.toml only configures black and ruff with a line length of 120, so there is no packaging metadata to rely on for version constraints.

The project is MIT licensed, which permits commercial use, modification and redistribution provided the copyright notice and permission notice are retained. That is a permissive licence and not legal advice; if you redistribute a modified harness inside a product, have counsel confirm the notice requirements. The licence does not cover the algorithms themselves, which carry their own terms, and several of those are not MIT. Checking each algorithm's licence before you ship it is your responsibility, not the harness's.

## Conclusion

Use ANN-Benchmarks if you need a reproducible, container-isolated comparison of ANN libraries on public datasets and you accept that the README declares it no longer actively maintained. Do not use it to pick a managed vector database for a specific production workload: the harness measures library behaviour on its own datasets, not your data, your filters or your latency budget. Before adopting it, check that the algorithm you care about has a Dockerfile under ann_benchmarks/algorithms, confirm install.py completes on your machine, and decide whether the current results/ directory or a fresh run is the basis for your decision.

## FAQ

### Is ANN-Benchmarks still maintained?

The README states that ann-benchmarks is no longer actively maintained and suggests submitting work to other benchmarks such as VIBE. The repository is not archived and the last push was on 2026-07-10, so the code remains available even though development has stopped.

### How do I install ANN-Benchmarks and run a benchmark?

Install the pinned Python dependencies from requirements.txt, then run install.py to build the Docker images for the algorithms. run.py drives the measurement and plot.py renders the recall versus queries-per-second chart from the recorded results.

### Which algorithms does ANN-Benchmarks evaluate?

The README lists Annoy, FLANN, scikit-learn's LSHForest, KDTree and BallTree, Weaviate, NMSLIB, hnswlib, FAISS, ScaNN, DiskANN, pgvector, Qdrant, Milvus, Elasticsearch, Vespa and several others. Each one is packaged in its own Docker container.

### What dataset format does ANN-Benchmarks use?

Datasets are pre-generated in HDF5 format, and create_dataset.py is the script for producing new ones. The harness loads a named dataset for a run, and the README describes the datasets as pre-generated rather than generated on demand.

## Sources

- [erikbern/ann-benchmarks on GitHub](https://github.com/erikbern/ann-benchmarks)
- [Issues](https://github.com/erikbern/ann-benchmarks/issues)
- [License: MIT](https://github.com/erikbern/ann-benchmarks/blob/main/LICENSE)
- [Project website](http://ann-benchmarks.com)
- [README](https://github.com/erikbern/ann-benchmarks/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/erikbern-ann-benchmarks
