hnswlib: the header-only nearest neighbor search that outgrew its parent
Header-only C++/python library for fast approximate nearest neighbors
At a glance
- What is it?
- hnswlib is an Apache-2.0 header-only C++ library with Python bindings for fast approximate nearest neighbor search, implementing HNSW with full support for incremental construction, element updates and deletions. It needs nothing beyond C++11, has significantly less memory footprint and faster build time than the full nmslib implementation, and its current release candidate adds a no-exceptions C++ API with Status returns.
- Who is it for?
- Use hnswlib when you need approximate nearest neighbor search embedded in a C++ project with minimal dependencies, or from Python with an index that survives pickling and accepts updates after construction. Use the full nmslib library when you need spaces beyond L2, inner product and cosine, or FAISS when GPU acceleration and its broader index zoo justify the weight.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 15 days ago.
- What is it written in?
- Mainly C++, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Header-only, and faster than the family flagship
hnswlib is a header-only C++ HNSW implementation with Python bindings, insertions and updates, and its headline properties are stated as engineering facts rather than aspirations, no dependencies other than C++ 11, significantly less memory footprint and faster build time compared to the current nmslib implementation. That last comparison is the interesting one, this is the lightweight extract that beat its own parent project at the parent's game, trading nmslib's breadth of spaces for speed and embeddability. Interfaces exist for C++ and Python directly, with external support for Java and R through the rcpphnsw project, and the C++ side accepts custom user-defined distances for spaces the maintainers never imagined. Construction is fully incremental, elements can be updated after insertion, deletions are supported by marking, and the Python index is picklable, four properties that many ANN libraries still treat as research questions.
Three distances, one of them not a metric
The Python bindings support exactly three spaces, squared L2 as 'l2', inner product as 'ip', and cosine distance as 'cosine', each with its equation printed in the documentation. The inner product entry carries a mathematical warning worth reading slowly, it is not an actual metric, and an element can be closer to some other element than to itself, a property the documentation immediately turns into an optimization, if you remove all elements that are not the closest to themselves from the index, you get a speedup, a pruning trick only a non-metric space permits. For anything beyond the three, the README points at the full nmslib library rather than pretending the header-only build covers every space, an honest boundary that explains when to graduate from this library to its parent.
A Python API of about six verbs
The Python surface is small enough to hold in your head. An index is created with hnswlib.Index(space, dim) and initialized with init_index, whose parameters each carry documented meaning, max_elements bounding capacity, M setting the maximum number of outgoing graph connections, ef_construction trading build time against accuracy, a random seed, and allow_replace_deleted toggling the deletion-replacement machinery. Insertion is add_items, taking a numpy array of vectors, optional integer labels, a thread count with -1 for default, and a replace_deleted flag, with two documented behaviors that matter in production, re-adding an existing label updates that element, slower than insertion but more memory- and query-efficient, and add_items is thread-safe with other add_items calls but not with knn_query. Deletion is mark_deleted and its undo unmark_deleted, capacity changes through resize_index, and set_ef tunes the query-time accuracy and speed trade-off, with ALGO_PARAMS.md explaining the theory behind each knob.
Deletion as marking, replacement as recycling
The deletion model deserves its own reading because it shapes long-running systems. mark_deleted does not remove an element, it marks it so search results omit it, and calling it on an already-deleted label throws, a deliberate guard against bookkeeping drift. The complementary half is replacement, when init_index was called with allow_replace_deleted, add_items can be told replace_deleted to fill vacated slots with new elements, which the documentation motivates as memory control, an index that only ever grows is an index that eventually does not fit, and recycling deleted slots keeps the structure at capacity without a rebuild. Between mark_deleted, unmark_deleted and the replacement path, an index can absorb churn, items going away and new ones arriving, which is the difference between a search library for static datasets and one that can sit behind a live service.
Version 0.10: a no-exceptions C++ API
The current release candidate, 0.10.0rc2, is dominated by one contribution, an optional no-exceptions C++ API in which the *NoExceptions methods return Status and StatusOr values, allowing the headers to compile with -fno-exceptions or -DHNSWLIB_ENABLE_EXCEPTIONS=OFF while the throwing methods remain the default. This matters because large codebases in games, embedded systems and some infrastructures ban exceptions outright, and libraries that only speak exceptions are unusable there. The same release hardens the edges, stream loadIndexNoExceptions now fails closed on an unopened or failed input without clearing a live index, saveIndex and loadIndex work over std::ostream and std::istream with a getInternalIdByLabel accessor alongside, CMake no longer wipes the caller's CMAKE_CXX_FLAGS, and CI now covers exceptions on and off across Clang, GCC and MSVC under ASAN and UBSAN, which caught a real Clang UBSan misaligned label store. The version string follows PEP 440 with an explicit note not to upload the candidate as a final 0.10.0.
The 0.7 to 0.9 arc: filters, races and recall fixes
Reading the older release notes in order sketches the library's maturation. Version 0.7.0 added filtering support with a Python interface the notes candidly describe as limited by the GIL, element replacement for size control, and, critically, fixes for data races and deadlocks in concurrent updates and insertions backed by a stress test for multithreaded operation, the difference between theoretically thread-safe and proven under contention. Version 0.8.0 brought multi-vector document search and epsilon search, C++ only for now, and disabled statistic aggregation by default after the notes observed nobody seemed to be using it, a speedup earned by deleting a feature. Version 0.9.0 is a correctness release, fixing incorrect bruteforce results with a filter, a missing normalization check in BFIndex, and throwing an exception when fewer than k elements are available instead of returning quietly wrong answers, the kind of fixes that build the trust an approximate library runs on.
Build, test, measure recall
The build story matches the header-only claim, C++ projects include the headers and need C++ 11 or newer, with the Python bindings built through setuptools over pybind11 and numpy, the setup script preferring -std=c++14 when the compiler offers it. Tests run through unittest discovery over the bindings test suite, with a test_all_build_types.sh script exercising the configuration matrix, and committed examples in both examples/cpp and examples/python show first usage without requiring documentation archaeology. Two dedicated documents serve the numerically serious, ALGO_PARAMS.md explaining the algorithm parameters, and TESTING_RECALL.md, whose very existence states the library's position that recall measurement is the user's job and deserves a guide. Releases run v0.9.0 in March 2026 and the September 0.10.0 candidates, with the last push on 2026-09-15, and against FAISS, the widely used alternative, the trade is scope, FAISS brings GPU indexes and a zoo of methods while hnswlib brings one algorithm, header-only, with nothing to link.
Editorial conclusion
Use hnswlib when you need approximate nearest neighbor search embedded in a C++ project with minimal dependencies, or from Python with an index that survives pickling and accepts updates after construction. Use the full nmslib library when you need spaces beyond L2, inner product and cosine, or FAISS when GPU acceleration and its broader index zoo justify the weight. Verify first which distance your embedding model assumes before choosing 'ip' over 'cosine', read ALGO_PARAMS.md before tuning M and ef_construction rather than copying defaults, and pin a stable release rather than the 0.10.0 release candidates until that line finalizes.
Frequently asked questions
What is hnswlib?
hnswlib is an Apache-2.0 header-only C++ library with Python bindings for fast approximate nearest neighbor search, implementing HNSW. It supports incremental construction, element updates and deletions, and has significantly less memory footprint and faster build time than the full nmslib implementation.
How do you install hnswlib?
For C++, copy the headers into your project, since the library is header-only with no dependencies beyond C++11. For Python, build the bindings from the repository, which requires numpy and pybind11 as build dependencies per the pyproject configuration.
Which distances does hnswlib support?
The Python bindings support squared L2 ('l2'), inner product ('ip') and cosine distance ('cosine'). Inner product is not a true metric, an element can be closer to another element than to itself, which enables a pruning optimization. Other spaces require the full nmslib library.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/nmslib-hnswlib)