# vectorlite: HNSW inside SQLite, with an index you have to save yourself

> vectorlite loads an hnswlib-backed approximate index into SQLite as a virtual table, which puts vector search within reach of any language that has a SQLite driver, and the design decision that matters most is that the index is memory-only until you write it out.

**1yefuwang1/vectorlite** — Fast, SQL powered, in-process vector search for any language with an SQLite driver

- Repository: https://github.com/1yefuwang1/vectorlite
- Website: https://1yefuwang1.github.io/vectorlite/
- Stars: 361 · Forks: 11
- Language: C++
- License: Apache-2.0
- Published: 2026-09-15 · Updated: 2026-09-15 · Language: en
- Canonical page: https://hysenlabs.com/projects/1yefuwang1-vectorlite

## The index lives in memory and dies with the connection

The single most important operational fact about vectorlite is that the index is never on disk by default. It is always held in memory, and the persistence is an explicit operation you write as an ordinary SQL insert into the virtual table itself:

```sql
-- Save the current in-memory index to a file (overwrites if it exists).
insert into {table_name}(operation, path) values ('save', '/path/to/index.bin');
-- Load a saved index into a freshly created table. Loading replaces the table's
-- current in-memory index; on any error the existing index is left unchanged.
insert into {table_name}(operation, path) values ('load', '/path/to/index.bin');
```

The lifetime is per database connection. The index survives schema changes such as VACUUM, ALTER TABLE, or DDL arriving from another connection, for as long as that connection lives, and it is gone when the connection closes unless you saved it. The same mechanism costs you two column names: `operation`, `path` and `distance` are reserved and cannot be used as the name of your vector column.

## Loading a saved file has three compatibility rules

Restoring is not symmetric with creating. On load, the vector dimension and the element type, for example `float32`, must match what is in the file, so an index built for 768 dimensions will not load into a table declared for 1536. The distance type is the exception and may differ from the one used to build the file, which is what lets you query a cosine index with squared L2 without rebuilding. And `max_elements` is allowed to be larger than the saved index, which is how a table grows after a load instead of being pinned to the capacity it happened to have. Two further details are worth planning around. The save operation overwrites an existing file without asking, and a load that hits any error leaves the current in-memory index untouched rather than half-replaced. Index files written by hnswlib itself can also be loaded directly, so an index built outside SQLite does not have to be exported through this extension first.

## max_elements has no default, so sizing the table is yours

Creating the virtual table splits its arguments into required and optional, and the split tells you what you must decide up front. Required are the table name, the vector column name, the dimension, and `max_elements`, which has no default at all: you state the capacity of the index when you create it, and there is no documented resize path in the visible documentation. Everything else carries a default: `distance_type` is l2, `ef_construction` is 200, `M` is 16, `random_seed` is 100, and `allow_replace_deleted` defaults to true. Because the HNSW parameters are exposed rather than fixed, tuning is a matter of editing the create statement, with a worked example for parameter tuning sitting in the examples directory. The random seed being configurable is worth noticing too, since it makes an index build reproducible, which is the property you want when a benchmark result has to be compared across runs.

## The SIMD claim is measured on one Intel CPU with AVX2

The distance computation is the performance story, and it uses Google's SIMD library rather than relying on compiler auto-vectorization, so the same binary can pick the best instruction set at runtime. The comparison given is specific rather than general: on the author's own machine, an i5-12600KF with AVX2 support, the implementation is reported as 1.5 to 3 times faster than hnswlib's when vectors have a dimension of 256 or more. That is one desktop CPU on one microarchitecture, and the threshold matters, since below dimension 256 there is less work to amortize the setup. The runtime answer is available without benchmarking, because `vectorlite_info()` prints the version and the best SIMD target Highway chose on the machine you are actually running on. The distance types are the three hnswlib provides, l2 as squared L2, cosine, and inner product, and the author attaches his own caveat against the last of them.

## rowid is mandatory because the vector is a pointer, not a record

The data model treats a vector as a pointer to metadata stored somewhere else, and rowid is what makes the link. Inserting into a vectorlite table requires you to supply rowid yourself, with the reasoning stated plainly: auto-generating one does not make sense, because there is nothing for the generated value to point at. Everything else behaves like a normal SQLite table for insert, update and delete, and a normal table is the right mental model for the metadata side. Two capabilities are gated on the SQLite version. Predicate pushdown, which pushes a metadata rowid filter down into the vector scan instead of filtering after it, requires SQLite 3.38 or newer, and update and delete statements that filter on rowid require the same floor. Below that version the vector operations still work, but the filtered paths do not.

## Other languages extract the binary out of a Python wheel

vectorlite is a loadable SQLite extension, so the distribution problem is solved by shipping the shared library for each platform and letting every language reach it. The pre-built targets are Windows x64, Linux x64, macOS x64 and macOS arm64, and the two obvious channels are a Python wheel and an npm package:

```shell
# For python
pip install vectorlite-py
# for nodejs
npm i vectorlite
```

For anything else, the instruction is to take `vectorlite.[so|dll|dylib]` out of the wheel for your platform, on the observation that a wheel file is a zip archive, and load it from the sqlite CLI:

```sql
-- Load vectorlite
.load path/to/vectorlite.[so|dll|dylib]
-- shows vectorlite version and build info.
select vectorlite_info();
```

The build configuration explains why a pre-built artifact is needed at all: the prebuilt CPython interpreters used for multi-platform wheels, on manylinux, macOS and Windows, ship a `sqlite3` module compiled without loadable extension support, so the library cannot be reached through the standard library's extension loading path there.

## The Python metadata demands 3.14 and claims up to 3.13

The packaging is scikit-build-core with Ninja as the generator and vcpkg supplying the toolchain file, with CMake 3.22 as the floor and Release as the build type. The Python distribution is named `vectorlite_py`, carries its own version of 0.3.0, and its wheel payload points at a path inside the bindings tree rather than a conventional package directory, with the wheel tagged for a generic py3 platform. The contradiction worth catching is in the interpreter support. `requires-python` is set to 3.14 or newer, while the classifier list enumerates 3.9 through 3.13 and stops there, so the declared floor is above the highest version the package advertises. A `setup.py` remains in the tree purely as a stub, delegating to the PEP 517 backend and noting that direct `python setup.py` invocations are deprecated.

## Two releases from 2024, a package at 0.3.0, a branch pushed in 2026

Version numbers here run on three separate tracks. The GitHub releases are v0.1.0 on 8 July 2024 and v0.2.0 on 19 August 2024, the Python package declares 0.3.0, and the branch itself was last pushed on 13 September 2026, so most of the recent work has never been tagged. The documentation matches that posture by warning that the project is in beta and that breaking changes may follow, and it says so in two separate places rather than once. The highlights list has its own small defect, renumbering itself partway down, which suggests a list that was appended to rather than rewritten. The tree is arranged like a native project that happens to ship Python: CMakePresets.json, vcpkg.json with a bootstrap script, separate debug and release build scripts, a wheel extraction script, a formatting script, clang-format and clangd configuration, and both `doc/` and `docs/` directories, alongside bindings, benchmarks, examples and scripts.

## Conclusion

vectorlite fits someone who already keeps data in SQLite and wants approximate vector search without adding a server or a second storage engine. It fits badly if you expect the index to behave like a normal SQLite index, because it is held in memory per connection and is lost when the connection closes unless you save it, and it fits badly for exact nearest neighbour work, where the honest answer is that the virtual table trades exactness for speed. Three things to check. Loading a saved index requires the vector dimension and element type to match the file, with distance type free to differ, so a dimension mismatch is a load that fails. Predicate pushdown and rowid-filtered updates and deletes need SQLite 3.38 or newer. And the Python metadata contradicts itself, declaring a requires-python floor of 3.14 while the classifiers stop at 3.13. The project calls itself beta and warns twice about breaking changes, the newest GitHub release is v0.2.0 from 19 August 2024 while the Python package sits at 0.3.0, and the branch was last pushed on 13 September 2026. Apache-2.0 licensed.

## FAQ

### What is vectorlite?

A runtime-loadable SQLite extension that adds approximate nearest neighbour vector search backed by hnswlib, exposed through SQL, so it works from any language that has a SQLite driver. It is pre-compiled for Windows x64, Linux x64, macOS x64 and macOS arm64, and is distributed as a Python wheel and an npm package.

### How do I install vectorlite for a language other than Python?

Take the `vectorlite.[so|dll|dylib]` file for your platform out of the Python wheel, since a wheel is a zip archive, and load it from the sqlite CLI with `.load`. Python and Node users can skip that step with `pip install vectorlite-py` and `npm i vectorlite`.

### Is a vectorlite index saved automatically?

No. The index is held in memory per database connection and is lost when the connection closes unless you write it out yourself. Persistence is an insert into the table using the reserved operation and path columns, with a save value to write a file and a load value to read one back into a freshly created table.

### Which SQLite version does vectorlite need?

Two features require 3.38 or newer: predicate pushdown for metadata filtering, and update or delete statements that filter on rowid. Nothing in the free-standing functions or the virtual table itself is given a version floor in the documentation.

### How accurate is vectorlite's search?

The virtual table holds an approximate index rather than scanning every row, so it is not fully accurate by design. Exact results are still reachable by calling the vector_distance function against a normal SQLite table and ordering by it, which returns fully accurate results at brute force cost.

## Sources

- [1yefuwang1/vectorlite on GitHub](https://github.com/1yefuwang1/vectorlite)
- [License: Apache-2.0](https://github.com/1yefuwang1/vectorlite/blob/main/LICENSE)
- [Project website](https://1yefuwang1.github.io/vectorlite/)
- [README](https://github.com/1yefuwang1/vectorlite/blob/main/README.md)
- [Releases](https://github.com/1yefuwang1/vectorlite/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/1yefuwang1-vectorlite
