valkey-search: vector and secondary indexing as a Valkey module
valkey-search is a C++ module which extends valkey with vector search and secondary indexing capabilities. It enables users to index and query data stored in Valkey using complex queries with filters while maintaining high performance and scalability.
At a glance
- What is it?
- valkey-search adds FT.CREATE, FT.SEARCH and FT.AGGREGATE to Valkey for vector, tag, numeric and full-text queries. It is a C++ module you build yourself, and the build toolchain is the first real hurdle.
- Who is it for?
- Adopt valkey-search if you already run Valkey and want vector or full-text queries without operating a separate search cluster, and if you can build the module from source with GCC 12 or Clang 16. Do not adopt it if you need a packaged binary, or if your data does not fit in memory, since the README states vectors are stored in-memory.
- Can I use it commercially?
- Yes. BSD-3-Clause is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 1 day ago.
- What is it written in?
- Mainly C++, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 28, 2026, and from our analysis. They are not legal advice.
DEEP OPEN-SOURCE ANALYSIS
What valkey-search adds to a Valkey deployment
Valkey on its own is a key-value store. You can put a vector in a hash field, but you cannot ask it for the ten nearest neighbours, and you cannot ask for all documents where category is "shoes" and price is under 100. valkey-search is a Valkey module that adds both: secondary indexes over Numeric, Tag and Text fields, plus vector search using Approximate Nearest Neighbor (ANN) with HNSW or exact matching with K-Nearest Neighbors (KNN).
The intended audience is teams running AI, search, analytics or recommendation workloads that already keep their data in Valkey. The README frames the value as avoiding a second system: index the data where it lives rather than exporting it to a dedicated search engine. Documents can be indexed from either Valkey Hash or Valkey-JSON data types, and the README points to separate hashes.md and valkey-json.md documents for each.
The command surface is deliberately small. Six commands are listed: FT.CREATE, FT.DROPINDEX, FT.INFO, FT._LIST, FT.SEARCH and FT.AGGREGATE. Anyone who has used RediSearch will recognise the naming, which is the point: the module is a search layer bolted onto the Valkey command space, not a new query language.
Query planning and the pre-filter versus post-filter trade-off
The interesting part of the design is how hybrid queries are handled. A hybrid query combines vector similarity with a filter on an indexed field. The README describes two textbook approaches and their failure modes. Pre-filtering filters the dataset first, then runs an exact similarity search; it works well when the filtered set is small but becomes costly as the set grows. Post-filtering runs the similarity search first and filters afterwards; it suits large filter-qualified sets but, in the README's words, may lead to "empty or lower than expected amount of results" because the top-k neighbours can all be filtered out.
valkey-search does not force you to choose. It uses a query planner that picks between pre-filtering and inline-filtering, where filtering happens during the similarity search itself. That is the right shape for the problem: the cost of the two strategies depends on selectivity, which the client usually does not know in advance, so pushing the decision into the engine is more useful than exposing a flag.
The README does not document how the planner estimates selectivity or what statistics it uses, so the quality of that choice is not something a reader can assess from the repository documentation alone. Treat the planner as the feature to test against your own data distribution rather than as a guarantee.
Building the module and running a first query
There is no published binary in the repository. You build from source. On Ubuntu or Debian the README lists the toolchain packages, and it is explicit that GCC 12 or higher, or Clang 16 or higher, is required.
sudo apt update
sudo apt install -y build-essential g++ cmake libgtest-dev ninja-build libssl-devIf your distribution ships an older compiler, the README gives the upgrade path to gcc-12 and g++-12 and registers them with update-alternatives. The build itself goes through a wrapper script around CMake.
./build.shRunning ./build.sh --run-tests executes the unit tests. The README also documents ./build.sh --run-integration-tests, which needs Python, a virtual environment and the Redis package repository key added to apt before it will run.
Once built, you load the module into the server with --loadmodule, pointing at the shared object. For JSON documents you load libjson.so as well.
valkey-server --loadmodule /path/to/libsearch.so --loadmodule /path/to/libjson.soAfter that, FT.CREATE defines an index over hash or JSON documents, and FT.SEARCH queries it. The README does not inline the FT.CREATE syntax; it points to QUICK_START.md for worked examples and to valkey.io/commands/#search for the full argument list, so read those before writing your first index definition.
Thread count, replica reads and where the scaling story stops
The README states that query processing and ingestion scale linearly with CPU cores in both Standalone and Cluster modes. Two knobs control this. By default the module matches worker threads to the number of CPU cores on the host, and you can override it at load time.
valkey-server "--loadmodule /path/to/libsearch.so --reader-threads 64 --writer-threads 64"The performance section attributes the throughput to in-memory vector storage plus three implementation choices: a threading model with lock-free execution on the read path, CPU cache efficiency, and SIMD vector processing. Those are claims from the project's own documentation, not measurements, and the README publishes no benchmark numbers, hardware configuration or recall-versus-latency curves to back them. The headline claims of single-digit millisecond latency, high QPS, billions of vectors and over 99% recall are stated without a reproduction recipe. If those numbers matter to your decision, you will have to generate them on your own hardware.
For horizontal query scaling, the README offers a second path: direct clients to read from replicas if replica lag is acceptable. That is a real trade-off, not a free win. Reading from a replica means a query can miss a document that was just written to the primary, which for a search index can surface as a result set that flickers between requests.
Rebuilding indexes on startup with skip-rdb-load
The most operationally interesting feature is skip index on RDB load. It is a non-modifiable configuration passed to the module at load time.
valkey-server --loadmodule /path/to/libsearch.so --skip-rdb-load yesWhen enabled, the index schema is read from the RDB file but the index contents are not; indexes start empty and are refilled by the backfill process. The README gives two reasons. First, index consistency: it is a recovery path from a corrupted or inconsistent vector index. Second, memory reclaim: if an index holds many deleted vectors, rebuilding at startup frees that memory.
The asymmetry is worth noting. Non-vector indexes (TAG, NUMERIC) keep working immediately, while vector indexes are unavailable until backfill finishes. So this flag trades startup latency on vector queries for consistency and reclaimed memory. The README does not say how long backfill takes, whether it blocks writes, or how to observe its progress, which makes it hard to plan a restart around. It also does not document any rollback procedure if the rebuild produces a worse index than the one on disk. Test the flag on a copy of production data before using it as a routine restart option.
Compared with Redis Search, and when to pick something else
The obvious alternative is RediSearch, the search module for Redis. The two share command names, which makes migration look cheaper than it is. The real difference is the substrate: valkey-search runs inside Valkey, so you inherit Valkey's licence and release cadence rather than Redis's, and the module is built from source in this repository rather than installed from a vendor distribution. If your organisation has already standardised on Redis and buys support, swapping the server underneath a search workload is a larger change than swapping a library.
A second alternative is a dedicated search engine such as OpenSearch or Elasticsearch, or a purpose-built vector database. Those give you a separate index with its own storage, its own scaling model and its own operational burden. valkey-search's argument against them is locality: the index lives next to the data, so there is no export pipeline and no second consistency problem. Its argument for them is that a dedicated engine does not require your entire corpus to sit in memory.
That memory requirement is the clearest wrong-tool case. The README states vectors are stored in-memory. Cluster mode horizontally scales the keyspace, but every node still holds its shard of vectors in RAM. If your corpus is large and your budget per gigabyte of RAM is the binding constraint, a disk-backed index is the better fit regardless of how good the query planner is. Similarly, if you need a managed service with a support contract, this repository does not offer one.
Maintenance, licensing and upgrade cost
The repository is not archived, and the last push was on 2026-07-15. Three releases are listed around that date: 1.1.1 and 1.0.3 on 2026-07-15, and 1.2.1 on 2026-07-07. The presence of 1.0.3, 1.1.1 and 1.2.1 as parallel patch lines suggests maintenance branches are kept alive alongside the newest minor version, which is a reasonable sign for teams that cannot jump minor versions quickly. The README does not state a support window for older lines, so how long 1.0.x keeps receiving fixes is not documented.
Licensing is BSD-3-Clause, per both the README and the LICENSE file at the repository root. That is a permissive licence, which matters here because you are linking a module into a server process rather than calling it over a network. The repository also carries a THIRD_PARTY_NOTICES file, which is where the licences of bundled dependencies will be recorded; if you redistribute a built libsearch.so, that file is the place to start. This is not legal advice, and the interaction between the module licence and the licences of whatever else you load into the same server is worth confirming with whoever handles compliance for your product.
Upgrade cost is dominated by the build. Because there is no packaged artifact in the repository, every upgrade means rebuilding against your compiler and re-running the integration tests, which the README notes need Python and the Redis apt repository configured. Budget for that on each release rather than treating it as a one-time setup.
Editorial conclusion
Adopt valkey-search if you already run Valkey and want vector or full-text queries without operating a separate search cluster, and if you can build the module from source with GCC 12 or Clang 16. Do not adopt it if you need a packaged binary, or if your data does not fit in memory, since the README states vectors are stored in-memory. Before committing, verify two things yourself: that your Valkey version is compatible, and whether you need the JSON module loaded alongside libsearch.so for FT.CREATE on JSON documents. Start with QUICK_START.md and the command reference at valkey.io/commands/#search, and check COMMANDS.md in the repository for the exact argument list of each FT command.
Frequently asked questions
How do I install the valkey-search module?
There is no packaged binary in the repository, so you build it from source with ./build.sh, which requires GCC 12 or higher, or Clang 16 or higher. You then start the server with valkey-server --loadmodule /path/to/libsearch.so, adding libjson.so if you need JSON support.
Does valkey-search work with Docker?
The README documents a VSCode Dev Containers setup for development, which builds the code inside a Docker container. It does not describe an official runtime image for deploying the module.
Which commands does valkey-search add to Valkey?
The README lists FT.CREATE, FT.DROPINDEX, FT.INFO, FT._LIST, FT.SEARCH and FT.AGGREGATE. Detailed argument descriptions are in the command reference at valkey.io/commands/#search.
Can I run valkey-search in Cluster mode?
Yes. The README states the module supports both Standalone and Cluster modes, and that Cluster mode provides horizontal scaling of the keyspace for large storage requirements.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/valkey-io-valkey-search)
Community notes