Open-source project
tinysearch/tinysearch avatar
tinysearch/tinysearch

tinysearch: a Rust and Wasm full-text search engine for static sites

🔍 Tiny, full-text search engine for static websites built with Rust and Wasm

2,975 stars94 forksRustApache-2.0

At a glance

What is it?
tinysearch compiles a full-text index into a WebAssembly module you ship alongside a static site. It is small, it has no runtime JavaScript dependencies, and it is not a drop-in replacement for lunr.js.
Who is it for?
Adopt tinysearch if your site is static, small to medium sized, and you want search that ships as one Wasm file with no JavaScript runtime. Do not adopt it if you need fuzzy matching, ranked relevance scoring, or an index that grows past a few thousand articles, because the README recommends it only for small to medium sites at roughly 2 kB uncompressed per article.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 18 days ago.
What is it written in?
Mainly Rust, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What tinysearch replaces on a static site

A static site generator produces HTML files and nothing else. There is no server to run a query against, so search has to happen in the browser, which means the entire index has to be downloaded by the visitor. The usual answer is lunr.js or elasticlunr, both of which the tinysearch README describes as too heavy for smaller websites because they load a lot of JavaScript. tinysearch takes the opposite approach: the index and the search logic are compiled ahead of time into a WebAssembly module, and the page loads that module instead of a JavaScript search library. The project is a Rust and WebAssembly port of the Python code from the article "Writing a full-text search engine using Bloom filters". It is aimed at people running Jekyll, Hugo, Zola, Cobalt, or Pelican who want a search box without adding a JavaScript dependency tree. The README gives a concrete size for one real site: the endler.dev index with 73 posts produces an optimized Wasm payload of 179 kB, 83 kB gzipped and 72 kB Brotli-compressed.

Exact posting lists versus the xor8 filter

The mechanism depends on which indexer you pick, and this is the part most people get wrong. By default, tinysearch stores one sorted vocabulary with exact posting lists that map each word to the articles containing it. Document IDs are delta- and varint-encoded, which keeps the index compact and compressible. Exact and prefix searches use the same index, and the README states that these do not introduce probabilistic false positives. That is the honest default: you get correct results, and you pay for them in index size. The optional xor8 indexer creates smaller probabilistic per-article filters instead. Smaller, but with a catch: it supports title prefixes, and body and metadata searches require complete words. So a user typing a partial word that appears only in the body gets nothing with xor8, where the exact indexer would have matched once the term reached three characters. The README notes that previously generated Xor-filter indexes use this same path, which matters if you are upgrading an existing site rather than starting fresh. There is no single right answer here. If your content is mostly titles and short posts, xor8 buys you size. If people search inside long articles, the default indexer is the one that will not silently miss matches.

Installing tinysearch and building a first index

tinysearch installs from crates.io as a binary. The README gives the command directly:

bash
cargo install tinysearch

If you want the Wasm output compressed further, install binaryen so that wasm-opt is available. On macOS the README suggests Homebrew, and binaryen is also available from the release page or your OS package manager:

bash
brew install binaryen

You need a JSON file describing the content to index. The repository ships one at fixtures/index.json that you can read before writing your own. The body field is optional, and the README notes you can skip it to index post titles only. Once the JSON exists, generate the Wasm search engine. The development form includes demo files, and the release form does not:

bash
tinysearch -m wasm -p wasm_output fixtures/index.json
tinysearch --release -m wasm -p wasm_output fixtures/index.json

Add -o to the release command to run wasm-opt, which requires binaryen to be installed. The README states that this creates a dependency-free Wasm module using vanilla cargo build rather than wasm-pack. To see it working before wiring it into a site, the repository provides a demo target that generates the Wasm files and starts a local server on port 8000:

bash
make demo

The README says to open http://localhost:8000/demo/ to try it. If you would rather not install Rust at all, the repository includes a Dockerfile and a Makefile target that builds an image tagged tinysearch/cli and runs it against the fixtures directory, writing output into docker_output/.

Choosing fields with tinysearch.toml

Which fields get indexed is not hardcoded. A tinysearch.toml file placed in the same directory as the JSON index controls the schema, with three keys: indexed_fields, metadata_fields, and url_field. The default, used when no configuration file is found, indexes title and body and takes the URL from a field named url. The repository ships an examples/tinysearch.toml and three worked configurations in the README: an e-commerce site indexing title, description, category, and tags while storing price, image_url, brand, and availability as metadata; a blog indexing title, body, and excerpt with author, publish_date, tags, and featured_image as metadata; and a documentation site indexing title, content, section, and keywords with version, last_updated, and contributor as metadata. The distinction between the two lists is the thing to internalize. indexed_fields are searchable; metadata_fields are returned with results and are not matched against the query. Putting a field in the wrong list is a silent failure, because nothing errors. Your search simply will not find the term, or your result cards will be missing data. The url_field key is also per-site: the e-commerce example uses product_url and the blog example uses permalink, so the JSON you generate has to agree with the config.

Where tinysearch stops working

The README is unusually direct about the ceiling. Because all search indices for all articles are bundled into one static binary, the project recommends using it only for small- to medium-size websites, and gives a working figure of around 2 kB uncompressed per article, roughly 1 kB compressed. That number is the whole argument. A 500-article site is on the order of a megabyte uncompressed before any compression, and every visitor downloads it whether or not they open the search box. There is no server-side filtering, no pagination of the index, and no lazy loading of shards. The other constraints are functional rather than architectural. With the default exact indexer, prefix matching starts once a query term reaches three characters, so shorter terms require an exact match. The xor8 indexer only supports prefixes in titles. And the search itself is not a relevance-ranked engine: the design is posting lists and filters, not scoring, so do not expect the ranking behavior of a system that computes term weights. If your users expect typo tolerance or fuzzy matching, tinysearch is the wrong tool, and so is any client-side index that has to fit in a page load.

How tinysearch differs from lunr.js

The README positions tinysearch as an alternative to lunr.js and elasticlunr, so that is the comparison worth making. lunr.js is a JavaScript library: you ship a JavaScript index and a JavaScript query engine, and the browser parses and executes both. tinysearch compiles the index and the query logic to WebAssembly ahead of time, and the README states the output module has no dependencies and is built with vanilla cargo build rather than wasm-pack. The practical difference is what the browser has to do at load time. A Wasm module is decoded and instantiated; a JavaScript search library has to be parsed, and its index deserialized. The other difference is the matching model. lunr.js builds a scored, ranked index with stemming and field boosting. tinysearch builds posting lists and, optionally, probabilistic filters, and it does not rank. So this is not a case of one being better. If you need relevance ordering or stemming, lunr.js does something tinysearch does not attempt. If you need the smallest possible payload and exact or prefix matching is enough, tinysearch is built for that trade.

Licence and upgrade cost

The repository carries two licence files, LICENSE-APACHE and LICENSE-MIT, and Cargo.toml declares the package as "Apache-2.0 OR MIT". That is a dual licence, and the choice between them is yours; the crate metadata and the repository contents agree on this, which is the thing to check before you rely on it. Nothing here is legal advice, and if you are redistributing the compiled Wasm module inside a commercial product, read both files rather than this paragraph. On upgrades, the version history shows v0.10.0 in September 2025, then v0.11.0 in August 2026 and v0.11.1 in September 2026. The last push to the repository was on 2026-09-14. The crate requires Rust 1.85 and edition 2024, so the toolchain floor is real if your CI pins an older Rust. Two upgrade hazards are visible: the README states that previously generated Xor-filter indexes use the same xor8 path, so an index built under an older release may need regenerating rather than reusing; and the library API is explicitly marked experimental with the note that it may change. If you use tinysearch only as a CLI to produce a Wasm file, the API caveat does not apply to you.

Editorial conclusion

Adopt tinysearch if your site is static, small to medium sized, and you want search that ships as one Wasm file with no JavaScript runtime. Do not adopt it if you need fuzzy matching, ranked relevance scoring, or an index that grows past a few thousand articles, because the README recommends it only for small to medium sites at roughly 2 kB uncompressed per article. Before committing, verify which indexer you need: the default exact indexer only starts prefix matching at three characters, while the xor8 indexer supports prefixes in titles only. Build the index from your own JSON with tinysearch --release -m wasm -p wasm_output your-index.json and check the resulting payload size.

Frequently asked questions

How do I install tinysearch?

Install it from crates.io with cargo install tinysearch. To optimize the WebAssembly output you can additionally install binaryen, which provides wasm-opt, and the README suggests brew install binaryen on macOS.

Does tinysearch need a server to run search?

No. tinysearch compiles the index and the search logic into a WebAssembly module that runs in the browser, which is why it is described as a search engine for static websites.

What is the difference between the default indexer and xor8 in tinysearch?

The default indexer stores one sorted vocabulary with exact posting lists and supports both exact and prefix searches without probabilistic false positives. The optional xor8 indexer creates smaller probabilistic per-article filters but only supports prefixes in titles, and body and metadata searches require complete words.

How large is the tinysearch WebAssembly payload?

The README gives one measured figure: the endler.dev index with 73 posts produces an optimized Wasm payload of 179 kB, 83 kB gzipped and 72 kB Brotli-compressed. The same section estimates around 2 kB uncompressed per article, about 1 kB compressed.

Official sources

  1. License: Apache-2.0
  2. Project website
  3. README
  4. Releases
  5. tinysearch/tinysearch on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/tinysearch-tinysearch.svg)](https://hysenlabs.com/projects/tinysearch-tinysearch)