Open-source project
toshi-search/Toshi avatar
toshi-search/Toshi

Toshi: A Rust Full-Text Search Engine Built on Tantivy

A full-text search engine in rust

4,256 stars136 forksRustMIT

At a glance

What is it?
Toshi is a Rust full-text search engine that wraps Tantivy in an HTTP API, aiming to occupy the same role relative to Elasticsearch that Tantivy occupies relative to Lucene. The README describes it as far from production ready, under continued development, and it avoids unsafe Rust throughout its own codebase.
Who is it for?
Toshi is worth evaluating for developers who want a Rust-native full-text search server with an HTTP API and who are comfortable running pre-production software. Its stable-Rust requirement and explicit avoidance of unsafe code in its own layer make it appealing for teams who want to audit or contribute to the codebase.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 94 days ago.
What is it written in?
Mainly Rust, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 25, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What Toshi Is and the Role It Fills

Tantivy is a Rust library for full-text search that takes its design cues from Lucene: inverted indexes, configurable tokenizers, and query DSL. Lucene is the search library at the core of Elasticsearch, but Lucene provides no HTTP server or document API. Elasticsearch wraps Lucene with a distributed HTTP layer and a JSON API.

Toshi does for Tantivy what Elasticsearch does for Lucene: it adds an HTTP API and a JSON document interface on top of the search library. The README states this relationship explicitly: Toshi strives to be to Elasticsearch what Tantivy is to Lucene.

The target user is a developer who needs full-text search without the JVM dependency that Elasticsearch carries and without the operational complexity of a multi-node cluster. Toshi compiles to a single binary and reads its configuration from a TOML file. The README notes it is far from production ready and that development continues at a measured pace.

The project avoids unsafe Rust. This is a deliberate design choice documented in the README: stable Rust was chosen for the guarantees and safety it provides. Underlying libraries (including Tantivy) may use unsafe internally, but the Toshi codebase itself does not.

How Toshi Uses Tantivy and Segments

Tantivy organizes an index into segments: chunks of indexed documents that are written to disk and later merged according to a merge policy. As documents accumulate, Toshi creates new segments and eventually merges smaller ones to maintain query performance. This segment lifecycle is the same approach Lucene uses.

Toshi exposes two merge policy options in its configuration. The log policy is the default, which uses logarithmic sizing to decide when segments should be merged. Three parameters tune it: min_merge_size, min_layer_size, and level_log_size. The nomerge option disables segment merging entirely, which reduces background I/O at the cost of growing segment counts over time.

The writer_memory setting controls how much memory (in bytes) Toshi allocates to the write buffer before flushing to disk as a new segment. The default is 200,000,000 bytes. The bulk_buffer_size setting controls the channel buffer for the JSON parsing threads during bulk ingest. Setting it to 0 makes the buffer unbounded.

The json_parsing_threads setting controls how many threads parse incoming document JSON during a bulk ingest operation. The default is 4. For high-throughput ingest, increasing this value uses more CPU to keep pace with incoming documents.

Building Toshi and Running a First Query

Building Toshi requires Rust and Cargo, available from rustup.rs. From the repository root:

bash
cargo build --release

Once built, run Toshi from the top level directory:

bash
./target/release/toshi

Toshi reads config/config.toml on startup. The default configuration looks like this:

toml
host = "127.0.0.1"
port = 8080
path = "data2/"
writer_memory = 200000000
log_level = "info"
json_parsing_threads = 4
bulk_buffer_size = 10000
auto_commit_duration = 10
experimental = false

Each setting has a specific role. The host is the hostname Toshi binds to on startup, and the port is the corresponding port. The path is where Toshi stores its data and indices on disk. The log_level controls the detail in log output. The auto_commit_duration controls how often an index automatically commits pending documents; setting it to 0 disables automatic commits, which means the caller must trigger commits manually after submitting documents.

Verify that Toshi is running:

bash
curl -X GET http://localhost:8080/

A successful startup returns the name and version as JSON. Once running, queries go to the index endpoint as HTTP POST requests with JSON bodies. The README notes that after the initial setup, the requests.http file in the repository root contains additional usage examples worth reviewing.

For a term query against a test_index:

bash
curl -X POST http://localhost:8080/test_index -H 'Content-Type: application/json' -d '{ "query": {"term": {"test_text": "document" } }, "limit": 10 }'

The limit parameter is optional and defaults to 10. Running the test suite requires cargo test.

Query Types: Term, Fuzzy, Phrase, Range, Regex, and Boolean

Toshi supports six query types. Term queries match documents that contain the exact term:

json
{ "query": {"term": {"test_text": "document" } }, "limit": 10 }

Fuzzy queries match terms within an edit distance and accept a transposition flag:

json
{ "query": {"fuzzy": {"test_text": {"value": "document", "distance": 0, "transposition": false } } }, "limit": 10 }

Phrase queries match an ordered sequence of terms in the same field:

json
{ "query": {"phrase": {"test_text": {"terms": ["test","document"] } } }, "limit": 10 }

Range queries match numeric fields with gte/lte/gt/lt bounds:

json
{ "query": {"range": { "test_i64": { "gte": 2012, "lte": 2015 } } }, "limit": 10 }

Regex queries match terms by pattern:

json
{ "query": {"regex": { "test_text": "d[ou]{1}c[k]?ument" } }, "limit": 10 }

Boolean queries combine must and must_not clauses from any of the other types:

json
{ "query": {"bool": {"must": [ { "term": { "test_text": "document" } } ], "must_not": [ {"range": {"test_i64": { "gt": 2017 } } } ] } }, "limit": 10 }

Distributed Mode and Current Production Readiness

Toshi includes experimental distributed mode settings in the configuration file:

toml
experimental = false

[experimental_features]
master = true
nodes = [
    "127.0.0.1:8081"
]

The README states plainly that these settings are not ready for use as they are very unstable or flat out broken. When experimental is set to false (the default), the distributed settings are ignored entirely. The Cargo workspace confirms this: toshi-proto and toshi-raft are commented out of the workspace members list, meaning they do not compile as part of the default build. The toshi-raft crate would handle distributed consensus, and toshi-proto would manage the protocol buffer definitions for inter-node communication. Neither is functional today.

Beyond the distributed limitation, the README opens with the statement that Toshi is far from production ready. This is an honest assessment that implies the API, data format, and configuration are subject to change without backward compatibility guarantees, and that there is no tested upgrade path between repository snapshots.

The project has no GitHub releases published. The current state of the master branch is the only version available. The repository includes a SECURITY.md at the top level for reporting vulnerabilities, which is standard practice but also signals that the project has thought about security disclosure even in its early state.

The repository also includes a doc.json and docs/ directory, which likely contains Tantivy schema documentation or API reference material, though the README does not describe their contents explicitly.

How Toshi Compares to Meilisearch

Meilisearch is another Rust-based full-text search engine with an HTTP API. It is further along in production readiness than Toshi, with versioned releases, official client libraries for many languages, and a focus on fast, typo-tolerant search for application developers.

The key architectural difference is the abstraction level. Toshi exposes Tantivy's query types directly through its API: term, fuzzy, phrase, range, regex, and boolean queries map closely to Tantivy primitives. Meilisearch abstracts the index and query model further, offering a simpler API where the developer sends documents and search strings, and the engine handles tokenization, synonym expansion, and ranking internally.

Developers who need direct control over query structure, who want to expose Tantivy's full query language through an HTTP interface, or who are building on top of Rust and want a search layer with no JVM dependency will find Toshi's approach more transparent. Developers who want a production-ready search service with a managed client library and an opinionated but simple API will find Meilisearch better suited to the task today.

License, Maintenance Status, and the Mascot

Toshi is released under the MIT license. The last push to the repository was on 2026-06-28. The README includes the caveat that Toshi is still under active development with a slow pace, so the six-month gap between major visible activity and the last push is consistent with the project's own framing.

The workspace covers four crates: toshi-server, toshi-client, toshi-types, and implicitly toshi-proto and toshi-raft (which are commented out). The toshi-client crate provides a programmatic client interface for the server. The toshi-types crate holds shared type definitions.

The project also documents its mascot: Toshi is a three-year-old Shiba Inu who personally reviews all code before it is committed. This note appears at the end of the README's What is a Toshi? section.

For contributors, the project uses cargo test for the test suite and has CI configuration in the .github/ directory. The ruff.toml file suggests Python tooling is also present in the repository for auxiliary scripts. No CHANGELOG file is mentioned in the top-level directory listing, so version history requires reading git log.

Editorial conclusion

Toshi is worth evaluating for developers who want a Rust-native full-text search server with an HTTP API and who are comfortable running pre-production software. Its stable-Rust requirement and explicit avoidance of unsafe code in its own layer make it appealing for teams who want to audit or contribute to the codebase. The distributed mode is experimental and described as very unstable in the README, so any architecture that requires multi-node search should rule Toshi out until that feature matures. The workspace structure shows toshi-proto and toshi-raft are commented out of the build, confirming the feature is not compiled by default. Verify that Tantivy's version matches Toshi's dependencies via Cargo.lock before building, since the project has no recent GitHub releases and the README's Rust version requirement of 1.39.0 may be outdated.

Frequently asked questions

How does Toshi differ from Elasticsearch?

Toshi wraps Tantivy in an HTTP API in the same way Elasticsearch wraps Lucene. The difference is technology: Toshi is a Rust binary with no JVM dependency, while Elasticsearch runs on the JVM. Toshi is also explicitly not production ready, while Elasticsearch is a mature distributed system.

How do I build and run Toshi from source?

Run cargo build --release from the repository root after installing Rust from rustup.rs. Start Toshi with ./target/release/toshi. It reads config/config.toml and binds to port 8080 by default. Verify it is running with curl -X GET http://localhost:8080/.

What query types does Toshi support?

Toshi supports six query types through its JSON HTTP API: term, fuzzy, phrase, range, regex, and boolean. Boolean queries combine must and must_not clauses from any of the other types. All queries accept an optional limit parameter that defaults to 10.

Official sources

  1. Issues
  2. License: MIT
  3. README
  4. toshi-search/Toshi on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/toshi-search-toshi.svg)](https://hysenlabs.com/projects/toshi-search-toshi)