Library / SDK
vcaesar/riot avatar
vcaesar/riot

Riot: a Go text indexing library forked from Bluge, with CJK and vector search built in

Go Open Source, Distributed, Simple and efficient Search Engine

6,051 stars468 forksGoApache-2.0

At a glance

What is it?
Riot is a Go library for text indexing and search, forked from Bluge, adding a gse-backed CJK layer, vector queries and aggregations. It suits Go services that want an embedded index rather than a separate search server.
Who is it for?
Adopt Riot if you are writing Go and want an embedded index with CJK tokenization or vector fields without running a separate search server. Do not adopt it if you need a query DSL over HTTP, a managed cluster, or a project with a long independent track record; it is a fork, and the README inherits Bluge's concepts without documenting migration from Bluge.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 10 days ago.
What is it written in?
Mainly Go, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What Riot is for, and who it is not for

Riot is a Go library, not a server. The README describes it as "A fast, modern text indexing library in Go, forked from Bluge", and the repository layout supports that: config.go, writer.go, reader.go, search.go and document.go sit at the top level, with index/, search/, numeric/, analysis/ and hnsw/ as subpackages. There is no HTTP listener in the tree, no REST API documented, and no dashboard. If your mental model of a search engine is Elasticsearch or Meilisearch, Riot is a different category of thing: you link it into your binary, you own the index directory, and you own the process lifecycle.

The audience is therefore narrow and specific. You are writing a Go service. You want full-text search over documents you already hold, and you would rather not operate a second system, ship a network hop, or write a client library for someone else's protocol. Riot fits that. It also fits the case where your corpus is Chinese or Japanese, because the gse subpackage wraps the index with a tokenizer from github.com/go-ego/gse plus query-string search and highlighting. That CJK layer is the most distinctive part of the project relative to its Bluge ancestry, and it is the reason a Go team might pick this fork over the upstream library.

Where it is the wrong tool: multi-tenant search with per-user access control, cross-language clients, or anything where search is a shared service consumed by applications you do not build. Riot gives you no query language over the wire and no server to point other teams at.

The indexing and search mechanism in Riot

The data flow is the classic writer/reader split inherited from Bluge. A Config describes where the index lives; riot.DefaultConfig(path) builds one rooted at a directory. riot.OpenWriter(config) returns a writer, and documents are added with writer.Update(doc.ID(), doc). A document is riot.NewDocument(id) chained with AddField calls, and fields are typed: the README lists Text, Numeric, Date, Boolean, IP, Geo Point and Vector. Because the writer is the only mutation path, the caller decides document identity, and Update is the upsert operation.

Reads go through a separate handle. riot.OpenReader(config) opens the same directory, and the README notes that if you index and search in the same process you should use writer.Reader() instead of riot.OpenReader, or riot.InMemoryOnlyConfig() for an index that never touches disk. That detail matters more than it looks: opening a reader over a directory a live writer holds is a concurrency question the README flags but does not fully answer.

Queries are built as values, not strings. riot.NewMatchQuery("riot").SetField("name") is wrapped in riot.NewTopNSearch(10, query), optionally with WithStandardAggregations(), and executed with reader.Search(context.Background(), request). The return value is an iterator, not a slice: you call Next() in a loop and VisitStoredFields to pull stored values back, which is why the README example prints only the _id field. The field and query type lists are long. Queries cover Term, Phrase, Match, Match Phrase, Prefix, Regexp, Wildcard and Fuzzy, plus Conjunction, Disjunction and Boolean, the range families (Numeric, Date, Term, IP), the geo families (Bounding Box, Distance, Polygon) and ANN and KNN for vectors. Scoring is BM25 with pluggable interfaces.

Aggregations are the part most likely to be underestimated. The README lists bucketing by Terms, Numeric Range and Date Range, and metrics including Min, Max, Count, Sum, Avg, Weighted Avg, Cardinality Estimation via HyperLogLog++ and Quantile Approximation via T-Digest. The go.mod confirms the dependencies: github.com/axiomhq/hyperloglog and github.com/caio/go-tdigest/v5. Those are approximate structures, so cardinality and quantile numbers carry the usual estimation error of HyperLogLog and t-digest; the README does not state error bounds.

Installing Riot and indexing your first document

Riot is a Go module, so installation is a module fetch. The README gives one command:

bash
go get -u github.com/vcaesar/riot

After that, the repository's test/readme_demo directory holds runnable versions of the basic indexing, querying and CJK programs. The README's indexing example is meant to be saved as write/main.go and run with go run ./write. It opens a writer over ./riot_index, builds a document with a single text field, and upserts it:

go
config := riot.DefaultConfig("./riot_index")
writer, err := riot.OpenWriter(config)
if err != nil {
	log.Fatalf("error opening writer: %v", err)
}
defer writer.Close()

doc := riot.NewDocument("example").
	AddField(riot.NewTextField("name", "riot"))

err = writer.Update(doc.ID(), doc)

The reader side is a second program, saved as read/main.go and run against the index just written. The query is a match query scoped to the name field, wrapped in a top-10 search:

go
query := riot.NewMatchQuery("riot").SetField("name")
request := riot.NewTopNSearch(10, query).
	WithStandardAggregations()
documentMatchIterator, err := reader.Search(context.Background(), request)

The README states the expected output of that pair is a single line, match: example. Note the two constraints the README attaches to this walkthrough: run the reader program against the index written by the writer program, and if you index and search inside one process, use writer.Reader() rather than riot.OpenReader, or riot.InMemoryOnlyConfig() if you do not want a directory at all.

Chinese and Japanese text through the gse wrapper

The gse subpackage is the reason to look at this fork rather than upstream Bluge. It wraps the index with a gse tokenizer and adds query-string search and highlighting, so you do not hand-build a query tree for each user input. Configuration is a single struct. gse.Option carries Index, Dicts and Opt, where Dicts is a comma-separated dictionary list such as "embed, ja" or "embed, zh" and Opt is a mode string such as "search-hmm". Setting Option{Lang: "en"} instead skips gse and uses a riot analysis/lang analyzer; gse.Langs() enumerates the available languages.

The README's CJK program indexes five documents into a temporary index and runs three query strings through gse.QueryString. The output shown includes per-hit scores and highlighted fragments, with matched terms wrapped in mark tags, for example a query for 搜索引擎 returning one hit from the document Riot 是用 Go 语言编写的全文搜索引擎. The README labels the timings as varying, so treat the microsecond figures in that block as illustrative rather than a benchmark.

The same wrapper handles structs. The examples/gse_struct directory holds a runnable version: gse.New(gse.Option{Lang: "en"}) with an empty Index means memory only, index.Index("article-1", Article{...}) takes a struct directly, and the query request exposes a Field property so gse.QueryString("started", true).Field = "title" restricts the search to one field. Results come back as result.Hits with hit.ID, hit.Fields and hit.Fragments. That is a noticeably friendlier surface than the raw reader API, and for a CJK or mixed-language corpus it is the part of Riot I would evaluate first.

Vector search, ANN and KNN, and the hnsw package

Vector is listed among the supported field types, and ANN and KNN appear in the query list. The repository has a top-level vector.go with vector_test.go, plus an hnsw/ directory and an examples/knn/ example. Hierarchical navigable small world graphs are the standard approximate nearest neighbour structure, so the presence of hnsw/ indicates the vector path is approximate rather than exhaustive, but the README does not describe the index parameters, the distance metric, or how recall degrades as the graph is built. Anyone adopting Riot for embeddings should read hnsw/ and examples/knn/ directly rather than relying on the README, which gives the vector feature a single line in a feature list.

Two practical points follow. First, vector search in Riot is library-level: you supply vectors as field values and query them in-process, so the embedding model, its version and any re-embedding pipeline are entirely your problem. Second, the combination of BM25 text scoring and vector scoring is not described in the README. If you need hybrid ranking, that is a design you would build on top of separate queries rather than something the documentation promises.

The Bluge relationship, and what a fork costs you

Riot is forked from Bluge, and the README says so in its first sentence. The dependency list shows how deep that inheritance runs: go-porterstemmer, mmap-go, segment, snowballstem and vellum all come from the blevesearch organisation, and the segment API is pulled from github.com/vcaesar/bluge_segment_api. The concepts a Bluge user knows, writer, reader, document, field, TopNSearch, aggregations, transfer almost unchanged.

The alternative worth naming is Bluge itself. The difference in approach is not architectural but directional: Bluge is the upstream project with its own release history and community, while Riot is a fork maintained under a single GitHub account that has added a gse-based CJK layer, vector and hnsw support, and a cobra-based cmd/ entry point. Choosing Riot means choosing those additions and accepting that the fork, not the upstream, is where your fixes would have to land. If you need CJK tokenization or vector fields, the fork is the point. If you need neither, upstream Bluge is the lower-risk dependency because you are not betting on a divergence being maintained.

A related cost: the README does not document how to migrate an existing Bluge index to Riot, and it does not state whether the on-disk formats remain compatible. The segment API being a separate module under the same author suggests the formats may have diverged. Verify that before planning any migration.

Maintenance, licensing and upgrade cost

The repository is not archived, and the last push was on 2026-09-21. Releases are frequent: v1.30.0 on 2026-09-21, v1.23.1 on 2026-09-13 and v1.23.0 on 2026-09-12. That cadence cuts both ways. You get fixes quickly, and you also get a moving target; the jump from v1.23.x to v1.30.0 in roughly a week suggests the version numbers are not signalling API stability in the way a slower project's would. Pin a version in go.mod and read the release notes before bumping.

The go.mod declares go 1.26.0, so your toolchain has to be at least that new. The dependency set is moderate but not trivial: roaring bitmaps, bitset, vellum for the term dictionary, mmap-go for file access, gse for tokenization, cobra for the command line, gonum as an indirect dependency, plus the segment API and ice modules from the same author. Each is a supply-chain surface you inherit.

The licence is Apache-2.0, per the LICENSE file at the repository root. Apache-2.0 permits commercial and closed-source use and includes a patent grant, and it requires that you preserve the licence and notices, including for the bundled third-party dependencies, which carry their own terms. That is a description of the licence text, not legal advice; if you redistribute Riot inside a product, have counsel review the notice file you ship.

Editorial conclusion

Adopt Riot if you are writing Go and want an embedded index with CJK tokenization or vector fields without running a separate search server. Do not adopt it if you need a query DSL over HTTP, a managed cluster, or a project with a long independent track record; it is a fork, and the README inherits Bluge's concepts without documenting migration from Bluge. Before committing, check the vector and hnsw packages for the ANN and KNN query path, confirm the gse dictionary bundles you need are present under gse/, and read the Apache-2.0 LICENSE together with the notices for the bundled dependencies.

Frequently asked questions

How do I install Riot in a Go project?

Run go get -u github.com/vcaesar/riot, which is the single installation command the README gives. Your toolchain must satisfy the go 1.26.0 directive in the module's go.mod.

Does Riot handle Chinese and Japanese text?

Yes, through the gse subpackage, which wraps the index with a gse tokenizer and adds query-string search and highlighting. Its Option struct takes a Dicts value such as "embed, ja" or "embed, zh", and setting Option{Lang: "en"} instead uses a riot analysis/lang analyzer.

Can I search and index in the same process with Riot?

The README says that if you index and search in the same process you should use writer.Reader() rather than riot.OpenReader, or use riot.InMemoryOnlyConfig() for an in-memory index. The read example is written to run against an index a separate writer program produced.

What query types does Riot support?

The README lists Term, Phrase, Match, Match Phrase, Prefix, Regexp, Wildcard and Fuzzy; Conjunction, Disjunction and Boolean; Numeric, Date, Term and IP ranges; Geo Bounding Box, Geo Distance and Geo Polygon; and ANN and KNN for vectors. Scoring is BM25 with pluggable interfaces.

Does Riot include vector search?

Vector is listed as a supported field type and ANN and KNN appear among the query types, with an hnsw/ package and an examples/knn/ example in the repository. The README does not document the HNSW index parameters or the distance metric, so those need to be read from the code.

Official sources

  1. Issues
  2. License: Apache-2.0
  3. README
  4. Releases
  5. vcaesar/riot on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/vcaesar-riot.svg)](https://hysenlabs.com/projects/vcaesar-riot)