Sonic: a schema-less search backend that returns IDs, not documents
🦔 Fast, lightweight & schema-less search backend. An alternative to Elasticsearch that runs on a few MBs of RAM.
At a glance
- What is it?
- Sonic is a Rust search backend that indexes identifier tuples instead of documents, runs in tens of megabytes of RAM, and speaks its own line protocol. Here is what it does, how to install it, and where it stops being the right tool.
- Who is it for?
- Adopt Sonic when your application already owns the documents and only needs a fast word-to-identifier index with suggestions and typo tolerance, and when you can accept a separate binary, a manually edited config file, and a protocol you have to learn. Do not adopt it as a general document store, a relevance-tuning platform, or a drop-in replacement for an existing Elasticsearch deployment with aggregations and faceting.
- Can I use it commercially?
- Yes, with conditions. MPL-2.0 is a weak copyleft licence: you can use it inside commercial and closed-source software, but if you distribute changes to its own files, you must publish those changes under the same licence.
- Is it still maintained?
- Yes. The repository received new commits within the last day.
- What is it written in?
- Mainly Rust, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Sonic indexes identifiers, not documents
The README is explicit that Sonic is "an identifier index, rather than a document index; when queried, it returns IDs that can then be used to refer to the matched documents in an external database." That single sentence decides the architecture of everything around it. You keep your rows in Postgres, MySQL, or wherever they already live. You push a text and an identifier into Sonic. Searches come back as identifiers, and your application joins them back to the source of truth.
The consequence is that Sonic never becomes your storage layer, which removes an entire class of operational problem: no reindex-on-schema-change migration, no duplicated copy of user content sitting in a second datastore that now has to satisfy deletion requests, no divergence between what the index says and what the database says. It also means Sonic cannot answer a query that needs the text itself, because it does not keep it. Sonic "doesn't store any direct textual data in its index, but it still holds a word graph for auto-completion and typo corrections."
The intended user is a team that already has a database and wants search over it without running a cluster. The README's own reference point is Crisp, which the project says indexes half a billion objects on a $5/month 1-vCPU SSD cloud server, and the README also states that Sonic responds to search queries in the microsecond range while consuming around 30MB of RAM. Those are the project's measurements, not independently reproduced here.
Buckets, collections and the Sonic Channel protocol
Data is organized in two levels. Search terms live in collections, and collections live in buckets. The README suggests you "may use a single bucket, or a bucket per user on your platform if you need to search in separate indexes." Bucket-per-user is the pattern that makes Sonic viable for multi-tenant products, because isolation is structural rather than a filter you have to remember to apply on every query.
Clients talk to the server over Sonic Channel, the project's own line-based protocol, specified in PROTOCOL.md. The README describes it as covering search, ingestion (push, pop, flush a collection, flush a bucket) and administrative actions. There are official client libraries listed in the README for applications that would rather not hand-write protocol frames.
Index-altering operations are handled asynchronously. The README says "a background tasker handles the job of consolidating the index so that the entries you have pushed or popped are quickly made available for search." Read that as an eventual-consistency contract: a push is accepted, then becomes searchable after consolidation. The README does not state a guaranteed consolidation latency, so if your product shows a newly created record in search immediately after creation, you cannot assume Sonic will have it yet.
Two features come from the word graph rather than from the documents. Suggest completes a partial word in real time, and typo correction kicks in "if there are not enough exact-match results for a given word in a search query." Text is also normalized before indexing: Sonic guesses the language and strips stop words, with the README claiming full Unicode compatibility across 80+ languages.
Installing Sonic from the Debian package and running a first push
The README gives two installation routes. Pre-built packages exist for Debian-based systems, and the project warns that it "only provides 64 bits packages targeting Debian 12 for now (codename: bookworm)" and that other Debian versions or Ubuntu may work but rely on a specific glibc version. The alternative is building from source with cargo.
Start by adding the APT repository and its signing key. Run these two commands, then refresh the package lists:
echo "deb [signed-by=/usr/share/keyrings/valeriansaliou_sonic.gpg] https://packagecloud.io/valeriansaliou/sonic/debian/ bookworm main" > /etc/apt/sources.list.d/valeriansaliou_sonic.list
curl -fsSL https://packagecloud.io/valeriansaliou/sonic/gpgkey | gpg --dearmor -o /usr/share/keyrings/valeriansaliou_sonic.gpg
apt-get updateInstall the package, then edit the configuration file the package pre-fills. The README points at /etc/sonic.cfg, and the repository root also carries a config.cfg you can read for the full set of keys:
apt-get install sonic
nano /etc/sonic.cfg
service sonic restartIf you prefer to build it yourself, the README says to install build-essential, clang, libclang-dev, libc6-dev, g++ and llvm-dev, then build with cargo. The binaries land in ./target/release:
cargo build --locked --releaseThere is also a Dockerfile in the repository. It builds with a Rust image, strips the resulting binary, and copies it into gcr.io/distroless/cc. The image declares EXPOSE 1491, so that is the port to publish when you run the container.
Once the server is up, ingestion and search go over Sonic Channel on that port. The protocol specification in PROTOCOL.md is the authoritative reference for the exact command frames; the README does not inline them, so read that file before writing a client by hand.
What Sonic does not do, and when it is the wrong choice
The most common mistake is treating Sonic as a smaller Elasticsearch. It is not, and the README says so in the same breath as the comparison: Sonic "can be used as a simple alternative to super-heavy and full-featured search backends such as Elasticsearch in some use-cases." The qualifier matters. If your search needs relevance scoring you can tune per field, faceted navigation, aggregations, highlighting, or stored document retrieval, Sonic gives you none of that. It returns identifiers and a relevance ordering; everything the user sees next is your application's job.
There is a second, sharper limitation that comes from the identifier model. Because Sonic returns IDs and the README describes results as identifier tuples, pagination and result-set stability depend on how those identifiers are ordered and how many matches a query yields. The README does not document a cursor or offset mechanism, so a product that needs deep paging, stable result windows across requests, or exact result counts is building those guarantees on top of a system that does not advertise them. Verify this against your own identifier scheme before designing the UI around it.
The asynchronous indexing path is a third boundary. Push and pop are consolidated by a background tasker, and the README gives no bound on how long that takes. A workflow that writes a record and immediately searches for it will occasionally miss. Finally, the deployment surface is narrow: 64-bit Debian 12 packages, or a source build with a specific toolchain. If you run on Alpine, on ARM, or on a distribution far from bookworm, plan on the cargo route and the C toolchain it needs.
Sonic against Meilisearch and Typesense
The realistic alternatives for a small self-hosted search backend are Meilisearch and Typesense. Both are document stores: you send them the JSON records you want searchable, they keep a copy, and they return matching documents with highlighting and faceting built in. Sonic inverts that. It keeps a word graph and identifier tuples, and your database stays the only place the document exists.
That inversion is the whole trade. With a document store you get richer query features and a simpler mental model, at the cost of a second copy of your data to keep in sync, reindex, and delete from. With Sonic you get a much smaller process and no data duplication, at the cost of implementing result enrichment, paging and presentation yourself.
There is also a protocol difference worth weighing. Meilisearch and Typesense expose HTTP APIs, so any HTTP client in any language can talk to them, and proxies, load balancers and observability tooling work without special handling. Sonic Channel is a bespoke line protocol with its own client libraries. That is lighter on the wire, and it is one more thing your team has to learn and version. If your organization already runs an HTTP-based search service, the migration cost to Sonic is not the server, it is the client code.
Licence, maintenance and upgrade cost
Sonic is licensed under MPL-2.0, declared in the workspace Cargo.toml and shipped as LICENSE.md. MPL-2.0 is a file-level copyleft licence: modifications to files that are part of the covered source must be made available under the same licence, while larger works that combine Sonic with separate files can be distributed under other terms. This is a general description of the licence family, not legal advice; if you patch Sonic and ship it, have counsel read the actual text.
The repository is not archived, and the last push was on 2026-09-15, the same day as the v1.9.1 release. The workspace also tracks a core crate, with core-v0.3.0 released alongside the server, so the protocol layer and the server version independently. rust-toolchain.toml and the Cargo.toml rust-version field pin the compiler expectations, and the README states the project was tested at rustc 1.98.0. Upgrading therefore has two axes: the server version and the toolchain, and a client library that has not tracked the protocol will be the thing that breaks first. Read CHANGELOG.md before moving a production instance, and check that your client library's release notes mention the same protocol revision.
Editorial conclusion
Adopt Sonic when your application already owns the documents and only needs a fast word-to-identifier index with suggestions and typo tolerance, and when you can accept a separate binary, a manually edited config file, and a protocol you have to learn. Do not adopt it as a general document store, a relevance-tuning platform, or a drop-in replacement for an existing Elasticsearch deployment with aggregations and faceting. Before committing, verify three things against your own data: that your identifier ordering is stable enough for the pagination you plan to expose, that the language you index is in Sonic's supported list, and that your client library speaks the current protocol version rather than an older one.
Frequently asked questions
What is Sonic and what does it actually do?
Sonic is a fast, lightweight, schema-less search backend written in Rust. It ingests search texts together with identifier tuples and, when queried, returns the identifiers that match, which your application resolves against an external database.
How do I install Sonic on Debian or Ubuntu?
The README documents pre-built packages for Debian-based systems. You add the Sonic APT repository and its signing key, run apt-get update, install with apt-get install sonic, edit /etc/sonic.cfg, and restart the service. The project warns that it only provides 64-bit packages targeting Debian 12 (bookworm) for now.
Does Sonic store my documents?
No. The README states that Sonic doesn't store any direct textual data in its index and that it is an identifier index rather than a document index. It does keep a word graph used for auto-completion and typo corrections.
What port does the Sonic Docker image expose?
The repository Dockerfile declares EXPOSE 1491, and the image is built on gcr.io/distroless/cc with the stripped sonic binary copied to /usr/local/bin/sonic.
Is Sonic a replacement for Elasticsearch?
The README describes Sonic as a simple alternative to super-heavy and full-featured backends such as Elasticsearch in some use-cases. It returns identifiers rather than documents, so features like faceting, aggregations and stored document retrieval are outside its scope.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/valeriansaliou-sonic)