usemoss/moss: a sub-10 ms retrieval runtime you embed in your own process
The retrieval layer for production AI systems. Lightning-fast (<10ms) search without vector databases. Built for browser, edge, on-device, and cloud.
At a glance
- What is it?
- Moss is an in-process semantic search runtime from usemoss, distributed as Python, TypeScript, Elixir and C SDKs plus a WebAssembly build for the browser. The pitch is that retrieval stops being a network round trip, and the trade-off is that index management moves into your application.
- Who is it for?
- Adopt Moss when retrieval latency is on the critical path of a conversation and you can accept that indexes are created and loaded from inside your own process, with credentials from moss.dev. Do not adopt it if you need a query language, joins or ad hoc analytics over the corpus, because the README describes a search runtime rather than a database.
- Can I use it commercially?
- Yes. BSD-2-Clause is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository received new commits within the last day.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The round trip Moss is trying to delete
The README frames the problem in one sentence: most retrieval stacks call out to a remote vector database, and "the round trip alone runs 200-500 ms - enough to break a real-time conversation." That number is the project's own framing of the status quo, not an independent measurement, but the shape of the argument is familiar to anyone who has put a hosted vector store behind a voice agent. The retrieval step is cheap in compute and expensive in geography.
Moss answers by moving search and embedding inside the calling process. The README states plainly that there is "no network hop on the hot path," so query latency lands in single-digit milliseconds. The intended audience is narrow and specific: voice bots, copilots, and agents that talk to humans in turn-taking time. If your retrieval happens once per document upload and never again, this is solving a problem you do not have.
What actually happens between create_index and query
The API surface is small enough to describe from the README alone. A client is constructed from a project id and project key. create_index takes a name and a list of documents, each with an id and a text field. load_index then pulls that index into the runtime. query takes the index name, a string, and options such as top_k, and returns a result object whose docs carry a score and the original text, alongside a time_taken_ms field.
Two details matter more than the method names. First, load_index is a separate step from create_index, which tells you the runtime holds index state in memory rather than querying a remote store on every call; the README does not document what happens to that state when the process restarts or how much memory a given corpus occupies. Second, embedding is part of the measured path. The benchmark table explicitly notes that "Moss includes embedding in the measurement" while the competitors use an external embedding service, so the comparison is end-to-end for Moss and search-only for the others. That is a fair way to measure a user-visible latency budget, but it is not a like-for-like comparison of search engines, and the README's own footnote is the place to see that.
Retrieval is hybrid, combining semantic and keyword search in a single query, with metadata filters using $eq, $and, $in and $near. Embedding models are built in, with a bring-your-own path. The repository layout backs this up: examples/python/custom_embedding_sample.py and examples/javascript/custom_embedding_sample.ts exist alongside the default samples, and examples/python/metadata_filtering.py covers the filter operators.
Installing the Python SDK and running a first query
The README's prerequisites are Python 3.10+ or Node.js 20+, and a free-tier account at moss.dev for a project_id and project_key. The environment template in the repository keeps credentials in two variables, MOSS_PROJECT_ID and MOSS_PROJECT_KEY, and notes that per-feature credentials for things like LiveKit or OpenAI live in the .env.example of the specific subdirectory that needs them.
Install the Python package from PyPI:
pip install mossThen create an index, load it, and query it. The example below is the README's quickstart with the placeholder credentials left as they are:
from moss import MossClient, QueryOptions
client = MossClient("your_project_id", "your_project_key")
await client.create_index("support-docs", [
{"id": "1", "text": "Refunds are processed within 3-5 business days."},
{"id": "2", "text": "You can track your order on the dashboard."},
{"id": "3", "text": "We offer 24/7 live chat support."},
])
await client.load_index("support-docs")
results = await client.query("support-docs", "how long do refunds take?", QueryOptions(top_k=3))
for doc in results.docs:
print(f"[{doc.score:.3f}] {doc.text}")What you should see is one printed line per returned document, each prefixed with a score to three decimal places, ordered by relevance. The README's own comment on that loop notes the results carry a time_taken_ms value, so the latency of that particular call is available on the result object rather than only in a benchmark table. If you prefer TypeScript, the equivalent is npm install @moss-dev/moss with camelCase methods (createIndex, loadIndex, query) and a topK option instead of top_k. The README does not document what create_index does if the index name already exists, so treat re-indexing as something to test on a throwaway index first.
Where Moss is the wrong tool, and what the docs leave open
Moss is not a database, and the README says so directly. There is no query language, no joins, no aggregation, and no mention of persistence guarantees, replication or backup. If your retrieval layer doubles as the system of record for the corpus, or if analysts need to run ad hoc queries against it, an in-process runtime is the wrong shape. The same applies to corpora that change constantly under concurrent writers: the create-and-load model implies a lifecycle you control, and the README does not describe incremental updates or how a partially loaded index behaves.
The documentation is also thin in places that matter for production. There is no stated memory ceiling per index, no guidance on how many indexes a single process can hold, and no documented behaviour when load_index is called for an index that does not exist. The benchmark table reports P50 through P99 on 100,000 documents with top_k=5 on an M4 Pro with 24GB, which describes one machine and one corpus size; extrapolating to a 10-million-document corpus or a memory-constrained edge device is not something the README supports. The repository does ship a benchmarks/ directory and the README links to it for reproduction, which is the honest way to handle that gap, but reproduction is left to the reader.
One more boundary: the SDKs are listed as Python, TypeScript, Elixir and C, with a separate WebAssembly package for the browser. If your stack is Go, the repository has examples/go/ samples, but the README's SDK list does not include Go, so check what those samples actually bind to before planning around it.
How this differs from running Qdrant or Chroma yourself
The obvious alternative is a self-hosted vector database. Qdrant and ChromaDB appear in the README's benchmark table, and the difference in approach is architectural rather than a matter of tuning. A self-hosted Qdrant is a separate service with its own process, its own persistence, and a network call from your application; you get durability, a query API and operational tooling in exchange for the round trip. ChromaDB sits closer to the application and is often embedded in-process in Python, which makes it the nearest neighbour to Moss in deployment shape, but the README's table measures it at 351.8 ms P50 against Moss at 3.1 ms, and again the footnote matters: ChromaDB's number excludes embedding in that comparison while Moss's includes it.
A hosted service like Pinecone pushes the trade in the other direction. You give up control of the runtime and pay a network hop, and the README measures that hop at 432.6 ms P50. The honest reading is that Moss is buying latency by accepting responsibility for index lifecycle inside your application, and that bargain only pays off when the latency is visible to a user. For a nightly batch job that embeds a corpus and writes vectors, none of this matters and a managed database is less work.
Licence, maintenance and the cost of upgrading
Moss is BSD-2-Clause, which is permissive and short: it allows use and redistribution with the copyright notice and disclaimer retained, and it does not carry the patent grant that BSD-3-Clause or Apache-2.0 include. That is a real difference if patent exposure is a concern in your organisation, and it is worth a look from whoever handles licensing rather than from this article. The BSD-2-Clause text does not impose copyleft obligations on your application, which is the usual reason teams pick it over GPL-family licences.
On maintenance, the repository is not archived and the last push was on 2026-09-08, roughly two weeks before this writing, so the project is being worked on now. The most recent releases listed are the Moss iOS SDK v0.6.0, v0.6.1 and v0.6.2, all dated 2026-06-29. Note what that means: the release feed is dominated by the iOS SDK, while the README leads with Python and TypeScript. Do not read the release list as a statement about the Python package's version, because no such version is published alongside the repository description.
Upgrade cost is hard to estimate from what is available. The README does not document a versioning policy, a deprecation window, or a compatibility matrix between SDK versions and the runtime. The presence of a ROADMAP.md and an AGENTS.md at the repository root suggests the project tracks its own direction in-repo, which is where to look before pinning a version. Because the SDK is embedded in your process, an upgrade is a dependency bump plus a redeploy rather than a server migration, which is the cheaper half of the bargain; the expensive half is that you own the index lifecycle, so any change to create_index or load_index semantics lands in your code.
Editorial conclusion
Adopt Moss when retrieval latency is on the critical path of a conversation and you can accept that indexes are created and loaded from inside your own process, with credentials from moss.dev. Do not adopt it if you need a query language, joins or ad hoc analytics over the corpus, because the README describes a search runtime rather than a database. Before committing, verify two things yourself: that your target language has a supported SDK (Python 3.10+, Node.js 20+, Elixir and C are listed), and what the free tier at moss.dev actually covers, since the README states only that one exists.
Frequently asked questions
What is Moss in AI?
Moss is a sub-10 ms semantic search runtime for conversational AI agents, distributed as an SDK that embeds in your application rather than as a separate database service. The README describes it as a retrieval layer with hybrid search, built-in embeddings and metadata filtering, plus a WebAssembly build that runs in the browser.
How do I install and use Moss?
Install the Python package with pip install moss or the TypeScript package with npm install @moss-dev/moss, then construct a client from a project_id and project_key obtained at moss.dev. You create an index with your documents, call load_index to bring it into the runtime, and query it with a string and options such as top_k.
Does Moss need a vector database or an OpenAI key?
No. The README states that embedding models are built in and that no OpenAI key is required, though you can bring your own embedding model instead. It also states explicitly that Moss is a search runtime rather than a database, so there are no clusters to tune or HNSW parameters to set.
Which programming languages does the Moss SDK support?
The README lists Python 3.10+, TypeScript and Node.js 20+, Elixir, and C via libmoss, with a separate WebAssembly package for browser use. The repository also contains examples/go/ samples, but Go is not in the README's SDK list.
Under what licence is the Moss repository released?
Moss is released under BSD-2-Clause, a permissive licence that requires the copyright notice and disclaimer to be retained and does not impose copyleft obligations on your application. Unlike BSD-3-Clause or Apache-2.0, it does not include an explicit patent grant.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/usemoss-moss)