Open-source project
ekailabs/contexto avatar
ekailabs/contexto

Contexto: a context engine for the turn after compaction

Context Engine for your long-running AI agents

620 stars22 forksTypeScriptApache-2.0

At a glance

What is it?
Contexto is a plugin for OpenClaw and Hermes that stores whole conversation episodes and retrieves them, so an instruction from turn two is still available at turn thirty-five. It indexes with hierarchical clustering and pulls context back with multi-branch beam search, and the storage layer is an interface you can replace with your own backend.
Who is it for?
Use Contexto if you run OpenClaw or Hermes sessions long enough to compact and a forgotten constraint has already cost you something, since the README says exactly that audience and explicitly excludes one-shot chats. Do not expect it to help a short session.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 118 days ago.
What is it written in?
Mainly TypeScript, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 4, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The failure it targets shows up at turn thirty-five

The README states the failure mode before describing the product. On turn two a user says to flag suspicious emails and not to delete anything. Thirty turns of tool calls, retries and compaction follow. By turn thirty-five, without help, the agent deletes twelve flagged emails, because the constraint was lost when the transcript was summarised.

With the plugin in place, the same turn produces four newly flagged emails and a retrieved context line reading that the user constraint is flag only, never delete. The instruction survives compaction.

That example is chosen carefully and it is worth reading what it leaves out. The failure is not that the model forgot, it is that the harness discarded the evidence. A summariser compressing turn two into turn thirty cannot preserve a negation reliably, because the compressed form of "do not delete" is easily rendered as "flag" and the prohibition is lost in the compression rather than in the model.

The stated symptoms are four: early instructions compacted away, summaries turning into summaries of summaries, unrelated topics blurring together, and an agent becoming less reliable the longer you use it.

The storage unit is a full episode, tool output included

Contexto's answer is to stop summarising and start indexing. The mechanism has four steps.

The agent buffers conversation turns as full episodes rather than as message text alone. When the prompt budget crosses the compaction threshold, the oldest episodes are ingested rather than summarised. Those episodes are clustered with hierarchical similarity so related work lands in the same branch. Retrieval then uses beam search to pull back the most relevant episodes for the current prompt.

The storage unit being a full turn, including tool output, is the detail that decides whether this works. A constraint that lived in a tool result rather than in a user message is still captured, and a conversation whose important content was a stack trace is not reduced to the sentence that preceded it.

The summary the README offers is short enough to be the design statement: old context is not gone, it is organized. Nothing is thrown away at compaction time; the history moves from the prompt into a tree.

Clustering is AGNES, retrieval is multi-branch beam search

Two algorithms carry the design, and both are named in the README.

Hierarchical clustering is AGNES. The stated property is that related episodes are grouped without predefined categories, which is what lets a conversation drift between topics without collapsing them into one undifferentiated stream. Topic separation is listed as a capability precisely because a naive index would blend unrelated work together.

Retrieval is multi-branch beam search, meaning a single pass can pull from several relevant branches at once rather than committing to one. That matters for the same reason: if your session covered a trip planning thread and a code review thread and the current prompt touches both, a single-branch retriever has to pick one and loses the other.

The fourth mechanism is the rebuild strategy, described as hybrid: periodic full rebuilds with cheaper incremental inserts in between. The tree therefore gets repaired on a schedule without paying for a full rebuild on every new episode.

Retrieval is also explainable, which the comparison table treats as a differentiator. A retrieved item can come back with a path such as `travel -> Japan -> visa docs`, so you can see why it was pulled rather than trusting an opaque score.

OpenClaw install is a plugin, a slot, and a key

The OpenClaw path is five commands, and the middle two are the ones people miss:

bash
openclaw plugins install @ekai/contexto
openclaw plugins enable contexto
openclaw config set plugins.slots.contextEngine contexto
openclaw config set plugins.entries.contexto.config.apiKey YOUR_KEY
openclaw gateway restart

Installing the plugin is not enough. It has to be enabled, then it has to be assigned the context engine slot, because the default slot still belongs to the built-in summarising behaviour. Only after the slot is set does the gateway pick it up, which is why the restart is the last command rather than an optional extra.

The configuration surface is one property. The whole table is a single row: `apiKey`, a string, required, described as your Contexto API key. There is nothing else to tune, which is either a strength or a limitation depending on how much you want to control retrieval.

Managed hosting is offered so you do not have to run retrieval infrastructure yourself, and the key comes from getcontexto.com.

Hermes takes a pip package and two lines of YAML

The Hermes path is a Python install rather than a plugin manager call:

bash
pip install contexto-hermes
contexto-hermes-install            # symlinks the plugin into hermes-agent

The second command is the interesting one, since it does not configure anything, it symlinks the plugin into the agent. Then the engine is selected in `~/.hermes/config.yaml`:

yaml
context:
  engine: contexto

and the key is exported before starting the gateway:

bash
export CONTEXTO_API_KEY=YOUR_KEY
hermes gateway run

So the two agents differ in mechanism and not in outcome. OpenClaw assigns a slot through its own config command; Hermes selects an engine name in a YAML file. Both end with a restart of the gateway.

For a fully local setup, using embeddings and summarization against your own OpenAI or OpenRouter key with the mindmap kept on disk, there is a separate quickstart document for the Hermes path.

Storage sits behind ContextoBackend, and the default is remote

The one design decision worth examining is the backend interface. The engine does not talk to a database directly, it talks to `ContextoBackend`, and the interface has two methods: `ingest`, which accepts a single webhook payload or an array of them and returns nothing, and `search`, which takes a query, a maximum result count, an optional filter and an optional minimum score, and resolves to a search result or null.

The default implementation is remote and calls `api.getcontexto.com`. The README states you can implement your own instead, and the shape of the interface is small enough that it is a realistic weekend of work for anyone who needs their episodes to stay in a store they control.

Two details make it less trivial than it looks. `search` returning null rather than an empty result leaves the decision about whether a miss is an error to the caller. And the optional filter argument means the interface anticipates scoping even though scoped context with access boundaries is still unchecked on the roadmap, which suggests the retrieval layer was designed for multi-tenant use before the access model shipped.

Everything on the roadmap is unchecked

Four items are listed, none of them done: horizontal scaling with sub-agent context delegation, scoped context with access boundaries, knowledge from external documents, and context sharing across agents.

Each one names a limitation of the current design. Without sub-agent delegation, a parent agent cannot hand a child a slice of its own context. Without access boundaries, there is one store and no permissions on it, which matters the moment two agents share an installation. Without external documents, everything in the tree came from a conversation. Without sharing, two agents on the same problem build two separate memories.

Read against the audience section, the exclusions are as informative. The intended users are OpenClaw and Hermes users whose sessions run long enough to compact, agents where forgotten constraints are costly, and teams that want reliability without prompt hacks. It is explicitly not for one-shot chats or very short sessions, which is the correct call given that the whole mechanism keys off the compaction threshold.

The last push to main is dated 2026-06-10, and the newest release tag is v0.1.11 from 2026-04-07, alongside a legacy snapshot tag from 2026-02-10.

The root package is an AI proxy, not the context engine

The repository is the Contexto plugin, but the manifest at the root belongs to something else. It is named `ekai-gateway`, version 0.1.0-beta.1, private, and described as an AI proxy with memory and dashboard, with workspaces covering `packages/*` and `packages/ui/*` and an author of Ekai Labs.

That explains the rest of the tree. The `Dockerfile` is a multi-stage build that installs workspace dependencies with pnpm from a frozen lockfile, then builds the dashboard, the memory package and the openrouter package in separate stages, rebuilding memory before openrouter because the latter depends on it, and finally produces an embedded static export of the dashboard for a single container image. A second file, `Dockerfile.selfhost`, sits beside it.

The `.env.example` matches that scope rather than the plugin's. It collects keys for Anthropic, OpenAI, xAI, OpenRouter, ZAI and Google, documents an Ollama base URL for local models, sets a SQLite database path, a dashboard port and an OpenRouter port, and notes that memory is embedded in the OpenRouter process rather than run separately. Releases are driven by multi-semantic-release, with a dry-run script beside it.

So the plugin you install and the gateway it ships inside are two layers of the same project, and the configuration you touch first is the plugin's single apiKey, not any of this.

Editorial conclusion

Use Contexto if you run OpenClaw or Hermes sessions long enough to compact and a forgotten constraint has already cost you something, since the README says exactly that audience and explicitly excludes one-shot chats. Do not expect it to help a short session. Before you install, decide whether the default remote backend at api.getcontexto.com is acceptable or whether you will implement `ContextoBackend` yourself, and note that the four capabilities on the roadmap, delegation, scoped access, external documents and cross-agent sharing, are all unchecked.

Frequently asked questions

What is Contexto?

An Apache-2.0 context engine plugin for OpenClaw and Hermes that stores full conversation episodes and retrieves the relevant ones, so earlier instructions survive compaction instead of being compressed away. It is aimed at long-running agents whose sessions compact.

How do I install Contexto with OpenClaw?

Run openclaw plugins install @ekai/contexto, enable it, set plugins.slots.contextEngine to contexto, set the plugin apiKey with openclaw config set, then run openclaw gateway restart. Assigning the slot is the step that makes the engine take over from the built-in summarising behaviour.

How do I install Contexto with Hermes?

Run pip install contexto-hermes and then contexto-hermes-install, which symlinks the plugin into hermes-agent. Select the engine in ~/.hermes/config.yaml with context and engine set to contexto, export CONTEXTO_API_KEY, and run hermes gateway run.

Can Contexto run without its hosted service?

The engine talks to storage through the ContextoBackend interface, whose default implementation calls api.getcontexto.com, and the README says you can implement your own. A quickstart document covers a fully local Hermes setup using your own OpenAI or OpenRouter key with the mindmap on disk.

Is Contexto useful for short conversations?

No. The audience section rules out one-shot chats and very short sessions, and lists OpenClaw or Hermes users whose sessions run long enough to compact, agents where forgotten constraints are costly, and teams wanting reliability without prompt hacks.

Official sources

  1. ekailabs/contexto on GitHub
  2. License: Apache-2.0
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/ekailabs-contexto.svg)](https://hysenlabs.com/projects/ekailabs-contexto)