Honcho: Memory Infrastructure for Stateful AI Agents
Memory library for building stateful agents
At a glance
- What is it?
- Honcho is an AGPL-3.0 FastAPI server from Plastic Labs that stores AI agent sessions and runs background reasoning to build queryable representations of users, agents, groups, and projects. It ships a Python SDK, a TypeScript SDK, a CLI, and integrations with MCP clients including Claude Code and Hermes.
- Who is it for?
- Honcho fits teams building products where agents need to track individual users across sessions and make inferences about their context over time. The AGPL-3.0 license is a concrete blocker for commercial SaaS products that cannot release their own source code: verify your legal position before integrating.
- Can I use it commercially?
- Yes, with strict conditions. AGPL-3.0 is a network copyleft licence: if people use a modified version over a network, for example as a hosted service, you must offer them its source code under the same licence.
- Is it still maintained?
- Yes. The repository last received commits 13 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 17, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The Problem: Agents That Lose Context Between Sessions
Most LLM agent frameworks handle a single conversation turn well, but they do not track how a user's preferences, knowledge gaps, or working context evolve across dozens of sessions. Each new conversation starts blank. The developer has to bolt on a retrieval layer, decide what to store, and write logic to inject past context into prompts.
Honcho addresses this by acting as a memory service that sits between the agent and the LLM. It receives messages and events, reasons over them in the background to extract conclusions about each participant, and exposes those conclusions through a query API. The README describes the target as engineers who want their agents to earn higher retention and build data moats by understanding users over time.
The project is particularly aimed at developers building coding assistants, tutoring applications, or long-running AI companions where per-user personalization matters. Honcho's integration list in the README includes Claude Code, OpenCode, OpenClaw, Hermes, and Cursor-compatible MCP clients, which positions it toward teams already using AI-assisted development environments as well as product developers building their own agent stacks.
How the Honcho Loop Works: Store, Reason, Query, Inject
The README describes a four-step loop. First, the application stores conversation turns, document events, or tool traces as messages on a session. Second, Honcho's background reasoning layer processes that queue asynchronously and updates per-peer representations. Third, the application queries Honcho for context, search results, or a natural-language answer about a peer. Fourth, the application injects the result into the next LLM call.
The data model is hierarchical. Workspaces contain peers. Peers participate in sessions. Messages live on sessions. A peer can be a human user, an AI agent, a group, a project, or an abstract idea, and Honcho builds a representation of each one from the messages they appear in. The multi-peer perspective feature allows the system to model what one peer knows about another, which the README describes as useful for scenarios where agents and users interact together.
Background reasoning is asynchronous. The README notes that newly added messages may take a moment before they are reflected in query responses. For low-latency reads, the representation endpoint returns the current snapshot without waiting for new reasoning to complete.
Installing Honcho and Running the First Session
The Python SDK installs from PyPI:
pip install honcho-aiThe TypeScript SDK installs from npm:
npm install @honcho-ai/sdkTo run a local server, install the CLI and call honcho start with the setup flag:
honcho start --setupThis starts a local stack. The CLI also provides honcho workspace inspect and honcho doctor for inspecting deployments. For the managed service, an API key is obtained at app.honcho.dev; the README states that new organizations receive $100 in free credits.
The Python SDK requires setting HONCHO_API_KEY for the managed service, or passing base_url="http://localhost:8000" for a self-hosted instance. The README's Python quickstart creates two peers named alice and tutor, opens session-1, and stores messages with session.add_messages(). After messages are stored, peer.chat() returns a natural-language answer about what Honcho has inferred about that peer. The TypeScript SDK follows the same peer/session model.
The server itself requires Python 3.13 or later, as declared in pyproject.toml. The Docker image builds from python:3.13-slim-bookworm and uses uv for dependency management.
Peer Representations and the Chat Endpoint
The core query primitive is the Chat Endpoint, which accepts a natural-language question about a peer and returns a natural-language answer. For example, a tutoring application can ask Honcho what learning styles a specific student responds to, and Honcho answers based on the session history it has processed.
The alternative to the Chat Endpoint is the representation endpoint, which returns a structured model of what Honcho knows about a peer. This is suited for programmatic use where the agent needs structured facts rather than a conversational answer.
Workspace-level queries run through honcho.chat() or honcho.chat_stream(), which draw on all sessions and all peers in the workspace rather than a single peer. The README shows a need table listing six capabilities: saving interaction history through session.add_messages(), asking what Honcho knows about a peer through peer.chat(), asking across a workspace through honcho.chat(), and additional endpoints for search results, session context, and prompt-ready context.
The optional surprisal-based dream prioritisation feature, controlled by the DREAM.SURPRISAL.ENABLED configuration flag, requires scikit-learn. Because scikit-learn pulls in scipy (roughly 199 MB installed on Linux), pyproject.toml keeps it as an optional dependency rather than a base one.
Self-Hosting, Infrastructure Dependencies, and Maintenance
The FastAPI server stores session data in PostgreSQL with the pgvector extension for vector similarity queries. It uses Redis for caching through cashews. The vector store can be configured to use TurboPuffer, LanceDB, or Qdrant depending on deployment requirements, though LanceDB is excluded on macOS x86-64 per pyproject.toml.
The Docker Compose example file and the Dockerfile in the repository support self-hosted deployments. The Dockerfile uses a multi-stage build with uv sync for reproducible dependency resolution. The alembic.ini and migrations/ directory handle database schema migrations.
Self-hosting Honcho therefore requires running PostgreSQL with pgvector, Redis, and optionally a vector store service. Teams evaluating the self-hosted path should budget for the operational overhead of these three additional services. The managed api.honcho.dev service removes that burden at the cost of external data transmission and the usage-based pricing model.
The repository was last pushed on 2026-09-17, and the current release is v3.2.1. The pyproject.toml shows active upstream dependencies including fastapi, sqlalchemy, pgvector, and sentry-sdk.
Where Honcho Is Not the Right Tool
Honcho's background reasoning layer adds latency and infrastructure complexity that is not justified when the use case is simple conversation replay. An application that only needs to show users their chat history, or that only needs to retrieve the last N messages for a context window, does not benefit from Honcho's reasoning pipeline. A plain PostgreSQL table or a Redis sorted set covers that requirement with far less operational overhead.
Honcho does not support multimodal inputs directly. The README describes the data model in terms of messages and events, without documentation of image or audio attachment handling.
The server requires Python 3.13 as a minimum. Teams running infrastructure on Python 3.11 or 3.12 cannot use the current server version without upgrading their runtime environment.
AGPL-3.0 and What It Means for Commercial Products
Honcho is licensed under AGPL-3.0 (GNU Affero General Public License version 3). The Affero variant of the GPL closes the so-called application service provider (ASP) loophole: if you deploy a modified version of Honcho as a network service, you must make the complete corresponding source code available to users of that service under AGPL-3.0 terms. This is distinct from Apache-2.0 or MIT, which do not impose this obligation.
For a SaaS product that integrates Honcho into its backend, this means the product's own source code may need to be released if it is considered a combined work with Honcho. Whether the Python and TypeScript SDKs, which are separate packages, carry the same copyleft effect depends on how they interact with the server: the README does not address this, and the licensing note at the top says AGPL-3.0 applies to the core service.
Teams building closed-source commercial products should review the AGPL terms and seek legal advice before deploying Honcho in production. The license file is at the repository root.
Editorial conclusion
Honcho fits teams building products where agents need to track individual users across sessions and make inferences about their context over time. The AGPL-3.0 license is a concrete blocker for commercial SaaS products that cannot release their own source code: verify your legal position before integrating. Python 3.13 or later is required. Teams that only need to store and replay conversation history do not need Honcho's background reasoning layer. The last push was on 2026-09-17, and version 3.2.1 is the current release.
Frequently asked questions
What is Honcho AI?
Honcho is a memory infrastructure service for AI agents, built by Plastic Labs. It stores conversation sessions, runs background reasoning to extract insights about each participant, and exposes queryable peer representations through a Python SDK, a TypeScript SDK, and a CLI.
How do I install Honcho?
Install the Python SDK with pip install honcho-ai or the TypeScript SDK with npm install @honcho-ai/sdk. To run a local server, install the CLI and run honcho start --setup. The managed service is available at api.honcho.dev.
How do I install Honcho locally?
Install the CLI package, then run honcho start --setup to start a local stack. Point the SDK at http://localhost:8000 by passing base_url when creating the Honcho client, or by setting the HONCHO_URL environment variable.
How do I use Honcho with Hermes?
The README lists Hermes as a supported MCP client integration. Honcho provides a hermes-plugin-honcho directory in the repository for connecting Honcho memory to Hermes-based agents. The setup path is documented in the Integrations section of the README.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/plastic-labs-honcho)