alash3al/stash: a self-hosted memory layer for MCP agents, backed by Postgres
Stash — persistent memory layer for AI agents. Episodes, facts, and working context stored in Postgres. MCP server included. Self-hosted, single binary, no cloud required.
At a glance
- What is it?
- Stash stores episodes, facts and working context in Postgres with pgvector, and exposes them to any MCP client over SSE. It is a single Go binary plus a database, and it needs an embeddings model and a reasoning model to consolidate anything.
- Who is it for?
- Adopt alash3al/stash if you already run Postgres, want agent memory to stay inside your own network, and are willing to pay for embeddings and consolidation calls on an OpenAI-compatible endpoint. Skip it if you have no MCP-capable client, no Postgres, or no tolerance for a pipeline that needs a chat model to turn raw episodes into facts.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 108 days ago.
- What is it written in?
- Mainly Go, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem Stash addresses: agents that start every session from zero
A chat model has no memory between sessions unless something outside it supplies one. Stash is that something. The README frames the problem bluntly ("Your AI has amnesia. We fixed it.") and then describes what it stores: episodes, facts and working context, kept in Postgres. The intended user is someone running an MCP-compatible agent (Claude Desktop, Cursor, Windsurf, Cline, Continue are the clients named) who wants recall to survive a restart without shipping their conversation history to a vendor.
The project is Apache-2.0 licensed and written in Go. It is not archived, and the last push to the default branch was on 2026-06-14, so treat it as a project that has moved recently rather than one with a long track record. The release list shows v0.2.9 through v0.2.11 between 2026-05-01 and 2026-05-29, which suggests a still-moving 0.2 line. Nothing in the README claims production hardening, and the version numbers are honest about that.
Episodes become facts: the nine-stage consolidation pipeline
The mechanism is a pipeline, not a vector store with a wrapper. Raw observations arrive as episodes. A nine-stage consolidation process turns them into structured knowledge: facts, relationships, causal links, patterns, contradictions, goal tracking, failure patterns, and hypothesis verification. The README states that each stage only processes data new since the last run, which is what keeps repeated consolidation from redoing the same work.
Two models are required for this. An embeddings model vectorizes content into Postgres through pgvector, and a chat-capable reasoning model does the extraction. The go.mod file confirms the shape of the stack: pgx for Postgres, pgvector-go for the vector type, mcp-go for the server, goose for migrations, and the OpenAI Go client for model calls. The reasoning model is configurable through STASH_REASONER_MODEL, and the .env.example lists gpt-4o-mini and gpt-4o among the OpenAI options. Because the base URL is configurable, any OpenAI-compatible endpoint works, including a local Ollama instance.
There is a decay mechanism too. STASH_DECAY_FACTOR defaults to 0.95 and STASH_EXPIRY_THRESHOLD to 0.1 in the compose file, which implies facts fade over time unless reinforced. The README does not explain the scoring formula behind those numbers, so the practical effect of changing them is something you would have to observe on your own data.
Installing Stash with Docker Compose and connecting an MCP client
The README's quick start is four commands. You clone the repository, copy the environment template, edit it with an API key and model names, and bring the stack up. The compose file builds the image from the Dockerfile and starts Postgres with the pgvector extension alongside the Stash server.
git clone https://github.com/alash3al/stash.git
cd stash
cp .env.example .env # edit with your API key + model
docker compose upThe .env file needs at least STASH_POSTGRES_DSN, STASH_VECTOR_DIM, STASH_MAX_RESULT_SIZE, STASH_OPENAI_API_KEY, STASH_OPENAI_BASE_URL, STASH_EMBEDDING_MODEL and STASH_REASONER_MODEL. The vector dimension must match the embedding model: the template notes 1536 for text-embedding-3-small and 3072 for text-embedding-3-large. Getting that wrong is the kind of mistake that shows up later as failed similarity queries rather than a startup error.
STASH_OPENAI_API_KEY=your-openai-api-key-here
STASH_OPENAI_BASE_URL=https://api.openai.com/v1
STASH_EMBEDDING_MODEL=text-embedding-3-small
STASH_REASONER_MODEL=gpt-4o-mini
STASH_VECTOR_DIM=1536Once the containers are healthy, the MCP server listens over SSE at http://localhost:8080/sse. The compose command that produces this is mcp serve with --host 0.0.0.0, --port 8080 and --with-consolidation, so background consolidation runs in the same process. Pointing a client at that URL is a small JSON edit. For Cursor, the README gives ~/.cursor/mcp.json:
{
"mcpServers": {
"stash": {
"url": "http://localhost:8080/sse"
}
}
}Claude Desktop, Windsurf and OpenCode use the same URL with slightly different file locations and, for OpenCode, a type field set to remote. After connecting, the README points to docs/GETTING_STARTED.md for the init, remember and recall calls that verify the loop works end to end. A fully local variant using Ollama is documented separately in docs/LOCAL_OLLAMA.md, where the example uses nomic-embed-text for embeddings, qwen2.5:3b as the reasoner, and STASH_VECTOR_DIM=768.
Where Stash is the wrong tool
Stash requires Postgres. There is no SQLite mode, no embedded store, and no file-based fallback in the compose file or the environment template. If you want a desktop agent to keep memory on a laptop without running a database, this is the wrong shape of project, and the hosted version at usestash.io is the answer the README itself offers.
It also requires an MCP client that speaks SSE. The README says it works with any agent that supports MCP over SSE, which is a narrower set than "any MCP client." Clients that only launch stdio servers need a bridge that the documentation does not describe.
The consolidation pipeline is the second constraint. Turning episodes into facts, relationships and contradictions is a model call, so every consolidation run costs tokens against whatever endpoint you configured, and the quality of the extracted knowledge is bounded by that model. The README does not document rollback for a consolidation run that produces bad facts, and it does not describe how to inspect or correct a wrong relationship once written. With a local Ollama reasoner the cost disappears but the extraction quality depends on the model you pick.
Finally, the cloud version is explicitly a separate codebase. The README states it is written from scratch and shares no code with this repository, and that feature sets differ in both directions. Migration between the two is not described, so choosing self-hosting is a commitment rather than a staging step.
How Stash differs from rolling your own vector store
The obvious alternative is a general-purpose vector database such as pgvector on its own, or a dedicated store, with your agent writing and querying embeddings directly. The difference is what sits between storage and retrieval. A raw vector store gives you similarity search over whatever text you embedded; you decide what to embed, when to re-embed, and how to reconcile two near-identical memories. Stash adds the consolidation stages and the decay parameters, so deduplication and expiry are handled by the pipeline. The compose file exposes STASH_CONSOLIDATION_DEDUP_THRESHOLD at 0.95 and STASH_CONSOLIDATION_SIMILARITY_THRESHOLD at 0.85, which are the knobs for that behaviour.
The trade-off is control. With pgvector alone you know exactly what is in the table and why. With Stash, facts are derived, and derived data can be wrong in ways that are harder to trace than a bad embedding. If your memory needs are simple (store this, fetch the nearest matches), the pipeline is overhead. If they are not, writing the deduplication and contradiction logic yourself is a real project, and the nine stages are the argument for using someone else's.
Maintenance, upgrade cost and the Apache-2.0 licence
The last push was on 2026-06-14, and the most recent release is v0.2.11 from 2026-05-29. That is a pre-1.0 project, which means configuration keys and defaults can move between minor versions. The upgrade path is Docker Compose plus migrations: goose is in the dependency list and the README says migrations run as part of the one-command startup, so pulling a new image and restarting should apply schema changes. The README does not document a downgrade path, and it does not describe how to take a consistent backup before an upgrade, so that is on you to arrange around the postgres_data volume.
Ongoing cost is model usage, not licensing. Embeddings are charged per token on a hosted endpoint, and consolidation calls a chat model, so a busy agent with frequent writes will spend more than a quiet one. The STASH_CONSOLIDATION_WINDOW default of 168h and STASH_CONSOLIDATION_BATCH_SIZE of 100 in the compose file bound how much work each run picks up, which is also how you bound the spend.
The licence is Apache-2.0, which permits commercial use and modification and requires that you keep the licence and notice files. It includes a patent grant. Nothing in the repository suggests dual licensing or a contributor agreement that would change that. This is a description of the licence text, not legal advice; if you are embedding Stash in a product, have counsel read it.
Editorial conclusion
Adopt alash3al/stash if you already run Postgres, want agent memory to stay inside your own network, and are willing to pay for embeddings and consolidation calls on an OpenAI-compatible endpoint. Skip it if you have no MCP-capable client, no Postgres, or no tolerance for a pipeline that needs a chat model to turn raw episodes into facts. Before committing, verify three things: that your embedding model's dimension matches STASH_VECTOR_DIM, that your client speaks MCP over SSE rather than stdio only, and that the consolidation stage behaves acceptably on your own data, since the README describes nine stages without documenting rollback for a bad consolidation run.
Frequently asked questions
How do I install alash3al/stash?
Clone the repository, copy .env.example to .env and fill in your API key and model names, then run docker compose up. The compose file starts Postgres with pgvector and the Stash MCP server together, and migrations run as part of that startup.
Does alash3al/stash need an OpenAI API key?
It needs an OpenAI-compatible endpoint, not OpenAI specifically. The .env.example shows OpenAI, OpenRouter and Atlas Cloud as options, and docs/LOCAL_OLLAMA.md covers a fully local setup using Ollama with nomic-embed-text and qwen2.5:3b.
Which MCP clients work with alash3al/stash?
Any client that supports MCP over SSE, pointed at http://localhost:8080/sse. The README gives configuration examples for Cursor, Claude Desktop, OpenCode and Windsurf, and names Cline, Continue, OpenAI Agents, Ollama and OpenRouter among the compatible setups.
What does the STASH_VECTOR_DIM setting have to match?
It must match the output dimension of your embedding model. The environment template notes 1536 for text-embedding-3-small, 3072 for text-embedding-3-large, and the Ollama guide uses 768 with nomic-embed-text.
Is the hosted version at usestash.io the same code as alash3al/stash?
No. The README states the cloud version is written from scratch and shares no code with this repository, and that feature sets differ in both directions.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/alash3al-stash)