Model or dataset
taylorsatula/mira-OSS avatar
taylorsatula/mira-OSS

mira-OSS: a self-hosted agent with decaying memory and a single conversation thread

This is the public release of MIRA OS. Discrete memories decay through momentum loss, tools auto-configure when dropped into tools/ folder, and the system prompt composes from modular trinkets. I would like to think I've made an elegant brain-in-box. You load it and send cURL requests - it talks back, learns, and uses tools. Contributions welcome.

479 stars44 forksPythonAGPL-3.0

At a glance

What is it?
MIRA OS is a Python, Postgres-backed agent that keeps one conversation forever, collapses old turns into first-person memories that decay on use-days, and adapts its own system prompt from feedback instead of retraining weights.
Who is it for?
Adopt mira-OSS if you want an agent whose memory is a database you control and whose behaviour you can inspect as text, and if you accept a single-thread model with no new-chat escape hatch. Skip it if you need per-user weight fine-tuning, multi-tenant chat isolation, or a documented public API contract, because the repository does not describe one.
Can I use it commercially?
Yes, with strict conditions. AGPL-3.0 is a network copyleft licence: if people use a modified version over a network, for example as a hosted service, you must offer them its source code under the same licence.
Is it still maintained?
Yes. The repository last received commits 29 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 27, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The one-thread constraint, and who it is built for

Most agent frameworks treat a conversation as a disposable session. MIRA treats it as the only session. The README states plainly that "There is one conversation thread forever" and that "There is no functionality to start a new chat." That is not an oversight, it is the design premise: the author argues the constraint "forces facing the hard questions of how to build believable persistence within a framework (forward pass transformers) that is inherently ephemeral."

So the audience is narrow and specific. This is for someone who wants a personal, long-lived agent running on their own hardware, is comfortable with Postgres and Docker, and cares more about continuity of recall than about clean session boundaries. It is not a chatbot SDK. If you are building a product where each end user gets an isolated assistant, the single-thread model is working against you from the first line of code.

The origin story is honest about scope: the README says the project began as a recipe generator that incorporated cuisine preferences, and "10,000 scope creeps later MIRA is a comprehensive best-effort approximation of a continuous digital entity." The author calls it "my TempleOS." That framing tells you what to expect from the documentation and the API surface: a personal system, not a platform.

How memory actually decays: use-days, not calendar days

Memory in MIRA is not a vector store you prune by hand. The README says memories "decay via formula" and points to a file in the repository, lt_memory/scoring_formula.sql, as the definition of that formula. Memories survive by earning their keep through access, explicit references, links to other memories and entities, or real temporal relevance.

The detail worth pausing on is the unit of decay: it runs on use-days rather than calendar days. The README's example is that going on vacation does not make MIRA forget you. That is a meaningful design choice, and it is also a constraint. It means the clock only advances when you talk to the system, so an idle instance is not quietly eroding its own history. The same rule governs the adaptation loop described later, which fires every 7 use-days.

Recall is mostly passive. Memories are loaded into the context window through semantic similarity, entity hubs, memory traversal, filtering, and reranking, and the README notes that most recall happens before Mira generates a response, so the model does not have to notice a gap and decide to search. There is still a manual path: a memory_tool for explicit searches, creating or linking memories, and exact control over what gets recalled.

For long document-shaped content, memories are the wrong container. That is what domaindoc_tool is for: stable encrypted documents that do not decay, with section-level version history, sharing, pinned sections, and nested subsections. MIRA can expand, collapse and subsection those documents on its own, and when a section is idle it "closes the drawer" so the body stops occupying the context window while the title, summary and tree location remain visible.

Subcortical retrieval and the Peanut Gallery observer

Two small systems sit on opposite sides of the main response, and they are the most interesting part of the architecture because neither one answers the user.

Subcortical runs on every user turn, before Mira answers. It reads the current message, the recent conversation, and the memories already in context. It then resolves pronouns and fragmentary references into concrete search phrases, extracts named entities for graph lookup, decides which old memories are still relevant enough to keep, and classifies whether the turn is straightforward or needs heavier reasoning. The output is invisible to the user. It is a retrieval and triage pass that prepares what the primary model gets to think with.

The Peanut Gallery runs after turns complete, asynchronously, roughly every five turns. It reads the recent conversation alongside a ledger of MIRA's commitments, tool calls and tool results, and most of the time it stays quiet. When the conversation has measurably drifted, it can place a short-lived concern or coaching directive into the context window. It can also nudge MIRA to take one bounded initiative when an answer was technically correct but pushed too much conversational work onto the user. Guidance expires after two turns or when the segment collapses. Standard guidance is repaired silently; critical guidance can authorize MIRA to break the fourth wall and correct itself.

The split is worth naming plainly: Subcortical gathers context before the answer, Peanut Gallery checks the work after it. The trade-off is latency and token spend on every turn for a mechanism whose output you never see directly. Whether that pays off depends on how long your threads run.

Collapsing history into first-person memory traces

Active conversation history stays live while it is relevant. Older material collapses into first-person summaries, and the README is explicit about why: third-person summaries "created epistemic distance: Mira read them as logs about someone else, not memories of work It actually did." So a collapsed segment reads as "I debugged the IndexError in process_batch.py" rather than a report about an assistant.

Absolute timestamps replace relative ones for the same reason. "On Jan 8" is written instead of "Yesterday," because relative time becomes a lie the moment the sun sets. It is a small rule with a large effect on how coherent old memories feel when they resurface months later.

When a new segment summary is generated, the model sees the previous five summaries as context. That lets the new one reference what came before with what the README calls hazy continuity, for example building on earlier API work or continuing last week's recipe experiments. Each summary knows vaguely where it came from without carrying the full weight of everything before it.

The current summary format produces a 3-4 sentence memory trace, a two-sentence precis, a short display title, and a complexity score. The trace is for remembering; the other fields keep the conversation manifest useful without stuffing the full history back into the context window. There is also a safety valve: if an active conversation grows too large before it collapses naturally, MIRA compresses older messages into one rolling continuation brief while leaving the most recent turns untouched.

Installing mira-OSS and sending a first request

The repository ships a .env.example that documents the two required keys and tells you what happens without them. Copy it to .env and fill in the keys. The file states that if these are set, first boot runs non-interactively, and that if they are omitted you can run the interactive setup wizard instead.

bash
cp .env.example .env
# edit .env, then set at minimum:
# MIRA_ANTHROPIC_KEY=sk-ant-...
# MIRA_PROVIDER_KEY=gsk_...

MIRA_ANTHROPIC_KEY is the Anthropic API key used by the primary model, and MIRA_PROVIDER_KEY is a generic provider key for fast inference, defaulting to Groq. Both are marked required in the example file. Everything else in that file is optional: MIRA_PROVIDER_NAME, MIRA_PROVIDER_ENDPOINT and MIRA_PROVIDER_MODEL for a non-default provider, MIRA_ANTHROPIC_BATCH_KEY for batch operations, MIRA_KAGI_KEY for search, and MIRA_DB_PASSWORD for the database.

If you skip the keys, the example file gives the wizard path directly:

bash
docker compose run -it mira

The README describes the intended first use as loading the system and sending cURL requests: "You load it and send cURL requests - it talks back, learns, and uses tools." The repository does not print a literal example cURL command or a port, so check deploy/ and the FastAPI application entry point at main.py for the actual bind address before you write your client. The Python dependencies in requirements.txt show the stack you are standing up: FastAPI and Starlette behind the hypercorn ASGI server, psycopg with connection pooling, pgvector for embeddings, valkey for caching, and sentence-transformers for the mdbr-leaf-ir-asym embedding model. The requirements file also notes that the spaCy model is installed separately:

bash
python -m spacy download en_core_web_lg

Tools are the one part with a genuinely low-friction story. The README says tools "auto-configure when dropped into tools/ folder," so adding a capability is a matter of placing a file in that directory rather than editing a registry.

Text-Based LoRA: prompt adaptation instead of weight updates

The README is blunt that "You cannot retrain an LLM's weights per-user. But you can retrain Its prompts." The author calls the resulting loop Text-Based LoRA, and the name is accurate about the mechanism: behavioural adaptation through text manipulation rather than gradient descent.

The loop has four moving parts. After each conversation segment collapses, an assessment extractor compares the conversation against MIRA's behavioural contract and records alignment, misalignment, and contextual passes, each with the specific evidence and the prompt section involved. Those signals accumulate in Postgres. Every 7 use-days, a pattern synthesizer folds the signals into a descriptive user model covering what works, what fails, where friction appears, and what needs a check-in. A separate critic checks the synthesis before it is loaded into the system prompt.

The synthesis is evolutionary rather than a replacement: each run builds on the previous one, and patterns can be reinforced, refined, revised, settled, or allowed to go dormant. This is where the system prompt earns its description as composed from modular trinkets, and it is also where the sharpest limitation lives. Everything the agent learns about you is stored as prompt text and database rows. If your problem is that the base model lacks a capability rather than a preference, no amount of prompt synthesis will fix it.

Licence, upgrade cost, and what the repository does not settle

mira-OSS is released under AGPL-3.0. That matters more here than for a library you import, because the software is a network service: if you run a modified MIRA and let other people interact with it over a network, the licence's source-availability obligations are the question to put to your own counsel. This article is not legal advice, and the repository's license.txt is the document that governs.

Maintenance is visible in the release history. The most recent release is v2026.06.25, following v2026.05.29 and v2026.05.21, which is titled "Security Hardening & Reliability Sync." The last push to main was on 2026-09-01. Upgrades are not trivial in the way a stateless service is: your memories, domaindocs, and accumulated behavioural signals live in Postgres, so any schema change to the memory tables is a migration you have to plan rather than a container you can simply replace. The repository keeps tests/ and scripts/ directories, but the README does not document a supported upgrade path or a rollback procedure. Treat a version bump as a database event.

The dependency list also carries real operational weight. You are running Postgres with pgvector, valkey, a CPU-only PyTorch wheel pulled from a separate index, sentence-transformers, scikit-learn for TF-IDF candidate discovery in the memory linking pipeline, and spaCy with a large English model for entity extraction. That is a self-hosted stack with several stateful services, not a single binary.

One thing the documentation does not resolve is multi-user operation. Domaindocs have sharing, so collaboration on documents is contemplated, but the single conversation thread is a global constraint, and the README does not describe how two people sharing one instance would get separate memories or separate behavioural models.

Where it fits against a conventional RAG stack

The obvious alternative is a conventional RAG pipeline over a chat framework: LangChain or LlamaIndex wired to a vector store, with a new session per conversation and a retriever that runs when the model decides to call it. The difference in approach is not the vector store, since MIRA uses pgvector too. It is who initiates retrieval and what happens to old turns.

In a standard RAG setup, retrieval is a tool the model invokes, and history is either truncated or summarized into a third-person running summary. MIRA inverts both. Recall is mostly passive and happens before generation, and collapsed history becomes a first-person memory with a decay score rather than a neutral summary. The practical consequence is that a conventional stack gives you predictable, debuggable per-session behaviour and clean multi-tenant isolation, while MIRA gives you continuity across months at the cost of a global thread and a memory system whose state you have to reason about as a database.

If you want the second behaviour without the single-thread constraint, the honest answer from the README is that MIRA does not offer a configuration for it. If you want the first, a plain RAG stack is less machinery for the same retrieval.

Editorial conclusion

Adopt mira-OSS if you want an agent whose memory is a database you control and whose behaviour you can inspect as text, and if you accept a single-thread model with no new-chat escape hatch. Skip it if you need per-user weight fine-tuning, multi-tenant chat isolation, or a documented public API contract, because the repository does not describe one. Verify first that your Anthropic and provider keys work with the defaults in .env.example, that pgvector and valkey come up under deploy/, and that the decay formula in lt_memory/scoring_formula.sql matches the retention behaviour you want before you put real conversations into it.

Frequently asked questions

Can I start a new chat in mira-OSS?

No. The README states there is one conversation thread forever and no functionality to start a new chat, and that this constraint is deliberate. Older material is collapsed into memories rather than archived into a separate session.

How does mira-OSS decide which memories to forget?

Memories decay via a formula kept in lt_memory/scoring_formula.sql and survive by being accessed, explicitly referenced, linked to other memories and entities, or staying temporally relevant. Decay runs on use-days rather than calendar days, so idle periods do not erode memory.

What keys does mira-OSS need to boot?

The .env.example marks MIRA_ANTHROPIC_KEY and MIRA_PROVIDER_KEY as required, with the provider defaulting to Groq. If they are omitted, the file says to run the interactive setup wizard with docker compose run -it mira.

Does mira-OSS fine-tune the model on my conversations?

No. The README states you cannot retrain an LLM's weights per-user, and describes Text-Based LoRA instead: assessment signals accumulate in Postgres and a pattern synthesizer folds them into a descriptive user model that is checked by a critic and loaded into the system prompt.

What do I need to run mira-OSS myself?

The dependency list points to FastAPI and Starlette behind the hypercorn ASGI server, PostgreSQL with the pgvector extension, valkey for caching, and sentence-transformers, with the spaCy model installed separately. The .env.example references a docker compose workflow.

Official sources

  1. License: AGPL-3.0
  2. Project website
  3. README
  4. Releases
  5. taylorsatula/mira-OSS on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/taylorsatula-mira-oss.svg)](https://hysenlabs.com/projects/taylorsatula-mira-oss)