Model or dataset
swarmclawai/swarmvault avatar
swarmclawai/swarmvault

SwarmVault: a local-first LLM Wiki with a typed knowledge graph for agent memory

The local-first LLM Wiki: open-source knowledge graph builder, RAG knowledge base, and agent memory store. Built on Andrej Karpathy's pattern. An Obsidian alternative for personal knowledge management, AI second brain, and durable Claude Code / Codex / OpenClaw memory.

704 stars85 forksTypeScriptMIT

At a glance

What is it?
SwarmVault turns documents, code, transcripts and URLs into a markdown wiki plus a local knowledge graph that Claude Code, Codex and other MCP clients can query. It is a CLI and desktop app built on Andrej Karpathy's three-layer LLM Wiki pattern, and it runs offline with no API keys.
Who is it for?
Adopt SwarmVault if you already accumulate source material faster than you can organise it and you want that organisation to live in plain markdown on your own disk, with a graph and a search index you can inspect. Skip it if you need a hosted multi-tenant knowledge service, or if you expect an LLM to produce a finished wiki without review: the approval bundles and the candidates directory exist precisely because the first pass is not trustworthy.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 91 days ago.
What is it written in?
Mainly TypeScript, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem SwarmVault picks, and who feels it

Most people who work with language models end up with the same pile: PDFs, transcripts, half-finished notes, a few URLs they meant to read, and a repository or two they keep meaning to document. Retrieval over that pile is usually either a folder of text files pasted into a prompt or a vector store nobody can inspect. Neither produces anything durable. The README frames the gap directly, describing SwarmVault as the tool that "turns docs, code, transcripts, notes, and URLs into a durable markdown wiki plus a local graph you can inspect, query, and hand to agents."

The audience is narrower than "anyone using AI". It is the person running Claude Code or Codex against a codebase and wanting the agent to remember decisions across sessions; the researcher who wants a book companion that compounds instead of a chat log that resets; the writer maintaining a personal knowledge base who has tried Obsidian and found that the linking work still lands on them. The README also names business intelligence and code documentation as intended uses, which is a wide net, but the three-layer model underneath is the same in each case.

The project is explicit about its lineage. It implements the pattern Andrej Karpathy described in a public gist, and the comparison table in the README lists what the gist described against what SwarmVault implements: CLI commands for ingest, query and lint, a typed knowledge graph, an interactive viewer, agent context packs, and a vault doctor. That framing is honest about the debt and also about the ambition. The gist is a set of ideas; this is a toolchain.

Raw, wiki, schema: how the three layers actually move

The architecture is three directories and one file. Sources land in raw/ as immutable copies, which the README states SwarmVault reads but never modifies. Generated and hand-written markdown lands in wiki/. The conventions that govern the wiki live in swarmvault.schema.md, which the README describes as co-evolved by you and the model.

The machine-readable output sits beside the wiki rather than inside it. state/graph.json holds the knowledge graph, and state/retrieval/ holds the local search index. Search is described as hybrid: SQLite full-text merged with semantic embeddings, so a query does not require fitting every page into a context window. That is the mechanism that makes the "does it scale past 100 pages" question answerable at all, and it is a more concrete answer than most tools in this space give.

The part worth scrutinising is how the graph stays honest. Every edge carries a tag: extracted, inferred, or ambiguous. Contradiction detection flags conflicting claims, and lint --conflicts audits for them on demand. New concepts are not written straight into the wiki; they land in wiki/candidates/ first. compile --approve stages changes into reviewable approval bundles.

That is a deliberate tax on the user. A tool that generated edges and wrote them in would feel faster on day one and be useless by month three, because you would have no way to tell a quoted fact from a model's guess. Tagging every edge is the design decision that makes the graph worth traversing. It is also the reason a first compile produces a review queue rather than a finished artifact, and anyone expecting the latter will be disappointed.

Installing @swarmvaultai/cli and running a first vault

The README gives a thirty-second path through npm. Node 24 or newer is required according to the badge, and the package is published as @swarmvaultai/cli. The first command installs the CLI globally, the second initialises a vault in the current directory, ingests a local file, directory or public GitHub repository, compiles the wiki and graph, writes share artifacts, and opens the local graph viewer.

bash
npm install -g @swarmvaultai/cli
swarmvault quickstart ./your-repo

If you have nothing to point it at, the README offers a demo instead. Both commands are documented as working with no API keys, because the built-in heuristic provider runs locally and offline.

bash
swarmvault demo

After that first compile, the README lists the commands that matter most. The one I would run first is next, which is read-only and reports whether the vault needs initialising, ingesting, compiling, querying, reviewing or refreshing. doctor checks the vault's health, and query runs a search against the compiled wiki and index.

bash
swarmvault next
swarmvault query "What are the key concepts?"
swarmvault graph serve
swarmvault doctor
swarmvault candidate list

What you should see on disk is a fixed layout: raw/ for source copies, wiki/ for generated pages and reports, state/graph.json for the graph, and state/retrieval/ for the index. The first-run share artifacts are written to wiki/graph/share-card.md, wiki/graph/share-card.svg and wiki/graph/share-kit/. If those files exist, the pipeline completed. candidate list is where you go to see what the model wanted to add but has not been approved into the wiki.

There is a second install route the README documents for people who do not want Node on the machine: a desktop app for macOS, Windows and Linux that bundles its own runtime, downloaded from the project's site. For server-side use, the repository ships a Dockerfile that builds the workspace and exposes the MCP server over stdio, with the entry point running swarmvault mcp.

dockerfile
FROM node:24-alpine AS build
WORKDIR /app
RUN corepack enable && corepack prepare [email protected] --activate

Where the pipeline breaks down

The heuristic provider is the default and it is offline, which is a real advantage for anyone who cannot send source material to a cloud API. It is also the weakest extractor in the tool. The README recommends pairing it with a local model via Ollama "for sharper extraction", which is an admission that the no-key path trades quality for convenience. If your sources are dense or your domain vocabulary is unusual, the heuristic output will need more review, and the review queue is where that cost shows up.

The approval workflow is the second place expectations diverge from reality. compile --approve stages changes into bundles; it does not merge them. Nothing reaches wiki/ until a human accepts it. For a personal vault that is fine. For a team vault with several contributors, the README does not document how approval bundles are assigned, merged or resolved when two people stage conflicting edits. The git-backed workflow (--commit) and watch mode with git hooks are mentioned, but the README does not describe conflict handling for approval bundles. Treat concurrent editing as an open question rather than a solved one.

Scale has a documented ceiling too. compile --max-tokens trims output to fit a bounded window, which means the compiler is making a budget decision about what to keep. The README does not explain what gets dropped when the budget is exceeded. That is the kind of silent truncation that produces a wiki which looks complete and is not, and it is the first thing I would test on a corpus larger than a few hundred pages.

Finally, this is the wrong tool if you want a hosted service. Everything is files on your disk, and the value depends on you pointing the tool at material worth keeping.

SwarmVault against Obsidian and against a plain vector store

The obvious comparison is Obsidian, and the README invites it by listing obsidian-alternative as a topic. The difference is who does the linking. Obsidian gives you a markdown vault and a graph view, and the connections between notes are yours to make. SwarmVault generates the pages, the cross-references and the graph from raw sources, then asks you to approve them. If you enjoy the linking work and your notes are the primary artifact, Obsidian is the better fit and always will be. If your raw sources are the primary artifact and the wiki is downstream of them, the generated approach saves the part you were not going to do.

The second alternative is a vector store bolted onto an agent. A vector index answers similarity queries and nothing else: you cannot ask it which concepts depend on which, and you cannot inspect why a passage was retrieved. SwarmVault's graph commands (graph query, graph path, graph explain, graph callers) exist for traversal rather than similarity, and the README positions them as the answer to scale. The trade is that a graph has to be built and maintained correctly, which is why the edge tagging and contradiction detection carry so much weight in the design.

There is also a lighter option inside the project itself. The README points at a standalone schema template, templates/llm-wiki-schema.md, described as zero install and usable with any LLM agent. If you want the pattern without the toolchain, that file is the place to start, and the CLI is what you graduate to when the manual version stops scaling.

Maintenance, releases and what the MIT licence means here

The repository is not archived. Its last push was on 2026-06-30, roughly three months before this writing, and the most recent release listed is v3.20.0 from 2026-06-12, with v3.19.0 and v3.18.0 appearing earlier the same day. Three releases in one morning suggests a release train rather than a single large drop, and the workspace version in package.json is 3.21.0, ahead of the newest published tag. That gap is worth noting: the repository state and the released artifact are not always the same thing.

The engineering surface is heavier than a typical CLI. It is a pnpm workspace with three packages (cli, engine, viewer), Biome for lint and format, lefthook for git hooks, and a check script that runs typechecking plus four verification scripts covering release sync, published manifests, README parity and the ClawHub skill. There are also live smoke lanes for heuristic, Neo4j, Ollama, OpenAI and Anthropic, which implies the project tests against real providers rather than mocks. For an adopter, that means the upgrade path is likely to be gated by those checks rather than by guesswork, but it also means contributing requires the full workspace toolchain, not just a Node install.

Upgrade cost depends on which surface you use. The CLI and the on-disk layout (raw/, wiki/, state/) are the stable contract; the graph schema in state/graph.json is the thing to watch across versions. The licence is MIT, which permits commercial use and modification, but the repository does not state a policy on schema migrations, so back up state/graph.json before upgrading a vault you care about. That is a precaution, not legal advice.

Editorial conclusion

Adopt SwarmVault if you already accumulate source material faster than you can organise it and you want that organisation to live in plain markdown on your own disk, with a graph and a search index you can inspect. Skip it if you need a hosted multi-tenant knowledge service, or if you expect an LLM to produce a finished wiki without review: the approval bundles and the candidates directory exist precisely because the first pass is not trustworthy. Before committing, run swarmvault demo on a throwaway directory and open state/graph.json to see whether the node and edge shapes match the questions you actually ask.

Frequently asked questions

Does SwarmVault need API keys to run?

No. The README states that no API keys are required for the first run, because the built-in heuristic provider runs locally and offline. Cloud providers are optional, and the README suggests pairing the tool with a local model via Ollama for sharper extraction.

What files does SwarmVault create in my project directory?

The README lists raw/ for immutable copies of ingested material, wiki/ for generated markdown pages and reports, state/graph.json for the knowledge graph, and state/retrieval/ for the local search index. First-run share artifacts go to wiki/graph/share-card.md, wiki/graph/share-card.svg and wiki/graph/share-kit/.

How does SwarmVault keep the LLM from inventing facts in the wiki?

Every edge in the graph is tagged extracted, inferred or ambiguous, and contradiction detection flags conflicting claims. New concepts land in wiki/candidates/ first, and compile --approve stages changes into reviewable approval bundles rather than writing them straight into the wiki.

Official sources

  1. License: MIT
  2. Project website
  3. README
  4. Releases
  5. swarmclawai/swarmvault on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/swarmclawai-swarmvault.svg)](https://hysenlabs.com/projects/swarmclawai-swarmvault)