Model or dataset
SamurAIGPT/llm-wiki-agent avatar
SamurAIGPT/llm-wiki-agent

LLM Wiki Agent: a coding-agent skill that writes and maintains your wiki

A personal knowledge base that builds and maintains itself. Drop in sources — Claude (or Codex/Gemini) reads them, extracts knowledge, and maintains a persistent interlinked wiki. Works with Claude Code, Codex, OpenCode, Gemini CLI. No API key needed.

3,591 stars413 forksPythonMIT

At a glance

What is it?
SamurAIGPT/llm-wiki-agent turns Claude Code, Codex, OpenCode or Gemini CLI into a maintainer of a persistent markdown wiki. It is a skill plus a folder convention, not a service, and that shapes both its appeal and its limits.
Who is it for?
Adopt it if you already work inside Claude Code, Codex, OpenCode or Gemini CLI and want your reading to accumulate as linked markdown instead of chat scrollback. Do not adopt it if you need a queryable vector index, a hosted service, or a team tool with access control: the wiki is files on disk and the agent is whatever CLI you opened.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 2 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What LLM Wiki Agent solves, and who it is actually for

Most note tools assume you will do the linking. You capture a paper, a meeting transcript, a chapter, and then you are the one who notices that the same company appears in three unrelated notes. LLM Wiki Agent inverts that. The README calls it "a coding agent skill": you drop documents into raw/, tell the agent to ingest them, and it produces a structured wiki under wiki/ with pages for sources, entities, concepts and syntheses. The promise in the README is that "every new source makes the wiki richer" and that "you never write it."

The audience is narrow and specific. You need an agent CLI already installed, because the project has no server and no API key of its own. It leans on Claude Code, Codex, OpenCode or Gemini CLI to do the reading and writing. If you live in one of those tools and your work involves accumulating sources over weeks (a research thread, a book, a set of customer calls), the fit is close. If you want a standalone app with a search box, this is not that.

The mechanism: raw/ in, interlinked markdown out

The data flow is visible in the repository layout. Sources land in raw/. The agent reads them and writes into wiki/, which the README documents as index.md (a catalog updated on every ingest), log.md (an append-only record of operations), overview.md (a synthesis revised on each ingest), plus sources/, entities/, concepts/ and syntheses/ directories. Entity and concept pages are "auto-created" and updated whenever a new source mentions them. Query answers are filed back as pages under syntheses/, so asking a question adds to the wiki rather than disappearing into a transcript.

Alongside wiki/ sits graph/. The README describes graph.json as persistent node and edge data that is SHA256-cached, and graph.html as a vis.js visualization you open in a browser. Edges come from [[wikilink]] syntax, with dotted edges for relationships the agent infers, and community detection clusters related topics. The caching detail matters more than it looks: re-running the graph build over unchanged pages should not force the agent to re-derive every relationship. Non-markdown inputs are converted at ingest time via markitdown, which the README lists as handling PDF, DOCX, PPTX, XLSX, HTML and EPUB among others. There is no separate conversion command to run. The agent does it when it reads the file.

Installing it and running a first ingest

There is no package to install and no service to start. The README's install section is a clone and a cd, and it explicitly says no API key or Python setup is needed because the agent CLI provides the model access.

bash
git clone https://github.com/SamurAIGPT/llm-wiki-agent.git
cd llm-wiki-agent

From there you open the directory in your agent. Each CLI reads a different config file, and the README lists which one: Claude Code reads CLAUDE.md plus .claude/commands/, while Codex and OpenCode read AGENTS.md and Gemini CLI reads GEMINI.md.

bash
claude      # reads CLAUDE.md + .claude/commands/ (slash commands available)
codex       # reads AGENTS.md
opencode    # reads AGENTS.md
gemini      # reads GEMINI.md

Once the session is open, the trigger is a plain instruction. The README's shorthand form takes a path under raw/.

code
ingest raw/papers/attention-is-all-you-need.md

What you should see afterwards is new markdown under wiki/: a summary page in sources/, and any entity or concept pages the agent decided the document warranted, with index.md and overview.md revised. In Claude Code the same operation has a slash-command form, /wiki-ingest, and the README notes the other three slash commands (/wiki-query, /wiki-lint, /wiki-graph) are Claude Code-specific while the natural language triggers work across all four agents. Plain English also works: the README gives "Ingest this paper: raw/papers/llama2.md" as an example. Batch and mixed formats are accepted, so a single instruction can name a .pptx and a .docx together.

The graph view and lint pass are where the design gets interesting

Two features carry more weight than their one-line descriptions suggest. The first is the lint pass. The README says it finds orphan pages, broken links, missing entity pages, and "data gaps with suggested sources to fill them", with an example suggesting a Mixtral paper because nothing covers mixture-of-experts. That is a maintenance loop most personal knowledge bases never get, because a human will not audit their own notes for structural holes. Whether the suggestions are good depends entirely on the model behind your CLI, not on this repository.

The second is contradiction flagging. The README states that when a new source contradicts an existing claim, it is flagged at ingest time rather than at query time. That ordering is the deliberate part. A contradiction surfaced during ingest lands next to the source that caused it, when you still remember the context. Surfaced later during a query, it arrives as an answer you have to trust. The trade-off is that ingest becomes slower and noisier: every new document is compared against everything already in the wiki, and on a large corpus that comparison is the agent's problem to manage, not the tool's. The README does not describe how the agent bounds that comparison, so on a wiki of several hundred pages the ingest step is the part most likely to degrade.

What it is not: no index, no service, no rollback story

The most common misreading is to file this under RAG. It is not a retrieval system. There is no embedding store, no chunking pipeline and no vector search in the repository layout or in pyproject.toml, whose dependencies are markitdown and tqdm. Retrieval happens because the agent reads markdown files, and that means the practical ceiling is how much the agent can hold in context at once. On a small wiki that is fine. On a large one, a query that needs to touch many pages will either miss material or cost a lot of tokens.

The second limitation is durability. The wiki is markdown on your disk, which is genuinely good for portability, but the agent edits files in place. The README does not document rollback, and there is no snapshot or versioning step described. If an ingest rewrites overview.md in a direction you dislike, your recovery path is whatever version control you put around the folder. Run the wiki under git from day one, or accept that a bad synthesis is permanent.

Third, the project is thin on its own runtime. pyproject.toml still carries the placeholder author "Your Name" and an empty description, and package-mode is false. That is a signal about what this is: a skill and a set of conventions more than a library with a stable Python API. Treat the .md config files and the folder layout as the interface, because those are what the README documents.

How it compares with Obsidian plus a plugin, or a RAG stack

The nearest alternative is Obsidian with a local model plugin, and the difference is who does the writing. In Obsidian you are the author; plugins help you search, link and visualize what you wrote. Here the agent writes the pages, and you review them. That is a real shift in effort, and it is also a shift in trust: an Obsidian vault is a record of your own judgement, while this wiki is a record of the agent's reading of your sources. If you want the first, this tool will feel like losing control of your notes. If the bottleneck is that you never get around to writing the notes at all, the trade is worth examining.

The other alternative is a conventional RAG stack: chunk your documents, embed them, query a vector store. That approach answers questions over a corpus without producing anything durable. LLM Wiki Agent produces files you can read, edit and diff, and it pays for that with context limits and slower ingest. Choose RAG when the corpus is large and the questions are ad hoc. Choose this when the corpus is something you are trying to understand over time and you want the understanding to persist outside a chat window.

Licence, dependencies and the supply-chain note in requirements.txt

The project is MIT licensed, which permits commercial use and modification provided the licence text is retained. Nothing in the repository suggests a dual-licence or a contributor agreement, so the usual MIT obligations apply. This is a description of the licence file, not legal advice.

The dependency file deserves a direct read. requirements.txt pins litellm to ~=1.83.10, with a comment stating that versions 1.82.7 through 1.82.8 were compromised in a supply chain attack in March 2026 and referencing issue 41 in the repository. That comment is the most operationally useful line in the whole project. It tells you the maintainers track upstream compromise, and it tells you to keep the pin rather than floating the dependency. networkx is pinned at ~=3.6.1, which lines up with the graph work. pyproject.toml separately declares markitdown[all] and tqdm, and requires Python >=3.10,<3.14, so a Python 3.14 environment is out of range.

Maintenance cost is mostly not in upgrades. The last push to the default branch was on 2026-09-08, and no releases are listed, so there is no changelog to read before updating. Updating means pulling the branch and re-reading the README, because the agent-facing contract lives in CLAUDE.md, AGENTS.md and GEMINI.md rather than in a versioned API. If you fork, keep those files under review: they are what your agent actually obeys.

Editorial conclusion

Adopt it if you already work inside Claude Code, Codex, OpenCode or Gemini CLI and want your reading to accumulate as linked markdown instead of chat scrollback. Do not adopt it if you need a queryable vector index, a hosted service, or a team tool with access control: the wiki is files on disk and the agent is whatever CLI you opened. Before committing, verify that your agent reads the config file it expects (CLAUDE.md, AGENTS.md or GEMINI.md), check that markitdown converts your real PDFs and slide decks cleanly, and decide where the raw/ folder lives, because every ingest copies its content into wiki/ pages that you then own.

Frequently asked questions

What is the LLM Wiki Agent and what does it do?

It is a coding agent skill: you place source documents in raw/ and instruct the agent to ingest them, and it writes a persistent interlinked markdown wiki under wiki/ with source, entity, concept and synthesis pages. The README describes it as a wiki that builds and maintains itself, with index.md and overview.md revised on every ingest.

Is LLM Wiki Agent a RAG system?

No. The repository layout and pyproject.toml show no embedding store, chunking pipeline or vector search; the dependencies are markitdown and tqdm. Retrieval happens because the agent reads markdown pages directly, which means the practical limit is how much the agent can hold in context.

What does LLM agent mean in this project?

Here the agent is the CLI you already run, not a component shipped by this repository. The README names Claude Code, Codex, OpenCode and Gemini CLI, each reading its own config file (CLAUDE.md, AGENTS.md or GEMINI.md), and states that no API key is needed.

Is LLM Wiki Agent worth using?

It depends on whether you already work in one of the four supported agent CLIs and want your sources to accumulate as linked files rather than chat history. The trade-offs are real: no rollback is documented, so the wiki needs version control around it, and ingest slows as the corpus grows because new sources are compared against existing claims.

Official sources

  1. Issues
  2. License: MIT
  3. README
  4. SamurAIGPT/llm-wiki-agent on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/samuraigpt-llm-wiki-agent.svg)](https://hysenlabs.com/projects/samuraigpt-llm-wiki-agent)