Model or dataset
nashsu/llm_wiki avatar
nashsu/llm_wiki

nashsu/llm_wiki: a desktop app that compiles your documents into an Obsidian-compatible wiki

LLM Wiki is a cross-platform desktop application that turns your documents into an organized, interlinked knowledge base — automatically. Instead of traditional RAG (retrieve-and-answer from scratch every time), the LLM incrementally builds and maintains a persistent wiki from your sources。

20,079 stars2,268 forksTypeScriptNOASSERTION

At a glance

What is it?
LLM Wiki turns the Karpathy llm-wiki pattern into a Tauri desktop application with a Rust chat agent, a persistent ingest queue and an optional LanceDB vector index. It fits if you want a browsable, versionable knowledge base rather than a query-time RAG pipeline.
Who is it for?
Adopt LLM Wiki if you want a durable, human-readable artifact you can open in Obsidian and diff in git, and you accept that every source has to be ingested through an LLM before it is useful. Skip it if your only goal is answering questions over a large corpus you will never read page by page, or if you cannot run a vision model for PDFs whose meaning lives in their figures.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Yes. The repository last received commits 1 day ago.
What is it written in?
Mainly TypeScript, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

DEEP OPEN-SOURCE ANALYSIS

The problem LLM Wiki addresses, and who it is for

Retrieval-augmented generation re-derives an answer on every query. You embed a corpus, retrieve the top-k chunks, and ask a model to synthesize something from them. Nothing accumulates. Ask the same question next month and the system starts from zero, with no memory of what it concluded last time and no artifact you can read on its own.

LLM Wiki takes the opposite position. The README describes it as a cross-platform desktop application that "incrementally builds and maintains a persistent wiki" from your sources, so that knowledge is "compiled once and kept current, not re-derived on every query." The output is a directory of Markdown pages with YAML frontmatter, [[wikilink]] cross-references and an index.md catalog. That directory is also an Obsidian vault, which the README lists as a deliberate compatibility goal.

The audience is narrow and specific. It suits researchers, writers and engineers who already keep notes and want the note-taking to be done for them. It suits people who read their knowledge base rather than only querying it. It does not suit someone who wants a chatbot over a document dump and has no intention of opening the resulting pages.

Ingest, index.md and the three-layer architecture

The design comes from Andrej Karpathy's llm-wiki.md gist, which the README credits as the foundational methodology. The repository keeps a copy of the source document at llm-wiki.md in the project root. Three layers are preserved: raw sources are immutable, the wiki is LLM-generated, and a schema layer holds rules and config. Three operations are preserved as well: ingest, query and lint.

What nashsu added is a two-step chain-of-thought ingest. The documentation states that the model analyzes first and then generates wiki pages, with source traceability and an incremental cache. The practical effect of the cache is that re-ingesting a folder does not reprocess files that have not changed. A separate persistent ingest queue serializes the work, and the README lists crash recovery, cancel, retry and progress visualization among its features. Serial processing is a real constraint: a large folder import will not parallelize across documents.

Parsing covers PDF, Office documents, EPUB and MOBI, Org mode, images, media, web clips and batches of URLs. PDF handling can run through a built-in path, a cloud service or a local MinerU instance. Images embedded in PDFs are extracted and captioned by a vision model, and those captions surface in image-aware search results with lightbox preview and jump-to-source. That last feature is the one that most clearly separates this from a text-chunking pipeline, and it is also the one that costs the most, since every figure needs a vision call.

Above the pages sits a knowledge graph built from four signals: direct links, source overlap, Adamic-Adar and type affinity. Louvain community detection runs over that graph to surface clusters with cohesion scores, and a Graph Insights view reports surprising connections and knowledge gaps. Vector search is optional and backed by LanceDB through any OpenAI-compatible embedding endpoint.

Installing LLM Wiki and running a first ingest

The repository is a Tauri application: a Vite and React frontend under src/, a Rust backend under src-tauri/, and a bundled MCP server in mcp-server/. The README points at an Installation section, but the cleaned text available here does not include prebuilt binary download links, so the reproducible path is to build from source. The package.json defines the scripts.

Install dependencies and start the desktop development build:

bash
npm install
npm run tauri dev

The build:desktop script is the one that matters for a production bundle, because it installs and compiles the MCP server before building the frontend:

bash
npm run build:desktop

Once the app is running, the workflow the README describes is: create a project, point a source folder at your documents, and let the ingest queue process them. The application watches raw/sources/ for external changes and keeps ingest and delete cleanup in sync, so dropping a file into that folder is enough. Scenario templates (Research, Reading, Personal Growth, Business, General) pre-configure purpose.md and schema.md if you would rather not write them yourself.

For a first real use, import a small folder of PDFs you have already read. Watch the activity panel, which reports file-by-file progress, then open a generated page and follow its source links back to the original. That round trip is the thing to validate before you commit a large corpus.

If you want the wiki reachable from an external agent, the app exposes a local JSON API on 127.0.0.1:19828 and ships an MCP server for hybrid search, file read, graph traversal and source rescan. A companion skill repository installs into Claude Code or Codex with a single npx skills add command, per the README.

Where LLM Wiki is the wrong tool

The cost model is inverted relative to RAG. In a retrieval system, ingestion is cheap and queries are expensive. Here, ingestion is the expensive step: every document goes through an LLM, and every embedded image goes through a vision model. A thousand-page corpus is a thousand-page bill before you have asked a single question. If your corpus is large and you only ever need a handful of answers out of it, a conventional retrieval pipeline will be cheaper and faster to stand up.

Quality is bounded by the model you route to ingest. The README lists flexible per-project model configuration and independent routing for Chat and Ingest, which is useful, but it also means a weak ingest model produces a weak wiki, and the errors are written into pages that later queries treat as ground truth. The lint operation exists to catch drift, but the README does not describe an automated rollback for a bad ingest run.

There is also a single-user assumption baked into the shape of the thing. Projects are exported and imported as complete archives, state is persisted locally, and the HTTP API binds to loopback. There is no documented multi-user mode, no server deployment story, and no access control beyond the fact that the port is local. Teams looking for a shared knowledge service should look elsewhere.

The licence is the quietest risk. Repository metadata reports NOASSERTION, and the README's licence section is referenced but not reproduced in the text available here. Read the LICENSE file in the repository root before you build anything on top of it.

LLM Wiki versus a RAG stack

The honest comparison is not feature-by-feature, because the two systems produce different artifacts. A RAG stack produces answers. LLM Wiki produces a document tree that happens to make answers easier.

Concretely: with a vector store and a retrieval chain, you control chunking, you can swap embedding models without touching the corpus, and you can re-index from source at any time at low cost. With LLM Wiki, the wiki pages are the index. They are generated text, not derived embeddings, so re-generating them means re-running the model over the sources. The architecture diagram in the repository shows the three-layer split, and the README is explicit that the wiki is the LLM-generated layer rather than a cache.

In exchange you get something a retrieval pipeline does not give you: a browsable artifact with cross-links, a chronological log.md of operations, and a knowledge graph you can inspect for gaps. The graph signals (Adamic-Adar, source overlap, type affinity) are computed over the generated pages, so they reflect what the model decided was related, not what a similarity metric found.

There is a middle path worth naming. LLM Wiki supports an optional LanceDB vector index and a Read Sources Only mode that answers exclusively from original imported material. That mode is closer to classical RAG, and it exists precisely because generated pages can drift from their sources.

Maintenance, upgrades and what the licence metadata leaves open

The last push to the default branch was on 2026-08-25, and v0.6.11 was released the same day. Preceding releases landed on 2026-08-21 and 2026-08-14, so the cadence over that window is roughly weekly. The repository is not archived. That is the extent of what the available facts support about maintenance; nothing here establishes a long-term support commitment or a deprecation policy.

Upgrade cost is the interesting part. The wiki directory is plain Markdown with YAML frontmatter, and the README states it works as an Obsidian vault. That means the generated content is not locked inside the application: you can read it, version it, and migrate it with ordinary file tools even if you stop using the app. What is not portable is the rest of the state. Conversations, settings, review items and project config are persisted locally, and the documented way to move them is the project export and import archive. A version bump that changes the archive format would be felt at that boundary, and the README does not describe a migration path for older archives.

The MCP server is a second surface to track. The build:desktop script compiles it from mcp-server/, and it is distributed with the app, so its protocol compatibility moves with the main release rather than independently.

On licensing: the project metadata reports NOASSERTION, which means no standard licence identifier was detected. The README links to a licence section and the repository root contains a LICENSE file, but the terms are not reproduced in the text available here. This is not a legal opinion, and none should be inferred from it. The concrete step is to open LICENSE and read it, particularly if you plan to redistribute the application or bundle the MCP server into another product.

Editorial conclusion

Adopt LLM Wiki if you want a durable, human-readable artifact you can open in Obsidian and diff in git, and you accept that every source has to be ingested through an LLM before it is useful. Skip it if your only goal is answering questions over a large corpus you will never read page by page, or if you cannot run a vision model for PDFs whose meaning lives in their figures. Before committing, verify two things: that your chosen provider handles the chat and ingest routes independently as the settings UI implies, and that the LICENSE file in the repository root actually grants the terms you need, because the project metadata reports NOASSERTION rather than a named licence.

Frequently asked questions

Is LLM Wiki better than RAG?

It solves a different problem. RAG retrieves chunks and answers from scratch on every query, while LLM Wiki compiles knowledge into persistent Markdown pages that you can browse and keep. The README also offers a Read Sources Only mode that answers exclusively from original imported material, which is closer to classical retrieval.

What is the LLM wiki pattern?

It is the methodology described in Andrej Karpathy's llm-wiki.md gist, which the README credits as the project's foundation. It defines a three-layer architecture of raw sources, an LLM-generated wiki and a schema, with ingest, query and lint as the three operations, index.md as the catalog and log.md as the operation record.

How do I install LLM Wiki?

The repository is a Tauri application, so the reproducible path is building from source: run npm install, then npm run tauri dev for development or npm run build:desktop for a production bundle. The build:desktop script also installs and compiles the bundled MCP server.

How do I use LLM Wiki?

Create a project, point a source folder at your documents, and let the persistent ingest queue process them. The app watches raw/sources/ for external changes, and the README states that scenario templates pre-configure purpose.md and schema.md for research, reading, personal growth, business or general use.

Does LLM Wiki work with Obsidian?

Yes. The README lists Obsidian compatibility among the elements kept from the original pattern, and states that the wiki directory works as an Obsidian vault. Pages carry YAML frontmatter and use [[wikilink]] syntax for cross-references.

What is LLM Wiki by Andrej Karpathy?

Karpathy's contribution is the underlying llm-wiki.md gist, an abstract design pattern for using LLMs to incrementally build and maintain a personal wiki. LLM Wiki is nashsu's concrete desktop implementation of that pattern, with substantial extensions such as a vision-model image pipeline and a knowledge graph.

Official sources

  1. Issues
  2. nashsu/llm_wiki on GitHub
  3. README
  4. Releases
For maintainers

Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/nashsu-llm-wiki.svg)](https://hysenlabs.com/projects/nashsu-llm-wiki)
Community notes

Community notes