# llm-wiki-compiler: a knowledge compiler that turns raw sources into a citation-traceable wiki

> llmwiki is a TypeScript CLI that compiles papers, notes, PDFs and web pages into interlinked markdown pages with citations, review gates and hybrid retrieval. It is built for teams who want durable, auditable knowledge rather than query-time chunk search.

**atomicstrata/llm-wiki-compiler** — The knowledge compiler. Raw sources in, interlinked wiki out. Inspired by Karpathy's LLM Wiki pattern.

- Repository: https://github.com/atomicstrata/llm-wiki-compiler
- Website: https://llmwiki.atomicstrata.ai
- Stars: 2,156 · Forks: 229
- Language: TypeScript
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/atomicstrata-llm-wiki-compiler

## The problem llm-wiki-compiler addresses: knowledge rediscovered on every query

Most retrieval setups re-derive the same conclusions every time someone asks a question. The raw files sit in a directory, a chunking step splits them, an embedding index stores the pieces, and the model reassembles an answer that is thrown away when the session ends. Nothing accumulates. The README frames llmwiki as the opposite approach: instead of re-discovering knowledge from raw files at query time, compile it once into durable pages that accumulate structure, provenance, review state, and retrieval metadata over time.

That framing comes from Karpathy's LLM Wiki pattern, which the project cites as its inspiration. The target user is not someone who wants a chatbot over a folder of PDFs. It is someone who wants a wiki: a set of pages that cite their sources by file and line range, that carry a review status, that can be linted for broken links and stale content, and that an agent can read as a stable context pack instead of a pile of loose files.

The README is explicit about when not to reach for it. Do not use llmwiki as a general static-site generator, a heavy ontology database, or a replacement for ad-hoc search over fast-changing raw logs. That third exclusion matters. If your source material changes hourly and nobody will ever review a compiled page, the compile step is pure overhead.

## How the compiler works: two-phase extraction, typed pages, and a profile contract

The pipeline is two-phase. An LLM pass extracts concepts from the raw sources, then a generation pass produces typed pages. The default page types are concept, entity, comparison, and overview. Output is markdown with wikilinks, so the result is browsable in tools that understand that convention, and the README lists Obsidian among the project topics.

The more interesting mechanism is Configurable Lifecycle Profiles. A validated .llmwiki/profile.json is the single contract for typed entities, fields and directed relations, lifecycle states and transition evidence, multi-stage workflows, hash-pinned artifacts, connector bindings, content tiers and retrieval behavior. What makes this more than a config file is where the rules are enforced: the README states that relation, evidence, artifact, and human/agent gates are enforced by the write path rather than left as prompt conventions. Invalid profiles and writes that bypass a declared gate fail closed. Standing lint catches drift after the fact.

Backward compatibility is handled by absence. A project without .llmwiki/profile.json uses the built-in default concepts-and-queries profile and keeps pre-1.0 behavior. So the profile is opt-in, and adding one changes what the compiler will accept as a valid write.

Retrieval is hybrid: semantic chunk search, BM25 reranking, and wikilink graph expansion build evidence packs for queries and agents. The CLI can query the wiki, the viewer browses it, and llmwiki serve exposes the same capabilities over MCP to compatible agents.

## Installing llmwiki and compiling a first wiki

The package is published on npm as llm-wiki-compiler and installs a binary named llmwiki. Note the engine requirement in package.json: node >=24. If your runtime is older, the install will not give you a working CLI.

```bash
npm install -g llm-wiki-compiler
llmwiki --help
```

After the global install, llmwiki --help should list the available commands. The README does not spell out an ingest command line in the excerpt available, but it does document the profile and template workflow, which is the fastest way to see a real project structure. Scaffolding your own profile starts with one entity type at a time.

```bash
llmwiki profile init research --entity paper
llmwiki profile validate
llmwiki workflow list
```

profile validate checks the generated .llmwiki/profile.json against the schema, and workflow list shows the multi-stage workflows the profile declares. If validation fails, the compiler will refuse writes rather than fall back to permissive behavior.

A more complete starting point is a built-in template. The autosci pack is a research system with papers, ideas, experiments, manuscripts, evidence artifacts, workflows, and Crossref ingestion.

```bash
llmwiki template list
llmwiki template inspect autosci
llmwiki template init autosci
```

template inspect prints what the pack contains before you commit to it, and template init writes it into your project. Templates contain configuration and examples, never executable plugin code, which is worth knowing if you are evaluating one before running it.

Once pages exist, the local viewer is the quickest way to judge whether the output is useful.

```bash
llmwiki view
```

The README describes llmwiki view as a read-only browser UI with search, page metadata, graph exploration, source-freshness badges, and citation chips. If the citation chips point at the wrong line ranges, the extraction stage is the place to look.

## Where llmwiki breaks down

The compile step assumes source material worth compiling. The README's own exclusion list is the honest starting point: fast-changing raw logs are a bad fit, because the value of a compiled page comes from reuse, and a page that is stale before anyone reads it has no reuse value.

The fail-closed design cuts both ways. Gates enforced by the write path mean an invalid profile or a write that skips a declared gate is rejected rather than warned about. That is the right default for auditable knowledge, but it also means a profile mistake blocks ingestion entirely instead of degrading gracefully. There is no documented permissive mode in the README excerpt.

Cost is another constraint the documentation does not quantify. Two LLM passes over a corpus, plus optional judge-model citation support in llmwiki eval, means token spend scales with source volume and recompilation frequency. The README lists many providers, including Anthropic, Claude Agent SDK local login, OpenAI Codex CLI local login, OpenAI-compatible servers, Ollama, GitHub Copilot, Atlas Cloud, OrcaRouter, and local OpenAI-compatible runtimes, so you can push inference to a local runtime, but the trade-off between local model quality and citation accuracy is not addressed.

Finally, the profile system is a real conceptual load. Declaring entity schemas, typed relations, lifecycle state machines and transition requirements is design work, not configuration. Teams that want a wiki in an afternoon should start with the default profile and never open profile.json.

## How llmwiki differs from RAG pipelines and static-site generators

The obvious comparison is a conventional RAG stack built on a vector database. The difference is where the work happens. A RAG pipeline indexes chunks and reconstructs an answer per query, discarding the result. llmwiki compiles pages once and keeps them, with citations to source files and line ranges that llmwiki lint validates. Retrieval still exists, and it is hybrid, but it retrieves over compiled pages rather than raw chunks. The cost profile flips: RAG pays per query, llmwiki pays per compile and then amortizes.

Against a static-site generator, the difference is that llmwiki is not one, and the README says so. A site generator renders whatever you give it. llmwiki decides what pages should exist from the sources, types them, links them, and tracks their provenance and review state. The output happens to be markdown, which is why Obsidian is a plausible front end, but the generator comparison stops at the file format.

Against a formal ontology or graph database, llmwiki is lighter and more opinionated. Relations are declared in the profile and the graph is derived from markdown pages, with GraphML available as an export format. If you need SPARQL-style querying or strict OWL reasoning, this is not that tool, and the README's warning about heavy ontology databases is aimed at exactly this mismatch.

## Maintenance, upgrade path, and what the MIT licence means here

The repository is not archived, and the last push was on 2026-09-10. Releases have moved quickly: v1.0.0 on 2026-07-11, v1.1.0 on 2026-07-16, and v1.2.0 on 2026-09-10. package.json on the default branch already reads version 1.3.0, which suggests unreleased work in progress. For adopters, that cadence means reading CHANGELOG.md before upgrading rather than assuming patch-level compatibility, especially around the profile schema.

The backward-compatibility story is the strongest part of the upgrade case. The README states that a project without .llmwiki/profile.json uses the built-in default profile and preserves pre-1.0 behavior, so projects that never adopt CLP are insulated from profile schema changes. Projects that do adopt it inherit the profile validation surface, and the publish tooling includes a release:check-docs script that runs before publishing, which suggests the maintainers treat documentation drift as a release blocker.

The licence is MIT, which permits commercial use, modification, and redistribution provided the copyright notice and permission notice are included. That is the extent of what can be said here; whether MIT fits your organisation's policy on generated content and model providers is a question for your legal team, not for this article. Note that the licence covers the compiler, not the sources you feed it or the model outputs you store.

## Conclusion

Adopt llmwiki if your source material is stable enough to be worth compiling once and reusing, and if you need citations, review state and retrieval metadata attached to every page. Do not adopt it as a static-site generator, an ontology database, or a way to search fast-changing raw logs; the README says so directly. Before committing, run llmwiki profile validate against a scaffolded profile and check that the lifecycle gates you declare match how your team actually reviews work, because invalid profiles and gate-bypassing writes fail closed.

## FAQ

### What is the LLM wiki and what does it do?

It is a knowledge compiler that takes raw sources such as papers, notes, READMEs, transcripts, PDFs, images, or web pages and produces an interlinked markdown wiki. The output carries citations, review state and retrieval metadata so agents and humans can browse, query, lint and reuse it.

### Is LLM wiki a RAG system?

Not in the usual sense. A RAG pipeline re-discovers knowledge from raw files at query time, while llmwiki compiles the knowledge once into durable pages. It does include hybrid retrieval (semantic chunk search, BM25 reranking, wikilink graph expansion), but that retrieval runs over compiled pages rather than raw chunks.

### Is LLM wiki worth adopting?

It is worth it when source knowledge is stable enough to compile, review and reuse, and when you need citation traceability and auditability. The README advises against using it as a static-site generator, a heavy ontology database, or a replacement for ad-hoc search over fast-changing raw logs.

### What is wikillm?

The README does not describe a project called wikillm. The project covered here is llm-wiki-compiler, published on npm as llm-wiki-compiler and installed as the llmwiki command.

## Sources

- [atomicstrata/llm-wiki-compiler on GitHub](https://github.com/atomicstrata/llm-wiki-compiler)
- [License: MIT](https://github.com/atomicstrata/llm-wiki-compiler/blob/main/LICENSE)
- [Project website](https://llmwiki.atomicstrata.ai)
- [README](https://github.com/atomicstrata/llm-wiki-compiler/blob/main/README.md)
- [Releases](https://github.com/atomicstrata/llm-wiki-compiler/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/atomicstrata-llm-wiki-compiler
