Model or dataset
Astro-Han/karpathy-llm-wiki avatar
Astro-Han/karpathy-llm-wiki

karpathy-llm-wiki: an Agent Skills package that compiles sources into a wiki your coding agent maintains

Agent Skills-compatible LLM wiki for Claude Code, Cursor, and Codex. Build a Karpathy-style knowledge base from raw sources, citations, and linting.

2,391 stars286 forksPythonMIT

At a glance

What is it?
Astro-Han/karpathy-llm-wiki turns Karpathy's LLM wiki idea into one installable Agent Skills skill: raw sources in, curated markdown pages out, with citations and linting. It is for people who want knowledge to compound instead of being re-derived at every query.
Who is it for?
Adopt it if you already keep notes as markdown, you use an Agent Skills tool, and you want ingest-time synthesis rather than query-time retrieval. Skip it if you need vector search over a large corpus, per-source provenance hashes, or a hosted service with a UI; the README lists those as deliberate omissions.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 69 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem: knowledge that gets re-derived instead of compiled

Most retrieval setups answer the same question from scratch. The chunks are embedded, the query is matched against them, and the model rebuilds a synthesis that already existed last week. Nothing accumulates. The README frames the alternative as an LLM wiki: a knowledge system where the model maintains structured wiki pages instead of re-searching raw documents on every question. New sources are compiled into durable markdown pages, cross-references are updated over time, and answers cite the pages that already hold the synthesized knowledge.

The audience is narrow and specific. You need to be someone who reads papers, blog posts and documentation faster than you can file them, and who already prefers plain markdown over a proprietary notes app. The skill is packaged for Claude Code, Cursor, Codex CLI, OpenCode and any other tool that follows the Agent Skills standard, so the practical requirement is that your agent can load a skill directory. If you work entirely inside a browser-based chat product with no filesystem access, this is not the shape of tool you are looking for.

Ingest, query and lint: the three operations the skill exposes

The README describes three operations. Ingest collects a source into raw/, triages it, then creates or updates wiki articles, or merely logs it when nothing is new. Query searches the wiki and answers with citations linking back to markdown pages. Lint checks index integrity, links and wiki health, producing auto-fixes plus a report of issues.

The layout is the mechanism. Sources land in raw/ as immutable material, one file per source under a topic directory, named with a date prefix. Compiled pages live in wiki/, also organized by topic, with a global index.md as the table of contents and an append-only log.md recording operations. The README's own tree shows a source at raw/topic/2026-04-03-source-article.md and a concept page at wiki/topic/concept-name.md.

What makes this different from a script that converts documents to markdown is the fan-out. According to the README, each new source can update multiple pages, strengthen cross-references, and record contradictions. That is the compounding claim, and it is also where the design carries risk: an ingest that touches several pages is an edit operation across your notes, so the wiki's consistency depends on the agent following the skill specification in SKILL.md rather than on any enforced schema.

Installing karpathy-llm-wiki and ingesting your first source

The README gives one install command for tools that support the Agent Skills standard, including Claude Code, Cursor and OpenCode.

bash
npx add-skill Astro-Han/karpathy-llm-wiki

Codex CLI is handled differently. The compatibility table says to copy the skill directory into the tool's own location rather than running the installer.

text
.agents/skills/karpathy-llm-wiki/

For any other tool, the README says to copy SKILL.md, references/ and scripts/ into that tool's skill directory. There is no build step and no server to start; the repository is markdown plus Python scripts and tests.

Once the skill is loaded, the first real use is a plain instruction to your agent, not a command. The README's example is to give it a URL, a file or pasted text. You should expect the source to appear under raw/ and one or more pages to appear or change under wiki/, with an entry appended to wiki/log.md. If the agent answers with a summary in the chat and writes no files, the skill has not been picked up; check that the directory landed where your tool looks for skills.

Asking a question is the same interface. The README's sample prompt is a natural-language question about a topic, and the documented output is an answer with citations that link back to your markdown pages. Linting is triggered the same way, by asking the agent to lint the wiki, and the README says it checks for broken links, missing index entries and stale cross-references.

Why there is no vector search, and when that choice breaks

The design boundaries section is the most opinionated part of the repository, and it is worth reading before you install anything. The author lists what was deliberately not built, citing three months of production logs and a survey of the ecosystem that includes LLM Wiki v2, llm-wiki-compiler, OKF and the agent-memory literature.

Vector and graph search are on that list. The stated reasoning is that at 50K to 100K tokens of curated wiki, grep and read are more reliable, and search tooling should be added only when recall measurably degrades. That is a defensible position for a personal knowledge base and a real limitation for anything larger. If your corpus is thousands of documents, or you need semantic recall across material you have not yet compiled, this skill is the wrong layer. It assumes the compilation step already happened.

Other omissions follow the same logic. Source-hash freshness tracking is skipped because raw/ is immutable, so hashes would guard against events that cannot happen. Persisted line-number citations are skipped because every observed fidelity error was a value absent from the source, which a whole-file grep catches. Numeric confidence scores are rejected as false precision. Per-article review dates are rejected because nobody can predict at compile time how fast a domain moves; maintenance runs through whole-wiki lint instead. Retract and bad-source machinery is deferred until it is actually needed, to be handled manually.

The honest reading is that several of these are bets on the wiki staying small and the sources staying trustworthy. The README says the retract case has not happened yet. When it does, you will be editing pages by hand, and the append-only log will be the only record of what changed.

karpathy-llm-wiki versus RAG and versus a hand-edited Obsidian vault

The README's own comparison table sets the skill against RAG. The difference is where knowledge lives and when synthesis happens. In RAG, knowledge lives in raw chunks and embeddings, and synthesis happens at query time; it is good for broad retrieval across large corpora. In an LLM wiki, knowledge lives in curated markdown pages, and synthesis happens during ingest and maintenance; it is good for compounding knowledge, summaries and durable cross-links. Neither is strictly better. RAG degrades gracefully as the corpus grows because retrieval is the whole design. This skill trades that scalability for pages a human can read and correct.

The comparison against a normal personal wiki is the one the README answers directly in its FAQ. An LLM wiki is maintained by the model, which updates summaries, cross-links, index entries and contradictions as new material arrives; a normal personal wiki depends on manual editing. So the real alternative is not another package, it is the vault you already have. If you use Obsidian and you enjoy maintaining the index and the links yourself, this skill removes work you were not trying to avoid, and it adds a dependency on an agent doing edits correctly. Where it wins is the case where your notes have outgrown your willingness to file them, and the cross-references you never wrote are exactly the ones you keep needing.

Maintenance, licence and what the repository does not cover

The last push to the repository was on 2026-07-23, and the repository is not archived. There are no retrieved releases, so there is no version number to pin and no changelog to read; upgrading means re-running the install command or re-copying the skill directory, and checking SKILL.md for changes yourself.

The licence is MIT, which is permissive and imposes no copyleft obligation on the wiki content you produce. That matters because the output of this tool is your own writing, stored in your own repository under raw/ and wiki/. Nothing in the README suggests the skill phones home or stores your sources anywhere other than the project directory. It also means there is no vendor to escalate to. If an ingest mangles a page, git is your undo, which is a reasonable argument for keeping the wiki in version control from day one.

The README states that the workflow is based on a real knowledge base with 94 articles and 99 sources, maintained daily since April 2026, and that the repository includes examples, templates and a design spec. The examples/ directory contains sample wiki pages, source files and operation logs. Those are the closest thing to a test suite for the workflow itself, and they are the first thing to read if you want to judge output quality before pointing the skill at your own material.

Editorial conclusion

Adopt it if you already keep notes as markdown, you use an Agent Skills tool, and you want ingest-time synthesis rather than query-time retrieval. Skip it if you need vector search over a large corpus, per-source provenance hashes, or a hosted service with a UI; the README lists those as deliberate omissions. Before committing, run the lint operation on a copy of your own notes and check whether the index and cross-links it produces match how you actually read your material.

Frequently asked questions

What is the Karpathy LLM wiki idea?

It is a knowledge system where the LLM maintains structured wiki pages instead of re-searching raw documents on every question, so new sources are compiled into durable markdown pages and answers cite those pages. The README credits the idea to Karpathy and links to his gist.

How do I set up karpathy-llm-wiki?

Run npx add-skill Astro-Han/karpathy-llm-wiki for tools that support the Agent Skills standard, including Claude Code, Cursor and OpenCode. For Codex CLI the README says to copy the skill into .agents/skills/karpathy-llm-wiki/, and for other tools to copy SKILL.md, references/ and scripts/ into the tool's skill directory.

How do I use karpathy-llm-wiki after installing it?

You give the agent a URL, a file or pasted text to ingest, then ask questions and receive answers with citations to your markdown pages. Linting is triggered the same way, by asking the agent to lint the wiki.

What is the difference between karpathy-llm-wiki and RAG?

In RAG, knowledge lives in raw chunks and embeddings and synthesis happens at query time. In this skill, knowledge lives in curated markdown pages and synthesis happens during ingest and maintenance, which the README describes as better for compounding knowledge and durable cross-links but not for broad retrieval across large corpora.

Official sources

  1. Astro-Han/karpathy-llm-wiki on GitHub
  2. Issues
  3. License: MIT
  4. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/astro-han-karpathy-llm-wiki.svg)](https://hysenlabs.com/projects/astro-han-karpathy-llm-wiki)