A book about Obsidian and an agent, shipped as two PDFs and a README
Obsidian + Claude Code: Rebuild Your Second Brain with AI · 橙皮书系列 · 用AI重建你的第二大脑
At a glance
- What is it?
- This repository contains no software. It is a self-published guide in Chinese and English, distributed as two PDF files, and its argument is a specific one: that a large language model should act as a compiler over a folder of Markdown files rather than as a retriever over an embedding index. Whether that is a good idea depends on how large your notes are and how much you trust the rewrite.
- Who is it for?
- Read this guide if you already keep notes in Obsidian and want a concrete method for making an agent maintain them, because the argument about treating a model as a compiler rather than a retriever is more useful than the specific tooling and the chapter structure is honest about what it assumes.
- Can I use it commercially?
- Not without permission. GitHub finds no licence file in the repository, and without a licence all rights are reserved by default: you may read the code but not reuse it. Check the README, or ask the authors, before using it.
- Is it still maintained?
- Yes. The repository last received commits 37 days ago.
- What is it written in?
- GitHub does not report a main language for this repository.
Answers come from the project's GitHub data, last synced on September 20, 2026, and from our analysis. They are not legal advice.
DEEP OPEN-SOURCE ANALYSIS
The repository is two PDFs, and that shapes how you should read it
Start with the file listing, because it is the most informative thing here. There are two PDF files at the top level, one Chinese and one English, both versioned v1.0.0. Then a readme in two languages, an assets directory holding what appears to be an animation, and a screenshots directory. There is no source directory, no build configuration, no package manifest and no plugin code, and the repository's primary language is not recorded at all. So this is a publishing project, not a software project, and the same author runs a series of them under a shared name, each one a free guide on an AI tool or an emerging technology. The readme says the PDFs are better downloaded than read through the repository's own viewer, which is a comment about how GitHub renders long documents rather than about the content. That constraint shapes the evaluation. You cannot run anything here, you cannot audit a technique by reading its implementation, and you cannot extend it except by writing your own version of whatever it describes. What you can do is read a method, check its internal consistency, and decide whether the trade-offs it describes match the ones you face. The screenshots directory and the animation suggest a book that cares about presentation, which is a differentiator in a series of free technical guides.
The core claim: a model as a compiler over files, not a retriever over vectors
The second of the three stated insights is the one that matters, and it is attributed to a named researcher rather than asserted on its own authority. The claim is that a large language model should be a compiler rather than a retriever, and the framing is that instead of building a retrieval system and having the model search your notes, you let the model maintain the knowledge base directly. The readme says this was validated at a scale of 400,000 characters. That number does real work in the argument and is worth reading carefully. Retrieval-augmented generation has a cost that grows with corpus size: an index to build, chunking decisions to get wrong, and retrieval that returns passages that are individually relevant and collectively misleading. A compiler model reads the files themselves and edits them, so there is no index to keep in sync and no recall problem to tune, and the failure mode changes from silently missing a relevant passage to visibly mangling a file. The cost is the mirror image. A retriever that misremembers has done nothing; a compiler that misremembers has rewritten your notes, and a wrong edit is more damaging than a missing one because it is harder to notice and it accumulates. The readme's other supporting observation is that several large projects, named in the first insight, chose Markdown files over a vector database for agent memory, and the author's conclusion from that is that Markdown is a native interface for agents. That is a reasonable reading of those choices, and the observation is doing rhetorical work in a book that wants you to stop building an index.
CLAUDE.md plus one index file per folder is the entire navigation model
The third insight is the most actionable and the most interesting, because it is a claim about how little you need. The readme says complex tool protocol setup is not required, and that a single instruction file at the vault root plus one index file in every folder is enough so the agent can move through the whole knowledge base efficiently. That is a strong claim and it is plausible for a specific reason: an agent reading files is not limited to a search index, it can list a directory, read an index, and follow references. The index files are the load-bearing part, and they are the part a reader has to write well. An index that lists each note with one line describing what is in it gives the agent a table of contents it can choose from, which is a different and more predictable primitive than similarity search. The cost is that the index is a second artefact per folder that must be maintained, and an index that is stale is worse than no index, because the agent trusts it. So the practical reading of this insight is that the structure you need is one file of standing instructions at the root and a maintained, human-written index in each folder, and the work is in keeping the indexes honest rather than in the setup. The chapter list suggests the guide covers this directly, in a section on designing the vault and another on having the agent maintain the knowledge base.
The chapter list is a syllabus, and it starts by naming the problem
The table of contents is short enough to read in full and tells you the argument's shape. There is a preface, then nine numbered sections. The first is titled as an information graveyard, and the framing in the readme's overview is that your notes application has become a graveyard of information and only AI can solve that. That is an aggressive opening and it is worth naming as rhetoric rather than accepting as diagnosis. A notes graveyard is usually a symptom of a capture habit that outruns a review habit, and no tool fixes that on its own. Having said that, an agent-maintained vault is a genuinely different proposition from manual filing, because the agent can act on the backlog rather than waiting for you, so the opening is not wrong so much as overstated. The second section is on why this particular notes application, and the third is a thirty-minute setup. The fourth brings in the agent, the fifth is vault architecture, the sixth is having AI maintain the knowledge base, the seventh is seven real workflows with concrete steps and prompts, the eighth covers the plugin ecosystem, and the ninth is an advanced section. The readme's overview expands on the last three: seven workflows you can copy directly, four plugins out of more than a thousand that you actually need, and advanced material on version control, custom skills, local models and multi-vault strategy.
The observations that carry rhetorical weight, stated plainly
Three claims appear in the overview that are best read as assertions to check rather than as findings to accept. The first is that three independent large projects converged on Markdown storage for agent memory and that this is not a coincidence. The inference is suggestive rather than evidential: independent convergence on a plain-text format is unsurprising when the alternative is a bespoke binary store, and a project choosing Markdown may be optimising for human debugging rather than for agent capability. The second is the acquisition and popularity figures attached to the named projects, and the readme asks you to find that notable. Whether a project is popular and whether an agent-memory format choice is correct are different questions, and the book leans on the first to support the second. The third is the claim that the setup effort is under forty minutes, which is a fair description of installing two pieces of software and connecting them, and is not a description of what happens in the following month as you write the index files and find the vault structure does not match how you actually think. None of this makes the guide bad. It makes it a book, with the rhetorical habits of one, and the chapters that include concrete steps and prompts are where the value will be for a reader who wants to act.
Licensing, distribution, and what a fork of this guide looks like
The licence section is a single paragraph and it is not a standard licence. It says the work is shared free for learning and exchange, that reprinting and quoting are welcome, and asks that the source be credited. That is a request rather than a grant, and the difference matters in one specific way: a request has no defined terms, so there is no clear answer about what a publisher may do with a translation, whether a company may include it in internal onboarding material, or whether a derivative may be sold. The repository records no standard licence and the language field is unrecorded, so a reader who wants to republish should ask rather than assume. The distribution channels are more interesting than the licence. The book is available as two PDFs, and the readme says it has also been synchronised to a hosted reading service where each chapter is a web page and the link can be handed to an agent as context:
https://www.workbuddy.cn/space/d/g5E78ejX962o977geKeOjjThat is a meaningful detail: it means the author has already decided that chapter-per-page is the right unit for agent consumption, and it is the same argument the guide makes about Markdown files, applied to the book itself. A collection of sibling guides on other tools is listed, and the whole set is gathered under a second index link, which is a small example of the index-per-collection pattern the guide recommends.
Editorial conclusion
Read this guide if you already keep notes in Obsidian and want a concrete method for making an agent maintain them, because the argument about treating a model as a compiler rather than a retriever is more useful than the specific tooling and the chapter structure is honest about what it assumes. Do not expect a tool from this repository, since it ships two PDF files, a readme in two languages, an assets folder and screenshots, and the primary language is not even recorded, so anything you implement you will build yourself. Three things to weigh. That the central claim is validated at a stated scale, which the readme puts at 400,000 characters, so if your vault is larger the argument has not been demonstrated and if it is much smaller the setup cost may exceed the benefit. That the convenience claim about not needing a complex tool protocol setup rests on one instruction file plus one index file per folder, which is a small amount of structure and also the whole of the navigation model. And that the licensing is informal, since the readme says the work is free for learning and exchange, welcomes reprinting with attribution, and no standard licence is declared, which is a weaker grant than a named licence and matters if you plan to translate or republish it. The last commit was on 2026-08-23 and there are no releases.
Frequently asked questions
Is obsidian-ai-orange-book software I can install?
No. The repository contains two PDF files, a readme in Chinese and English, an assets directory and a screenshots directory, with no source code, build configuration or plugin. Anything you would run has to be written by you, following the method the guide describes.
What is the main argument of the guide?
That a large language model should act as a compiler over your notes rather than as a retriever over an embedding index, so the agent maintains the knowledge base directly instead of querying a retrieval system. The readme says this approach was validated at a scale of 400,000 characters.
What setup does the guide say is necessary?
The readme says complex tool protocol setup is not required, and that one instruction file at the vault root plus one index file per folder is enough so the agent can move through the whole knowledge base efficiently. It also claims the setup can be done in about forty minutes.
How is the book licensed?
Informally. The readme says the work is shared free for learning and exchange, that reprinting and quoting are welcome, and that attribution is requested. No standard licence is declared and the repository records no licence field, so there are no defined terms for translation or republication.
Where can I read the guide online?
The readme points to a hosted reading service where each chapter is a separate web page and says you can hand the link to an agent as context. It also links an index gathering the author's other guides in the same series.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/alchaincyf-obsidian-ai-orange-book)
Community notes