Self-hosted service
axoviq-ai/synthadoc avatar
axoviq-ai/synthadoc

Synthadoc: an LLM wiki engine that compiles documents into Markdown

Synthadoc: An open-source LLM knowledge compilation engine that turns raw documents into structured, local-first wikis. A transparent, human-readable alternative to traditional RAG, which can be self-managed and self-improved without the use of any tools.

1,372 stars132 forksPythonAGPL-3.0

At a glance

What is it?
Synthadoc turns PDFs, spreadsheets, slide decks and transcripts into a local Markdown wiki with cross-references and source citations. It is a compile-time alternative to retrieval at query time, and it is licensed AGPL-3.0.
Who is it for?
Adopt Synthadoc if you want a knowledge base you can read, diff and back up as plain Markdown, and if you are willing to run an LLM on every ingest. Skip it if your corpus changes hourly and you need answers seconds later, because compilation is a batch job and the README does not document rollback.
Can I use it commercially?
Yes, with strict conditions. AGPL-3.0 is a network copyleft licence: if people use a modified version over a network, for example as a hosted service, you must offer them its source code under the same licence.
Is it still maintained?
Yes. The repository last received commits 5 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 25, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What Synthadoc is for, and who it is not for

Synthadoc is a knowledge compilation engine. The README describes it as domain-agnostic and says it reads raw source documents (PDFs, spreadsheets, PPTs, web pages, images, videos, Word files, TXTs and AI session transcripts in .jsonl) and uses an LLM to synthesize them into a persistent, structured wiki. The output is not an index or an embedding store. It is Markdown files on disk.

The framing comes from an Andrej Karpathy gist the README quotes: "The LLM should be able to maintain a wiki for you." Synthadoc takes that literally. Most tools in this space retrieve and summarize at query time. Synthadoc compiles at ingest time, so every new source is supposed to enrich and cross-link the whole corpus rather than append another chunk. The wiki is the artifact, and it stays readable and browsable with no process running.

The README names three audiences. A solo researcher or freelancer can run it on Gemini Flash or a local Ollama model, which the README presents as zero ongoing cost. A team of three to twenty people gets a shared internal wiki where the system resolves contradictions automatically. Enterprises get per-department wikis on separate ports, an audit trail for every ingest and cost event, a hook system for CI/CD integration and OpenTelemetry for ops dashboards.

That last row is where the design gets interesting and where the risk sits. A wiki that rewrites itself when new documents arrive is only as trustworthy as the model doing the rewriting. The README claims contradictions are detected and surfaced, and that every answer cites its sources. Those two features are the whole argument for the approach, and they are also the two things a prospective user should test hardest before trusting the output.

How compilation works: ingest, synthesis, and the Markdown artifact

The mechanism visible in the repository is a pipeline, not a query engine. Documents go in, an LLM writes and rewrites wiki pages, and the result is a directory of Markdown. The README says cross-references are built automatically, contradictions are detected and surfaced, orphan pages are flagged, and every answer cites its sources.

The dependency list in pyproject.toml tells you what the ingest layer actually parses. There is pypdf and pdfminer.six for PDFs, python-docx for Word, python-pptx for PowerPoint, openpyxl for spreadsheets, beautifulsoup4 for web pages, and youtube-transcript-api for video transcripts. So the format coverage in the README maps to concrete libraries rather than a generic extractor. Retrieval, where it happens inside the engine, uses rank_bm25, a lexical scorer, not a vector database.

Two dependencies stand out for anyone reasoning about scale. networkx and python-louvain are both present, which points to graph construction and community detection over the wiki pages. That is consistent with automatic cross-referencing and orphan detection: you need a graph to know a page has no inbound links. The persistence layer is aiosqlite, so state lives in SQLite rather than a server database.

Model access goes through anthropic and openai clients, with httpx underneath. The README mentions Gemini Flash and Ollama as options, so the provider layer is not limited to the two SDKs. The engine also ships a FastAPI and uvicorn web UI, a Typer CLI, and an MCP server (the mcp dependency, pinned below version 2.0), which is how the README says you connect Claude. File locking comes from filelock, which matters if two ingest runs touch the same wiki.

The important architectural consequence: because the output is Markdown, the wiki has no runtime. You can open it in Obsidian, put it in git, or sync it to a drive. The README states there is no cloud account and no vendor lock-in. That is a real property of the design, not marketing: the artifact outlives the tool.

Installing Synthadoc and compiling a first wiki

The package is on PyPI as synthadoc and requires Python 3.11 or newer, per pyproject.toml. The README's installation section is the place to check for the current command, but the standard path for a published Python package applies:

bash
pip install synthadoc

After installation the CLI is available. The README documents a command reference organized by use case and a user quick-start guide under docs/user-quick-start-guide.md, which is where the concrete ingest and build commands live. The repository also ships a wiki/ directory and an obsidian-plugin/ directory, so there is a working example layout to compare against.

The configuration surface is YAML, since pyyaml is a dependency and the README has a Configuration section. Provider credentials come from the anthropic or openai SDKs, so the usual environment variables for those clients apply. The README does not spell out every key in the excerpt available here, so read docs/ before assuming a default.

For a first run, point the tool at a small folder rather than your whole archive. The README's end-to-end example is an M&A due diligence walkthrough under docs/example/aquaflow/, which is a realistic shape: a bounded set of documents, one domain, one wiki. Start there.

The README also describes four interfaces: CLI, Obsidian, Web UI and MCP. The Web UI runs on FastAPI and uvicorn; the README mentions per-department wikis on separate ports, so port selection is a configuration concern rather than a fixed value. The MCP path is documented in the quick-start guide under the heading about connecting Claude via MCP. Pick one interface for the first run and ignore the others until the wiki compiles cleanly.

Where Synthadoc breaks down

The failure mode that matters most is model dependence. Every cross-reference, contradiction flag and citation in the wiki is produced by an LLM. If the model is weak, the wiki is confidently wrong in a format that looks authoritative because it is Markdown with links. The README's claim that contradictions are detected and surfaced describes intent, not a guarantee, and it does not say what happens when the model misses one.

The second constraint is latency. Compilation at ingest time means the cost is paid up front, per document, with an LLM call. A corpus that changes hourly is a poor fit: you are either recompiling constantly or reading a stale wiki. Query-time retrieval systems handle that case better precisely because they do not maintain a persistent artifact.

Third, the README does not document rollback. There is an audit trail for every ingest and cost event, and filelock is in the dependency list, but nothing in the available documentation describes reverting a bad compile. Since the output is plain Markdown, git is the obvious mitigation, and the README does suggest backing the wiki up with git. That is a workaround, not a feature.

Fourth, the project is classified as Development Status 4 - Beta in pyproject.toml. The version history supports that: v1.3.1 landed on 2026-08-26, with v1.3.0 and v1.2.1 in the two weeks before it. Three releases in a month is a fast cadence, which is good for fixes and bad for anyone pinning behaviour.

Finally, scale is unproven here. SQLite and BM25 are sensible for a personal or departmental wiki. The README does not state a corpus size at which they stop being sensible, and nothing in the documentation describes a benchmark. Treat the enterprise row of the audience table as a deployment pattern (separate ports, hooks, OpenTelemetry) rather than a demonstrated capacity claim.

Synthadoc against query-time retrieval

The clearest alternative is a conventional retrieval-augmented generation stack: chunk documents, embed them into a vector store, and retrieve relevant chunks at query time for an LLM to summarize. Tools in that family keep an index, not a document. The index is opaque, it cannot be read in an editor, and the answer exists only for the moment of the query.

The difference in approach is when the LLM does its work. RAG pays at query time and amortizes nothing: ask the same question twice and you pay twice, and two users asking related questions get two unlinked answers. Synthadoc pays at ingest time and keeps the result. Ask a question and you read a page that already exists, with links to related pages and citations to sources.

That trade has a cost the RAG side does not carry. A RAG index can be rebuilt cheaply from the same chunks. A compiled wiki is a set of authored documents, and recompiling can change them. If you need byte-stable answers, a retrieval system is easier to reason about.

There is also a middle option worth naming: using an LLM to summarize individual documents into notes, then linking them by hand or with a tool like Obsidian. Synthadoc automates the linking and the contradiction checking, which is the part humans skip. Whether the automation is better than your own judgement on a 200-page corpus is an empirical question the README cannot answer for you.

The honest positioning is this: Synthadoc is for corpora that are read more often than they change, and for people who want the knowledge artifact to be inspectable. If your documents are a firehose, keep the retrieval stack.

Licence, maintenance and upgrade cost

Synthadoc is licensed AGPL-3.0-or-later, stated in pyproject.toml and in the LICENSE file at the repository root. AGPL is the network-copyleft variant: the practical implication is that if you modify Synthadoc and let users interact with it over a network, the licence's source-disclosure obligations are generally understood to extend to that interaction. This is a description of the licence's intent, not legal advice. If you plan to embed Synthadoc in a hosted product, have counsel read the licence text rather than this paragraph.

For internal use the licence is usually a non-issue, and the README's enterprise framing (local wikis, per-department ports, audit trail) is compatible with keeping everything in-house. The friction appears only if you redistribute or expose a modified engine.

The maintenance picture is active. The last push to the default branch was on 2026-08-26, the same day v1.3.1 was released, and the repository is not archived. The recent cadence is roughly weekly: v1.2.1 on 2026-08-13, v1.3.0 on 2026-08-19, v1.3.1 on 2026-08-26. There is a CI workflow at .github/workflows/ci.yml and a tests/ directory, so changes are gated by something.

Upgrade cost is where the beta label bites. A tool that rewrites your wiki on every ingest can also change how it rewrites it between versions. The dependency set is broad (roughly thirty packages, including two LLM SDKs, two PDF parsers and an MCP server pinned below 2.0), so a major bump in any of them can ripple. Pin the version, keep the wiki under git, and recompile a sample corpus before upgrading the whole archive. The README's audit trail and the filelock dependency suggest the project takes concurrent and repeated runs seriously, but neither substitutes for a diffable backup.

Editorial conclusion

Adopt Synthadoc if you want a knowledge base you can read, diff and back up as plain Markdown, and if you are willing to run an LLM on every ingest. Skip it if your corpus changes hourly and you need answers seconds later, because compilation is a batch job and the README does not document rollback. Before committing, verify three things on a small corpus: that your chosen model resolves contradictions acceptably, that the per-department port layout in the docs fits your network, and that AGPL-3.0 is compatible with how you intend to distribute anything you build on top of it.

Frequently asked questions

What is Synthadoc called when an AI writes and maintains its own wiki?

The README describes Synthadoc as a knowledge compilation engine: it uses an LLM to synthesize raw documents into a persistent, structured wiki at ingest time rather than retrieving at query time. It cites an Andrej Karpathy gist about an LLM maintaining a wiki as its inspiration.

Does Synthadoc require a cloud account or a vector database?

No. The README states there is no cloud account and no vendor lock-in, and the output is plain Markdown. Retrieval inside the engine uses rank_bm25, a lexical scorer, and state is stored with aiosqlite, so there is no vector store in the dependency list.

Which file formats can Synthadoc ingest?

The README lists PDFs, spreadsheets, PPTs, web pages, images, videos, Word files, TXTs and AI session transcripts in .jsonl. The dependencies match that list with pypdf, pdfminer.six, openpyxl, python-pptx, python-docx, beautifulsoup4 and youtube-transcript-api.

Which Python version and licence does Synthadoc use?

pyproject.toml requires Python 3.11 or newer and declares the project under AGPL-3.0-or-later. The classifiers mark it as Development Status 4 - Beta.

How does Claude connect to Synthadoc?

Through MCP. The repository depends on the mcp package (pinned to versions below 2.0), and the README's quick-start guide has an appendix titled about connecting Claude via MCP. The README also lists MCP as one of four interfaces alongside the CLI, Obsidian and the Web UI.

Official sources

  1. Official README
  2. Project repository
  3. Release notes
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/axoviq-ai-synthadoc.svg)](https://hysenlabs.com/projects/axoviq-ai-synthadoc)