# Molio turns your documents into a Markdown knowledge base your agents can write back to

> A local-first TypeScript workspace where PDFs, Office files, images, web clips and whole books are converted to Markdown, indexed by a wiki engine into entities and cross-links, and then used as the working context for Claude Code, Codex, Gemini CLI or Qwen Code, with every task result deposited back into the same vault.

**zhuzhaoyun/Molio** — A local-first personal knowledge layer for AI agents. Build evolving knowledge spaces with LLM Wiki, knowledge graphs, and agent workflows.

- Repository: https://github.com/zhuzhaoyun/Molio
- Website: https://molio.cn
- Stars: 406 · Forks: 55
- Language: TypeScript
- License: NOASSERTION
- Published: 2026-09-18 · Updated: 2026-09-18 · Language: en
- Canonical page: https://hysenlabs.com/projects/zhuzhaoyun-molio

## Everything lands on disk as Markdown, which is the whole migration story

The supported-input table is the clearest statement of what this tool is for. Markdown, TXT, HTML and CSV are read as-is, because that is the default shape of a knowledge base here. PDF goes through layout analysis, OCR and table reconstruction before arriving as Markdown. Word, PPT and Excel take the same path with heading hierarchy and tables preserved. Images in PNG, JPG and TIFF are read by OCR. Web pages arrive by one-click clip, or by article extraction for WeChat posts. An Obsidian vault is opened in place, with its original files treated as read-only. Books of a million words or more go through a chunked preprocessing and digest pipeline that is resumable.

The commit at the end of all of that is that nothing stays in a proprietary shape. Files on disk are plain Markdown, which is the basis for the claim that there is no migration and no lock-in.

One cost is stated openly. PDF and Office conversion runs through a bundled docling skill, and the first conversion downloads that tool plus roughly 500MB of models. Later runs reuse the cache, so the expense is paid once per machine rather than per file.

## The Wiki engine is what makes accumulated Markdown readable by an agent

Markdown on disk is not yet something an agent can navigate. That gap is what the Wiki engine exists to close. It extracts entities and concepts from the unified Markdown layer, then weaves dense cross-links and layered indexes across them.

The phrasing used for this is worth taking seriously, because it sets an order of operations: data becomes a foundation only after it is processed. Nothing in the repository claims you can skip the indexing step and point an agent at a folder of PDFs.

The screenshot captions name two surfaces that follow from this design. The Knowledge Space view is a vault file tree with Markdown rendering, so the processed layer is visible as files rather than hidden in an index. The Agent Workbench view carries multi-agent support with streaming responses, which is the other half: the agent is reading a space that was built for reading.

What the engine does not claim is anything about accuracy of the extracted entities or concepts. No evaluation, no precision or recall figure and no failure cases appear alongside it, so treat the graph as navigation infrastructure rather than as a verified knowledge base.

## Four named agents work the space, and their output is deposited as Markdown

The second of the three stages is agent execution, and four tools are named for it: Claude Code, Codex, Gemini CLI and Qwen Code. They run inside your knowledge space doing research, writing, answering and analysis, and what they see is the accumulated material rather than a blank slate. A unified GUI lets you pick an agent and watch streaming output.

The third stage is what makes this more than a viewer. Every task output is written back as Markdown, described as a reusable long-term asset, and the knowledge graph keeps growing so the next task starts from a higher base. Publishing is the last step of that loop: typesetting through doocs/md distributes to more than thirty platforms in one click.

So the unit of value is a Markdown file that survives the session. That design has a direct consequence. Anything an agent writes into the vault becomes part of the context for the next agent, including its mistakes, which is why the deposit step is the part worth being deliberate about.

There is also a phone path: a QR code lets you chat with your base from your phone through WeChat.

## Only the web console carries a complete channel today

A single instance can serve several channels in parallel, and most can be onboarded straight from the Web console. The support table is uneven, and the unevenness is stated rather than smoothed over.

The Web Console is the default and the only channel with text, image, file and voice all supported; group chat is marked not applicable for it. WeChat supports text, image and file, with voice and group chat planned. Feishu supports text, image and file plus group chat, with voice planned. Telegram has nothing supported yet and everything planned. Slack and Discord have text, image and voice marked not applicable, with file also planned.

Read that as a roadmap column rather than a compatibility list. If your channel is WeChat or Feishu you can start today on text, images and files. If it is Telegram, Slack or Discord, the table describes intent, not function.

The QR code entry point belongs to this part of the system too. It is the mobile front door to the same base, which is why voice being unsupported on WeChat is the gap that would matter most for phone use.

## The resource library is the shortcut, and its catalog is weighted toward Chinese classics

For anyone starting from zero, the answer offered is not the tool but the catalog. A resource library at molio.cn/resources.html publishes ready-made structured knowledge graphs that you import in one click instead of building.

The listed highlights make the shape of that catalog clear. Dream of the Red Chamber covers characters and imagery across all 120 chapters. Jin Ping Mei covers people, commerce and social structure. Shiji covers people, institutions and thought. Zizhi Tongjian covers 1,362 years of rise and fall. History of Ming covers empire, institutions and court politics. Zhouyi covers all 64 hexagrams with line-change and image-number analysis. Laozi and Zhuangzi, Chinese and Western philosophy, traditional Chinese medicine classics and an obstetric ultrasound knowledge base follow, then knowledge engineering covering ontology, RAG, GraphRAG and agent engineering, and a Guangdong gaokao application planning base.

The stated domain coverage is literature, history, philosophy, traditional Chinese medicine, medicine and AI engineering, with new resources added regularly.

The demo video points at the Zizhi Tongjian base as the worked example. The argument made for these assets is that their structure is something an AI cannot produce on demand, which is also the reason the free path into this project may be the catalog rather than the code.

## Docker on port 3100 is the documented deployment, and it targets NAS hardware

The compose file describes a single container holding the daemon, the web interface and the Claude Code CLI, built for linux/amd64 and linux/arm64. The platforms named in its comments are servers and NAS boxes, listing Synology, QNAP, Terramaster, TrueNAS and Unraid by name.

Setup is two commands after copying the environment file:

```bash
cp .env.example .env
docker compose up -d
```

Then open http://<IP>:3100 and configure the AI model under Settings and Runtime. Port 3100 comes from `MOLIO_PORT`, which defaults to that value.

Three volumes are declared. Application state, a SQLite database plus configuration, persists in `molio-data` mounted at `/home/molio/.molio`. Claude Code authentication and configuration persist in `molio-claude` at `/home/molio/.claude`. Your actual documents come from the host path in `MOLIO_VAULT_PATH`, mounted at `/vaults`, defaulting to `./vaults`. On first start with no base present, a default knowledge base is created pointing at `/vaults`, so the web interface opens straight into it.

Two things to check before trusting that compose file. The service pulls a prebuilt `latest` image from a registry hosted on Alibaba Cloud in Guangzhou rather than building locally, and `latest` is not pinned to a digest. A local build is available by commenting out the image line and using `docker compose up -d --build`.

## Node 24 and pnpm 11.5.0, with a package.json version behind the release tags

The root package.json names pnpm@11.5.0 as its package manager and requires Node 24 or newer through its engines field. TypeScript is pinned at ^5.8.3 and concurrently at ^10.0.1 in development dependencies, and the repository carries a `.nvmrc` alongside them.

Layout is a pnpm workspace: `apps/` holds the daemon, web, desktop and cloud targets, `packages/` holds shared contracts, and the scripts run through filters such as `@molio/daemon` and `@molio/web`. `pnpm -r build` builds everything, `pnpm -r typecheck` typechecks, and the test script chains the cloud, daemon, desktop and web suites before running the landing-page tests with `node --test`.

One discrepancy is worth flagging. The version field in package.json reads 0.1.0 while the published tags run 0.3.57, 0.3.58 and 0.3.59, the last on 2026-09-30. The root manifest is private, so that field is not the release number, but it does mean you cannot infer which release you cloned from the repository root.

The container build shows the practical consequences. It starts from `node:24-slim`, installs `python3`, `make` and `g++` because better-sqlite3 is a native module, enables pnpm through corepack, and filters out the desktop package during install since its Electron and Playwright dependencies are large and unnecessary for a server.

## Model access is configured in the web UI, with three provider routes documented

The environment example says outright that the AI model and API key are best set in the web interface under Settings and Runtime, and that the file exists for advanced users and automated deployment.

It documents three routes anyway. An Anthropic API key. An Alibaba Cloud Bailian Token Plan, which sets `ANTHROPIC_BASE_URL` to a token-plan endpoint in Beijing. And DeepSeek through an Anthropic-compatible endpoint, where the same token and key variables are set and the base URL points at the DeepSeek Anthropic path.

The DeepSeek example goes further than a base URL, mapping the model variables explicitly: a default model plus Sonnet, Opus and Haiku overrides, including two with a 1M context suffix. A separate block shows the same four variables mapped to a different model family instead, which is the mechanism for pointing the Anthropic-shaped variables at something else.

File permissions are handled by detection rather than configuration. The container runs as a non-root user because Claude Code requires it, and at startup it reads the owner of the vault directory and aligns the container user to match, printing the uid it settled on. Manual `PUID` and `PGID` exist for the cases where that guess is wrong, namely several mounted directories with different owners or a deliberate fixed identity.

## Conclusion

Adopt it if your material is documents you own rather than rows you query, and if you want an agent to accumulate what it learned rather than lose it at the end of a chat. Think twice before deploying the Docker image on a NAS, because the compose file pulls a prebuilt image from a Chinese registry and ships no pin for its digest. The first thing to verify is the conversion cost you will pay once: the bundled docling path downloads about 500MB of models on the first PDF or Office file, and the vault you point at is mounted read-write, so decide where that lives before you import anything.

## FAQ

### What file formats can Molio import into a knowledge space?

Markdown, TXT, HTML and CSV are read as-is. PDF, Word, PPT and Excel go through layout analysis, OCR and table reconstruction into Markdown, images in PNG, JPG and TIFF are read by OCR, web pages can be clipped or extracted, and an Obsidian vault is opened in place with its originals left read-only.

### Does Molio need a server, and what are the deployment requirements?

Everything runs on your own machine rather than through a third-party server. A Docker Compose file is provided for linux/amd64 and linux/arm64, targets NAS platforms, serves on port 3100 by default via MOLIO_PORT, and persists state, Claude Code credentials and your documents in three volumes.

### Which AI agents can work inside a Molio knowledge space?

Claude Code, Codex, Gemini CLI and Qwen Code are the four named agents, selectable in a unified GUI with streaming output. Their task results are written back into the space as Markdown and publishing goes through doocs/md to more than thirty platforms.

### How much does Molio cost to set up the first time a PDF is converted?

PDF and Office conversion runs through a bundled docling skill, and the first conversion downloads that tool plus roughly 500MB of models. Later runs reuse the cache, so that download happens once per machine rather than once per file.

### Which messaging channels does Molio support today?

The Web Console is the default and supports text, image, file and voice. WeChat and Feishu support text, image and file, with Feishu also covering group chat. Telegram, Slack and Discord are listed as planned rather than supported.

## Sources

- [Issues](https://github.com/zhuzhaoyun/Molio/issues)
- [Project website](https://molio.cn)
- [README](https://github.com/zhuzhaoyun/Molio/blob/main/README.md)
- [Releases](https://github.com/zhuzhaoyun/Molio/releases)
- [zhuzhaoyun/Molio on GitHub](https://github.com/zhuzhaoyun/Molio)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/zhuzhaoyun-molio
