# taiwan-md: an AI-native knowledge base with 1122 articles and 13 languages

> A statically built site where Traditional Chinese is the single source of truth, translations are machine-assisted, and any MCP client can query it over stdio for free.

**frank890417/taiwan-md** — 🇹🇼 讓全世界完整認識台灣 | An open-source, AI-friendly knowledge base about Taiwan

- Repository: https://github.com/frank890417/taiwan-md
- Website: https://taiwan.md
- Stars: 1,199 · Forks: 187
- Language: HTML
- License: not declared
- Published: 2026-10-06 · Updated: 2026-10-06 · Language: en
- Canonical page: https://hysenlabs.com/projects/frank890417-taiwan-md

## Traditional Chinese as the single source of truth

The pitch in the repository description is that this is the world's first AI-native open knowledge base about Taiwan, and the first thing the README does with that claim is pick a side. Content is written in Traditional Chinese by default, which the README justifies rather than assumes: it calls it the world's oldest writing system still in daily use and Taiwan its last major home.

Everything downstream follows from that. The statistics table marks the Chinese column as the SSOT and counts 1122 articles there, with English at 1108, Korean at 1106, Vietnamese at 1107, French at 1105, Spanish at 1104, Portuguese at 1077, Russian at 1068, Japanese at 1035, Arabic at 1054, Indonesian at 1014, Hindi at 1015 and German at 835. The gap between 1122 and 835 is not a backlog so much as a visible cost of the model, and the project is upfront about it by publishing per-language numbers rather than a single headline.

The README also draws a line against what this is not. It is not a Wikipedia clone and not a tourism guide. What it describes instead is a curated literary exhibition, with every page answering why this matters, written in an essay-driven style using Noto Serif TC and explicitly inspired by a Taiwanese publication. That editorial framing explains the three-layer depth structure: a 30-second overview, a 5-minute read, and the full article.

## Making the knowledge base readable by an AI client

The AI-native claim is mostly plumbing, and the plumbing is documented. The README lists `llms.txt` and `robots.txt` as first-class outputs, structured Markdown as the SSOT format, and JSON-LD structured data plus Open Graph cards and RSS feeds for the site itself.

The more interesting piece is the MCP connector. You can point any MCP client at it with one command:

```bash
claude mcp add taiwanmd -- npx -y taiwanmd mcp serve
```

What the README promises about that is specific and worth repeating: free, no API key, no account, running on your machine over stdio, with your questions not sent anywhere. It names Claude Code, Claude Desktop, Cursor and Codex CLI as clients, and points to `cli/CONNECTOR.md` for the details. For clients that cannot run Node, there is a remote endpoint at `mcp.taiwan.md`.

The same npm package also reads locally. `npx taiwanmd` prints a ladder of four rungs with live article and language counts, so you can see the size of the corpus from a terminal before committing to anything. The CLI is described as supporting terminal-native reading, quiz, search and retrieval augmented generation, which is a wider feature list than most content sites would attempt.

## The contributor node that stops at a pull request

The fourth rung is the one that raises the most questions, because it turns a static site into something that writes to itself. A contributor node wakes once a day, picks up one open task such as a missing translation, a broken link or a metadata gap, does it, and opens a pull request.

```bash
claude plugin marketplace add frank890417/taiwan-md
claude plugin install taiwanmd
```

What the README is careful about is the boundary. A node's output always stops at a pull request: it never pushes to the repository, never merges, and never posts anywhere as a maintainer. The project describes that as the design rather than a limitation, with merge staying a human decision, and points to `docs/pipelines/CONTRIBUTOR-NODE-PIPELINE.md` for the contract.

The zero-code contribution story is part of the same ladder. Alongside the plugin, the site offers forms, AI prompts or email as ways to contribute, and the repository has `CONTRIBUTING.md` and a `CODE_OF_CONDUCT.md`. The image pipeline is CC BY-SA 4.0 with Wikimedia Commons images cached locally, which matters for a project that wants to be remixable rather than merely readable.

## Quality control as a build step, not an editorial habit

The feature list includes an article health scanner that describes its job as automated detection of hollow AI content, citation gaps, broken links and structural drift. The README also promises that every article includes references and data attribution, and `SOURCES.md` sits in the repository root alongside `SECURITY.md` and `ROADMAP.md`.

The implementation detail is in `package.json`, where the prebuild chain is long and mostly Python invoked from npm scripts. The build runs sync, status, aliases, then `run-p` over api, map, dashboard, changelog, contributors, supporters, china-terms, fork-graph, spores, buildperf, langswitch, og, i18n, content-dates, gitinfo, content-stats, i18n-progress, llms, and stats, before finishing with search, latest and redirects. Node 22.12 or newer is required and the project is an ES module.

So the scanner is not a suggestion to editors. `prebuild:dashboard` runs an article-health script across everything with a quality baseline and an immune profile, writes a health sweep, and feeds it into the dashboard generation. The site you read is downstream of that check, and a dashboard exists to show the check's results in public. The repository tree explains the rest of the shape: `knowledge/` for content, `i18n/` for translations, `src/` for the Astro site, `supabase/` and `workers/` for backend pieces, `cli/` for the connector, and `bench/` for benchmarks.

## Fourteen categories and a knowledge graph over them

The category table gives 14 in total, and the counts show where the writing effort went. History has 44 articles, Geography 63, Culture 61, Food 55, Art 37 and Music 33. Each row carries a highlights column listing specific subtopics rather than vague descriptors, which is a reasonable proxy for editorial consistency.

On top of the articles sit two D3.js visualisations. One is a force simulation knowledge graph with zoom, drag and cross-category bridges, holding 220 or more nodes. The other is a bidirectional tidy tree mindmap of 146 or more official Taiwan websites, which is the more useful of the two for anyone trying to find a primary source rather than a summary.

There is a resources section and a contribute section alongside them, and the site also exposes a dashboard described as a real-time health monitor with quality scores and growth charts, backed by analytics. The repository is not archived, last pushed 2026-09-20, with 1185 stars, 186 forks and only 4 open issues. The stats table shows 75 contributors and a pace of 36 articles in the past 7 days and 132 in the past 30.

## Reading the release notes as an argument about machine authorship

The releases are unusually long and unusually reflective, and they are the best documentation in the repository for how the project thinks about its own reliability. v1.15.0, tagged 2026-08-11, covers 1733 commits in 17 days, with Traditional Chinese articles moving from 863 to 889 and translations rising from 5675 to 8764. Coverage for six languages among new July content went from 27.3% to 81.9%.

That same release reports a naked rate of 19.6%, meaning 177 of 905 Chinese articles had zero footnotes, and an immune score of 60 with a note about a yellow light for 38 days. A project tracking its own citation gaps in public, with the numbers moving in the release title, is doing something few content projects do.

v1.14.0, tagged 2026-07-26, describes the infrastructure turn: the routine flywheel moving off one contributor's laptop onto a Mac mini that never closes its lid, contributor machines becoming split nodes, and the translation line becoming a compute pool where local GPUs and idle cards are scheduled. It also records the language count going from 6 to 12, with Arabic and Russian opening the site's first right-to-left reading direction on 2026-07-25. v1.13.0, tagged 2026-07-16, is about moving the editorial desk onto the public site, where each long-form article has a making-of page showing which gate caught it. German arrived on 2026-09-05 and is the language still filling out at 835 articles.

## Conclusion

taiwan-md is an interesting case because its unusual choices are all visible in the repository rather than hidden behind a site. Traditional Chinese as the single source of truth, with everything else derived from it, is the decision that shapes the rest: the translation pipeline, the quality scanner, and the contributor node all exist because one language is canonical and twelve others are projections of it. The parts that will decide whether it stays useful are the health metrics the project publishes against itself, including the bare citation rate it tracks as an immune score, and the 2025 to 2026 history of that number under pressure. The README is unusually self-critical, which is a good sign, but a knowledge base that projects 14,586 article versions has an obvious failure mode of confident translation with no sources behind it. Install the CLI with `npx taiwanmd` before anything else, read `SOURCES.md` and `CLAUDE.md`, and judge any single article against its references.

## FAQ

### How does taiwan-md serve content to an AI assistant?

Through an MCP connector. `claude mcp add taiwanmd -- npx -y taiwanmd mcp serve` starts a local stdio server with no API key and no account, and the README states that questions stay on your machine. Claude Code, Claude Desktop, Cursor and Codex CLI are all listed as clients, with a remote endpoint at mcp.taiwan.md for clients that cannot run Node.

### Why is the content written in Traditional Chinese first?

The README treats Traditional Chinese as the single source of truth and justifies it as the world's oldest writing system still in daily use, with Taiwan as its last major home. Every other language is a translation of that source, which is why the per-language counts differ and why German, the newest addition, sits at 835 articles against 1122 in Chinese.

### How can I contribute to taiwan-md without writing code?

The README lists forms, AI prompts and email as zero-code contribution routes, alongside the `npx taiwanmd contribute` command that scaffolds an article for you to edit and open as a pull request. Content is licensed CC BY-SA 4.0, and images come from Wikimedia Commons with local caching, so contributions are remixable rather than locked down.

### What does the contributor node in taiwan-md actually do?

It wakes once a day, picks one open task such as a missing translation, a broken link or a metadata gap, completes it, and opens a pull request. The README is explicit that a node never pushes to the repository, never merges and never posts anywhere as a maintainer, so merge stays with a human. The full contract is in `docs/pipelines/CONTRIBUTOR-NODE-PIPELINE.md`.

## Sources

- [frank890417/taiwan-md on GitHub](https://github.com/frank890417/taiwan-md)
- [Issues](https://github.com/frank890417/taiwan-md/issues)
- [Project website](https://taiwan.md)
- [README](https://github.com/frank890417/taiwan-md/blob/main/README.md)
- [Releases](https://github.com/frank890417/taiwan-md/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/frank890417-taiwan-md
