# kordoc: Converting Korean HWP, HWPX and PDF Documents to Markdown

> kordoc is a TypeScript CLI and MCP server that turns HWP, HWPX, HWPML, PDF, XLS(X), DOCX and image files into Markdown, with a lossless patch path back into HWPX and HWP 5.x. The idea is sound and the format coverage is unusually wide, but the documentation is written almost entirely in Korean and the accuracy numbers come from the author's own corpus.

**chrisryugj/kordoc** — , HWP HWPX PDF Office Markdown . CLI MCP | Convert Korean documents (HWP, HWPX, PDF, Office) to Markdown, CLI and MCP server with form filling and diff.

- Repository: https://github.com/chrisryugj/kordoc
- Website: https://www.npmjs.com/package/kordoc
- Stars: 1,974 · Forks: 364
- Language: TypeScript
- License: MIT
- Published: 2026-08-08 · Updated: 2026-08-18 · Language: en
- Canonical page: https://hysenlabs.com/projects/chrisryugj-kordoc

## The document swamp kordoc was built for

Korean public-sector work runs on HWP and HWPX, the formats produced by Hangul (Hancom Office). A civil servant who spent seven years in that environment wrote kordoc, and the README says so directly. The problem is not that these files are binary. It is that the tooling around them is closed: most converters handle HWPX, which is a ZIP of XML, and give up on HWP 5.x, which is a compound binary with compressed streams and layout caches. Add scanned PDFs, legacy XLS, DOCX from other ministries and the occasional PNG of a stamped form, and any automation project starts with a month of parsing work before it does anything useful.

kordoc targets that whole set. The package description lists HWP3-5, HWPX, HWPML, PDF, XLS(X) and DOCX, plus PNG, JPG and WebP images through built-in OCR. The audience is narrow and specific: developers building retrieval or agent pipelines over Korean administrative documents, and the agents themselves, since the project ships an MCP server alongside the CLI. If you never touch a .hwp file, this is not your tool.

## How parsing, patching and rendering fit together

There are three distinct paths in the codebase, and conflating them is the easiest way to misjudge the project.

The first is extraction. A file goes in, Markdown comes out, and the API can also return a structured result with markdown, pages, blocks and metadata. Page numbering is the interesting part. When the source was saved by Hancom, kordoc reads the saved layout cache (linesegarray in HWPX, PARA_LINE_SEG in HWP5) and reports real page boundaries, setting pageMode to layout. When the file was generated by kordoc or by an AI, there is no cache, so page boundaries fall back to section approximations and pageMode reads section. The release notes for v4.7.3 state that using the pages filter in approximate mode emits a PAGE_BOUNDARY_APPROXIMATE warning. That distinction matters more than most of the feature list, because it tells you whether pagination in your output can be trusted.

The second path is the round trip. Markdown produced by kordoc can be edited and handed back to patchHwpx or patchHwp, which replace only the changed paragraph and table cell text inside the original file rather than regenerating it. The README claims original formatting is left untouched at the byte level. This is the strongest idea in the project: most converters are one-way, and one-way conversion is useless when the output has to go back into a government approval chain.

The third path is rendering, which turns HWPX into SVG using the same layout cache, or a pure TypeScript reflow engine when no cache exists. The README notes that reflow was validated against measured line breaking and that Hancom-saved files still replay their cache. Rendering is where the project is most exposed to edge cases, and the release notes read like a list of them.

## Installing kordoc and converting your first HWPX file

The documented install path is a single npx command. It assumes Node.js 18 or newer and runs on macOS, Linux and Windows. The setup wizard is interactive: it lists supported AI clients, marks the ones it detects on your machine, patches the relevant config file, and asks you to restart the client. There is no manual JSON editing in the documented flow.

```bash
npx -y kordoc setup
```

After a restart, the README says 15 document tools become available to the client, including parse_document, parse_table, fill_form, patch_document and generate_document. If you only want the command line, the README says no installation is needed at all:

```bash
npx kordoc <파일>
```

For the Claude Code plugin route, the README gives two slash commands instead of MCP registration:

```bash
/plugin marketplace add chrisryugj/kordoc
/plugin install kordoc@kordoc
```

The README states that the skill activates automatically when a .hwp or .hwpx file is mentioned or a government document generation or form-filling request is made, and that it calls the CLI internally with npx -y kordoc@^4, so no separate install is needed.

Two failure modes are documented up front. A MODULE_NOT_FOUND error pointing at dist/cli.js means a broken global install is still on disk; the README's fix is npm uninstall -g kordoc followed by npx -y kordoc@latest setup. On Windows PowerShell, a PSSecurityException about npx.ps1 is the default execution policy blocking unsigned scripts, which the README says is unrelated to kordoc. Running the command from cmd avoids it, or you can set the policy once with Set-ExecutionPolicy -Scope CurrentUser RemoteSigned. Note the two spellings: setup for the wizard, plain npx kordoc for conversion.

## Where kordoc breaks or misleads

The seal placement feature is the clearest example of an honest limitation. kordoc seal finds anchor text such as "(인)" and floats a stamp image over it. The README states that positions inside nested tables, text boxes and paragraphs with tabs or line breaks are approximate, that the result reports this through warnings, and that you should verify in Hangul and nudge with --dx or --dy. A tool that tells you when it is guessing is more trustworthy than one that does not, but it also means seal placement is not a hands-off operation.

Table extraction has a structural constraint that is easy to miss. The README explains that a table spanning multiple pages is, in the current intermediate representation, a single block belonging to the starting page, so the markdown for the intervening pages is an empty string. The array length still equals the page count, so a naive consumer that concatenates page markdown will silently lose nothing, but one that assumes every page has content will misbehave.

Then there is the accuracy question. The README cites strong numbers: 1,673 of 1,673 tables and 27,714 of 27,714 cells lossless on a 291-document HWPX corpus, and 98.6 percent matching with 65.2 percent exact on six HWPX-to-PDF pairs. Those are the author's own measurements against his own corpus, produced by benchmark scripts in bench/ that run as a gate in prepublishOnly. They are a reasonable signal of care. They are not independent verification, and a corpus of government documents will not represent your documents. The 65.2 percent exact match figure in particular is worth reading twice: more than a third of cross-format table comparisons did not match exactly.

Finally, the documentation is Korean-first. README-EN.md exists, but the main README, 사용법.md and the changelog are in Korean, and the changelog is where most of the behavioural detail lives. An English-only team will be reading release notes through a translator for every upgrade.

## kordoc against a general-purpose converter

The obvious comparison is a general document converter such as Pandoc or a Python library like python-docx paired with a PDF text extractor. The difference is not quality, it is the target format. Pandoc has no HWP or HWPX reader, because those formats have essentially no audience outside Korea. It handles DOCX, XLSX and PDF through other tools well, and it has a stable, documented, widely understood interface.

kordoc's approach is the opposite: deep support for one country's document formats, including the binary HWP 5.x that general tools skip, at the cost of a smaller and faster-moving API. The round-trip patch capability has no equivalent in the general tooling, because general converters treat conversion as one-way by design. If your documents are DOCX and your output is a static site, Pandoc is the safer choice. If your documents are HWPX approval forms that must come back out looking identical, kordoc is addressing a problem the general tools do not attempt. The two are not really substitutes; the decision is made by the input format, not by preference.

## Licence, upgrade cost and what the repository does not say

kordoc is MIT licensed, and the repository carries a NOTICE file and a THIRD_PARTY directory, with a check-notices.mjs script wired into prepublishOnly. That suggests third-party components are tracked deliberately, which matters for anyone shipping this inside a commercial product. The built-in OCR uses PP-OCRv5 korean running on local CPU, so there is no API key and no external service in that path, which removes a category of data-handling review. I am not a lawyer and this is not legal advice: read NOTICE and THIRD_PARTY yourself before redistributing.

The upgrade cost is real. Version numbers move fast. The published package.json in the repository is at 4.13.2, while the most recent release listed is v4.10.0 from 2026-08-27, so the repository is ahead of the tagged release. Patch releases have historically bundled large sets of fixes: the 3.16.1 patch is described as correcting 55 defects found in an integration review, mostly in the category of quietly wrong output behind a success message. That is a healthy thing to fix and an unhealthy thing to have shipped. Pin your version, and read the Korean changelog before upgrading, because behaviour changes are documented there and not in the README.

The last push to the repository was on 2026-08-27. The README does not document a rollback procedure for patchHwp or patchHwpx, and it does not describe a compatibility matrix between kordoc versions and the Hancom versions that produced your source files.

## Conclusion

Adopt kordoc if your pipeline is already Korean-document shaped: government forms, HWPX reports, scanned PDFs with tables, or an agent that needs to read .hwp files. Skip it if you need a stable public API surface, English-first documentation, or a converter whose accuracy claims you can reproduce without the author's corpus. Before committing, run npx -y kordoc setup against one real document from your own archive and check the pageMode value in the JSON output: layout means the page boundaries came from Hancom's saved layout cache, section means they are an approximation.

## FAQ

### What is kordoc and which document formats does it support?

kordoc is a TypeScript CLI and MCP server that converts Korean documents to Markdown. The package description lists HWP3-5, HWPX, HWPML, PDF, XLS, XLSX and DOCX, plus PNG, JPG and WebP images handled through built-in OCR.

### How do I install kordoc?

The README gives npx -y kordoc setup, which runs an interactive wizard that patches your AI client's config file and asks you to restart it. For command-line use only, the README says no installation is needed and npx kordoc <file> works directly. Node.js 18 or newer is required.

### Does kordoc convert Markdown back to HWPX without losing formatting?

The README describes patchHwpx and patchHwp as replacing only the changed paragraph and table cell text inside the original file, leaving the original formatting untouched. Row additions and deletions in tables are supported from v3.7, and filling empty cells in HWP 5.x from v3.8.

### Can kordoc read scanned Korean PDFs without an API key?

Yes, according to the README. Since v4.2 the OCR runs as local CPU inference using PP-OCRv5 korean, with no external service and no API key, and it selects only pages whose text layer is broken before running OCR.

## Sources

- [Official documentation](https://www.npmjs.com/package/kordoc)
- [Official README](https://github.com/chrisryugj/kordoc#readme)
- [Project repository](https://github.com/chrisryugj/kordoc)
- [Release notes](https://github.com/chrisryugj/kordoc/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/chrisryugj-kordoc
