# html-to-markdown: a Rust core HTML converter with bindings for 16 languages

> xberg-io/html-to-markdown turns messy HTML into CommonMark or Djot through one convert() call, with the same output across Rust, Python, Node.js, Go and a dozen more runtimes. Here is how the tiered parser works, how to install it, and where it stops being the right tool.

**xberg-io/html-to-markdown** — High performance and CommonMark compliant HTML to Markdown converter. Maintained by the Kreuzberg team. Kreuzberg is a fast, polyglot document intelligence engine with a Rust core. It extracts structured data from 98+ document formats using streaming parsers and built-in OCR.

- Repository: https://github.com/xberg-io/html-to-markdown
- Website: https://docs.html-to-markdown.xberg.io
- Stars: 877 · Forks: 72
- Language: HTML
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/xberg-io-html-to-markdown

## The HTML you actually have is not the HTML parsers expect

Most HTML to Markdown failures are not conversion failures. They are parsing failures that surface as conversion failures. A page with an unclosed <td>, a stray CDATA block, a custom element the parser has never seen, or a Windows-1252 byte sequence in the middle of a UTF-8 document will either throw or silently drop content. The README frames the project around exactly this: "Feed html-to-markdown the HTML you actually have." The list of inputs it names is unclosed tags, CDATA, custom elements, broken entities, nested tables and mixed encodings.

The intended audience is not someone converting a blog post once. It is someone running a corpus job: a RAG pipeline that has to ingest scraped pages, a migration that has to move a CMS archive into Markdown, a data extraction job where the source pages were generated by five different template engines over ten years. The topics on the repository (rag, text-extraction, hocr, text-processing) point the same direction. If your input is well-formed and your output is one file, almost any converter works and this one is more machinery than you need.

## Tiered dispatch: byte scanner, DOM walker, then html5ever repair

The mechanism the README describes is a three-tier dispatch. A byte scanner handles the common cases first, a DOM walker handles structure, and html5ever repair is the fallback for input the earlier tiers cannot resolve. The claim attached to this design is byte-equal output across tiers, which matters more than it sounds: it means the fast path and the repair path are supposed to agree, so a corpus does not develop two dialects of Markdown depending on how broken each page happened to be.

That is a real design commitment and it is the part I would test first. Byte-equal output across tiers is a strong invariant to hold across a 16-language surface, and the README does not document how it is enforced beyond the statement itself. The benchmark claim is 19 to 116 MB/s on the Wikipedia/mdream corpus, with per-group regression thresholds enforced on every PR. Those numbers come from the project's own harness; the README does not state the machine, the input size distribution, or whether the range reflects different document shapes or different tiers.

The output side is not only Markdown. Setting output_format to "djot" emits Djot instead. Tables are GFM style with padded cells, alignment and pipe escaping. Metadata extraction parses the <head> into structured data covering Open Graph, Twitter cards, JSON-LD, microdata, RDFa and header hierarchy. There is a Visitor API, feature-gated, for transforming the converted Markdown AST, and preprocessing presets named standard, strict and lenient. Inline image mirroring of data URIs and remote references is opt-in, which is the right default: silently fetching remote images during a corpus job would be a surprise.

## Installing html-to-markdown and converting your first document

The project ships per-language packages rather than one binary you call. For Rust, the README gives the crate name directly:

```bash
cargo add html-to-markdown-rs
```

For Python the distribution on PyPI is html-to-markdown, and for Node.js the package is scoped as @xberg-io/html-to-markdown. There is also a separate WASM package, @xberg-io/html-to-markdown-wasm, which is what the browser demo runs on. The other bindings are published under their own registries: io.xberg:html-to-markdown on Maven Central for Java, github.com/xberg-io/html-to-markdown/packages/go/v3 for Go, XbergIo.HtmlToMarkdown on NuGet, xberg-io/html-to-markdown on Packagist, html-to-markdown on RubyGems, html_to_markdown on Hex, h2m on pub.dev, plus Swift via SPM, Zig, and a C ABI. The README does not give install commands for most of these, only the registry links.

The entry point is one function. The README states that convert() returns a structured result with content, warnings, and optional metadata. The warnings field is the part worth building around: a converter that repairs broken HTML has to decide what to do with markup it cannot represent, and the result object gives you a place to find out. The README does not enumerate what triggers a warning or what the warning objects contain.

Configuration is where the trade-offs live. The preprocessing presets are standard, strict and lenient, and the README says you can build your own. Output format is selected with output_format, set to "djot" for Djot. If you need to reshape the Markdown after conversion, the Visitor API is feature-gated, meaning it is not compiled in by default and the README does not show the feature flag name. That is a gap: the feature is advertised in the feature table without the flag needed to enable it.

## Sixteen bindings is a distribution problem, not a feature

The headline number is 16 languages over one Rust core, and the repository layout supports the claim: crates/ holds html-to-markdown, html-to-markdown-cli, html-to-markdown-ffi, html-to-markdown-node, html-to-markdown-php, html-to-markdown-py, html-to-markdown-rs-jni and html-to-markdown-wasm, while packages/ holds the per-language wrappers including dart, swift, elixir, r and ruby. The workspace version is pinned at 3.13.0 in Cargo.toml, while the root package.json still reads 3.13.0 and the most recent tagged release in the list is v3.12.3. Those are different numbering tracks, and that is normal for a polyglot monorepo, but it means you cannot assume the npm package and the crate are at the same version at any given moment.

This is the real cost of the design. Sixteen bindings means sixteen release pipelines, sixteen sets of platform wheels or prebuilt artifacts, and sixteen places where a version can lag. The README does not describe how the bindings are kept in sync, and the presence of a .alef-generation.toml plus an auto-generated README header ("This file is auto-generated by alef") suggests the project has invested in generation tooling to manage exactly this. That is a reasonable answer, but it shifts the risk rather than removing it: if the generator is the source of truth, a binding can be correct in the repository and stale in the registry.

The CLI crate exists in the workspace, and the repository lists a cli-proxy/ directory at top level. The README does not document CLI usage, flags, or a binary name, so if you want a shell command rather than a library call, the README is not where you will find it.

## Where html-to-markdown is the wrong choice

It handles HTML. The README does not claim PDF, DOCX, LaTeX, RTF or EPUB input, and the description of the parent project (Kreuzberg, a document intelligence engine covering 98+ formats) is a separate thing from this repository. If your pipeline needs one converter for every format you ingest, this is a component, not the whole answer.

The second boundary is output fidelity. Markdown cannot represent everything HTML can. Nested tables, styled spans, absolutely positioned layout, form controls and script-driven content have no CommonMark equivalent. The project's promise is that content is not lost, not that structure is preserved. If your downstream consumer depends on the table structure of a complex layout table, a GFM table with padded cells is a different object than what you started with.

The third is that the repair behaviour is deliberately automatic. The README states you never choose a parsing strategy or tune anything to get correct output. That is a good default and a bad debugging story. When a page converts to something unexpected, there is no documented way to force a specific tier, inspect which tier ran, or diff the tiers against each other. The warnings array is the only documented signal, and the README does not say what it contains. If you need to explain to a reviewer why a particular page produced a particular Markdown file, the tool does not currently give you the vocabulary.

Finally, the project is young in release terms. The release list shows v3.12.1, v3.12.2 and v3.12.3 all landing within three days in September 2026. Rapid patch releases after a major version bump usually mean regressions are being found in the wild. That is not a reason to avoid it, but it is a reason to pin a version rather than tracking latest.

## How it differs from Pandoc and from JavaScript-only converters

Pandoc is the obvious comparison and the difference is architectural. Pandoc is a document converter with an internal AST that many formats read into and write out of. HTML is one reader among dozens. That breadth is the point, and it comes with a different behaviour on malformed input: Pandoc's HTML reader is not built around a tiered repair strategy for scraped pages, and the README here makes no attempt to compare the two. If you need HTML in and Markdown out and nothing else, Pandoc is a larger dependency for a narrower slice of its capability. If you need HTML, LaTeX and DOCX in the same pipeline, Pandoc already does that and this project does not.

The closer comparison is to single-language JavaScript converters. Those are typically one parser, one output path, and a JavaScript runtime requirement. The difference here is the Rust core plus the tiered dispatch, which is what allows the same conversion to run in a Go service, a Python ingestion job and a browser through WASM without three separate implementations drifting apart. That consistency is the actual product. If your stack is entirely Node.js and your HTML is clean, a smaller JavaScript library will do the job with less surface area.

The Djot output is worth noting as a differentiator in the other direction. Djot is a Markdown successor format, and being able to emit it from the same converter means you can evaluate a migration without standing up a second toolchain. The README states the switch is output_format = "djot" and says nothing further about which Djot constructs are supported.

## Licence, releases and what maintenance actually costs you

The licence is MIT, stated in both the README badge and the Cargo.toml workspace metadata. That is permissive: you can use it commercially, modify it and redistribute it, provided the copyright notice and permission notice travel with it. There is an ATTRIBUTIONS.md at the repository root, which suggests bundled or vendored dependencies carry their own notices. MIT gives you no patent grant, which is the standard trade-off and worth knowing if your legal team cares about that distinction. None of this is legal advice; read the LICENSE file.

On maintenance: the repository is not archived and the last push was on 2026-09-10, nine days before this writing. The release cadence is active, with three patch releases in the first half of September 2026. The project is maintained by the Kreuzberg team, and Kreuzberg is described as a document intelligence engine with a Rust core, so this converter is a component inside a larger commercial-adjacent effort rather than a standalone hobby project. That is usually good for longevity and occasionally bad for priority, since the roadmap is set by the parent product's needs.

The upgrade cost is the part to plan for. A polyglot monorepo means bumping versions in more than one manifest, and the version numbers do not move in lockstep: the Rust workspace is at 3.13.0 while the newest tagged release in the list is v3.12.3. If you consume more than one binding, treat them as separate dependencies with separate upgrade decisions. The toolchain floor is stated: rust-version = 1.88 and edition 2024 in Cargo.toml, node >=18 and pnpm >=11 in package.json, and Python >=3.10 in pyproject.toml. Those are the constraints that will actually block an upgrade in a locked-down environment.

## Conclusion

Adopt html-to-markdown if you need one HTML to Markdown implementation shared across several runtimes, or if the HTML you feed it is broken enough that a strict parser gives up. Skip it if you need a general document converter that also reads PDF, DOCX or LaTeX; that is a different class of tool, and the README only claims HTML in. Before committing, run your own fixtures through convert() in the language you will actually ship, check the warnings array rather than assuming silence means success, and confirm the binding for your runtime is published at the version you expect, since the packages release independently.

## FAQ

### Can I convert HTML to Markdown with html-to-markdown?

Yes. The README describes a single convert() entry point that takes HTML and returns a structured result with content, warnings and optional metadata. The same call is exposed across the language bindings, from Rust and Python to Node.js, Go, Java and the WASM package.

### How do I convert HTML to Markdown with html-to-markdown in Python or Node.js?

The Python distribution is html-to-markdown on PyPI and the Node.js package is @xberg-io/html-to-markdown on npm, with a separate @xberg-io/html-to-markdown-wasm package for the browser. The README links each registry but only gives an explicit install command for the Rust crate, cargo add html-to-markdown-rs.

### What is html-to-markdown compared with Pandoc?

Pandoc is a general document converter with readers and writers for many formats, while html-to-markdown only takes HTML in and emits CommonMark or Djot. The README does not make the comparison itself; the difference visible in the repository is scope, plus the tiered repair path this project uses for malformed HTML.

### What is html-to-markdown?

It is a CommonMark-compliant HTML to Markdown converter maintained by the Kreuzberg team, built on a Rust core with bindings for 16 languages. It also extracts page metadata such as Open Graph, Twitter cards and JSON-LD during the same pass.

### What is convert html to markdown in the context of html-to-markdown?

The conversion is done by the convert() function, which the README describes as the single entry point. It returns content, warnings and optional metadata, and the same function name is used across the language bindings.

## Sources

- [License: MIT](https://github.com/xberg-io/html-to-markdown/blob/main/LICENSE)
- [Project website](https://docs.html-to-markdown.xberg.io)
- [README](https://github.com/xberg-io/html-to-markdown/blob/main/README.md)
- [Releases](https://github.com/xberg-io/html-to-markdown/releases)
- [xberg-io/html-to-markdown on GitHub](https://github.com/xberg-io/html-to-markdown)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/xberg-io-html-to-markdown
