Model or dataset
harlan-zw/mdream avatar
harlan-zw/mdream

mdream: an HTML to Markdown converter tuned for LLM token budgets

☁️ The fastest HTML to markdown convertor on GitHub. Optimized for LLMs and supports streaming.

963 stars61 forksTypeScriptMIT

At a glance

What is it?
mdream is a zero-dependency HTML to Markdown converter from harlan-zw, shipped as a Rust NAPI engine, a pure JS engine, a crawler and a Docker image. Its selling point is token efficiency and streaming, not fidelity to every HTML edge case.
Who is it for?
Adopt mdream if your pipeline already produces HTML and you need Markdown that costs fewer tokens, especially at site scale where the crawler and @mdream/action packages do the work for you. Skip it if you need a general-purpose DOM sanitizer or a converter with a long tail of custom element rules; the preset-driven design is deliberately narrow.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 1 day ago.
What is it written in?
Mainly TypeScript, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem mdream targets: HTML that costs too many tokens

Feeding raw HTML into a language model is expensive. Tags, attributes, inline styles and navigation chrome all consume context window without carrying meaning. The mdream README frames the project around exactly that gap: it describes itself as a zero-dependency, LLM-optimized converter whose output is tuned for token efficiency, and claims up to 2x fewer tokens than Turndown, node-html-markdown and html-to-markdown, and 70-99% fewer tokens than raw HTML. Those numbers come from the project's own benchmark directory, so treat them as the author's measurements rather than an independent result. The audience is developers building retrieval pipelines, documentation mirrors, or `llms.txt` artifacts who already have HTML and want compact Markdown on the other side. It is not aimed at people converting HTML for visual fidelity in a rendered document.

Two engines, one API: Rust NAPI, WASM and the pure JS core

The repository is a pnpm workspace with a `packages/` directory and a separate `crates/` tree, which tells you the converter exists in more than one implementation. The main `mdream` package bundles a Rust NAPI engine for Node plus a WASM build for edge runtimes, and the README lists its size as 60kB gzip with the Rust WASM engine. `@mdream/js` is the pure JavaScript engine, described as having full hook access and zero native dependencies, with tree-shakable conversion available through a `/core` entry point. That split matters for deployment: the JS package is the one you pick when a native binary is awkward, at the cost of the speed advantage the Rust path provides. The README puts the core JS bundle at 10kB gzip. Conversion is driven by presets, and the repository points to a `minimal` preset source file, which is the preset used throughout the README examples. Streaming is listed as a first-class feature, intended for large documents and real-time pipelines, so output can be consumed incrementally rather than buffered whole.

Installing mdream and converting a real page from the command line

The README's own first example pipes a URL straight into the CLI through `npx`, so no install step is required to try it. You can also install the package globally or as a project dependency from npm under the name `mdream`.

bash
curl -s https://en.wikipedia.org/wiki/Markdown | npx mdream --preset minimal

This fetches the Wikipedia Markdown article and prints Markdown to stdout. The `--preset minimal` flag selects the minimal preset referenced in the package source. Expect headings, links and lists in the output, and expect presentational markup to be dropped.

Relative links and images break when you convert a page in isolation, which is why the README adds an origin flag. Combining it with `tee` writes the stream to a file while you watch it:

bash
curl -s https://en.wikipedia.org/wiki/Markdown \
 | npx mdream --origin https://en.wikipedia.org --preset minimal \
  | tee streaming.md

The README states that `--origin` fixes relative image and link paths. For local files, the same pattern applies to a file on disk piped into the CLI, with `tee` writing the result to a `.md` file. The README also suggests piping the output into `glow` if you want to read it in a terminal.

Where mdream is the wrong tool

Preset-driven conversion is a trade-off, not a free win. The minimal preset strips markup aggressively, and the README's token comparisons are made against that style of output. If your downstream consumer needs tables preserved exactly, or relies on attributes that the preset drops, you will spend time on hooks in `@mdream/js` to recover behaviour the default path removes. The README does not document a rollback or round-trip path, so there is no supported way to reconstruct the original HTML from mdream output. The performance claims are also engine-specific: the 1.8MB in roughly 5.2ms figure is attributed to the Rust engine, and the README's own comparison separates Rust NAPI from JS and Rust from Rust. A pure JS deployment should not assume those numbers. Finally, the crawler package pulls in Playwright Chrome for its Docker image, which is a much heavier dependency than the converter itself and changes the operational profile of a deployment.

Turndown and node-html-markdown: a different conversion model

Turndown is the reference point the README benchmarks against, and the related searches show people arriving with `TurndownService` in mind. The architectural difference is configuration style. Turndown exposes a service object you extend with rules per element, which gives fine control over how any tag becomes Markdown, and it runs in the browser as well as Node. mdream inverts that: you choose a preset, and the presets encode the opinions about what to keep. The README positions mdream as faster and leaner than Turndown, node-html-markdown and html-to-markdown, and its output as more token-efficient. If your existing pipeline already has a set of Turndown rules encoding your house style, migrating means re-expressing that intent as preset selection plus hooks, and the README does not claim rule-for-rule parity. node-html-markdown sits in the same benchmark group and is also a pure JS converter, so the practical choice there is bundle size and output shape rather than runtime.

Crawling a whole site and generating llms.txt artifacts

The converter is one package among several. `@mdream/crawl` is described as a site-wide crawler that generates `llms.txt` artifacts from entire websites, and `@mdream/action` produces `.md` and `llms.txt` artifacts from static `.html` output inside GitHub Actions. `@mdream/vite` generates `.md` files for Vite sites during the build, and the repository ships `examples/nextjs/`, `examples/nuxt/` and `examples/vite-ssr-vue/` directories. There are also two Dockerfiles at the repository root, `Dockerfile.core` and `Dockerfile.crawl`, and the README describes pre-built images: a small core converter and a crawl image that includes Playwright Chrome. A native Rust crate with its own CLI is published on crates.io, and the README mentions browser usage through unpkg or jsDelivr with no build step. The practical consequence is that you can pick the integration point that matches your stack instead of wrapping the converter yourself.

Licence, maintenance and what an upgrade actually costs

The repository is MIT licensed, and the root `package.json` carries `"license": "MIT"` with a `LICENSE.md` file at the top level. MIT is permissive, so embedding the converter in a closed product is not the friction point; the friction is attribution and the usual warranty disclaimer, which is a question for your own legal review rather than something this article can settle. On maintenance, the last push was on 2026-09-08 and the most recent release is v1.7.2 from the same day, with v1.7.1 on 2026-09-02 and v1.7.0 on 2026-08-19. The repository is not archived. Upgrade cost is shaped by the workspace layout: the release script runs `bumpp` across `packages/*/package.json` and executes a script that syncs the Cargo version, so the npm packages and the Rust crate are versioned together. If you pin the crate and the npm package independently, expect to move both. The v1 release notes are linked from the README, which is where breaking changes from the pre-1 line would be recorded.

Editorial conclusion

Adopt mdream if your pipeline already produces HTML and you need Markdown that costs fewer tokens, especially at site scale where the crawler and @mdream/action packages do the work for you. Skip it if you need a general-purpose DOM sanitizer or a converter with a long tail of custom element rules; the preset-driven design is deliberately narrow. Before committing, run the same page through the minimal preset and the default preset and compare the output, because that choice changes what survives conversion more than any other setting.

Frequently asked questions

How do I install mdream?

You do not need to install anything to try it: the README pipes a URL into `npx mdream --preset minimal`. For repeated use, install the `mdream` package from npm, or use the Rust crate published on crates.io if you are working in Rust.

Does mdream work in the browser without a build step?

Yes. The README lists browser CDN usage through unpkg or jsDelivr as one of the supported targets, alongside the Node, edge, CLI, Docker and GitHub Actions paths.

What is the difference between the mdream package and @mdream/js?

The `mdream` package bundles the Rust NAPI engine plus WASM for edge, and is described as performance-first with declarative config. `@mdream/js` is the pure JavaScript engine with full hook access and no native dependencies, and it offers tree-shakable conversion through a `/core` entry point.

Can mdream convert an entire website to Markdown?

The `@mdream/crawl` package is a site-wide crawler that generates `llms.txt` artifacts from entire websites, and the README also documents a Docker crawl image that includes Playwright Chrome.

Official sources

  1. harlan-zw/mdream on GitHub
  2. License: MIT
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/harlan-zw-mdream.svg)](https://hysenlabs.com/projects/harlan-zw-mdream)