# html-to-markdown: a Go HTML converter built around plugins

> A Go library that turns HTML into readable Markdown, with a plugin registry, per-tag rendering control, and its own documented rules for escaping.

**JohannesKaufmann/html-to-markdown** — ⚙️ Convert HTML to Markdown. Even works with entire websites and can be extended through rules.

- Repository: https://github.com/JohannesKaufmann/html-to-markdown
- Website: https://html-to-markdown.com
- Stars: 3,822 · Forks: 225
- Language: Go
- License: MIT
- Published: 2026-10-07 · Updated: 2026-10-07 · Language: en
- Canonical page: https://hysenlabs.com/projects/johanneskaufmann-html-to-markdown

## One function for the easy case, a converter for everything else

The whole library has a two-level API, and the split is the first thing worth understanding. `ConvertString` is the small wrapper, and it wires up a default converter with the base and commonmark plugins already registered:

```go
package main

import (
	"fmt"
	"log"

	htmltomarkdown "github.com/JohannesKaufmann/html-to-markdown/v2"
)

func main() {
	input := `<strong>Bold Text</strong>`

	markdown, err := htmltomarkdown.ConvertString(input)
	if err != nil {
		log.Fatal(err)
	}
	fmt.Println(markdown)
	// Output: **Bold Text**
}
```

When you want control, you build the converter yourself with `converter.NewConverter` and pass plugins explicitly. The README is unusually firm about a detail that trips people up: if you construct the converter directly, you are responsible for registering the base and commonmark plugins, or the output will quietly lack basic formatting. That warning matters because the failure is silent rather than an error.

Version 2 is the current line and lives on `main`. The v1 code is preserved on a separate `v1` branch, so the major version bump is a real fork in the history rather than a rename. The module path carries the suffix, which means the import is `github.com/JohannesKaufmann/html-to-markdown/v2` and not the shorter v1 path. Getting that import path right is the only mechanical step between you and a working program.

## Tag types decide whether a node collapses at all

Whitespace handling is where most HTML to Markdown converters fall apart, because it depends on knowing whether a node is a block or an inline element. html-to-markdown makes that knowledge explicit through a tag type registry rather than hiding it in a lookup table. A `collapse/` directory in the tree name suggests where the work happens.

You can declare that a tag should be removed from the output entirely, or register a renderer for a specific tag at a specific priority:

```go
conv.Register.TagType("nav", converter.TagTypeRemove, converter.PriorityStandard)

conv.Register.RendererFor("b", converter.TagTypeInline, base.RenderAsHTML, converter.PriorityEarly)

conv.Register.RendererFor("article", converter.TagTypeBlock, base.RenderAsHTMLWrapper, converter.PriorityStandard)
```

Those three lines cover the interesting cases. `nav` disappears. `b` is treated as inline and rendered as raw HTML, so `<b>` survives into the output. `article` is a block whose children are converted to Markdown while the wrapper tag itself is kept, which is what `RenderAsHTMLWrapper` means in practice.

Priority is the mechanism for overriding defaults, and the README notes that some tags are removed automatically out of the box, with `<style>` as the example. Registering a different renderer at `PriorityEarly` keeps a tag that would otherwise be dropped. For anyone building custom elements or web components, this registry is the extension point that matters, because a custom element has no tag type until you tell the converter what it is.

## Escaping gets its own document because it is genuinely hard

Markdown has a handful of characters that mean something structural, and the correct decision about whether to escape one depends on context. A hyphen in the middle of a sentence and a hyphen at the start of a line are not the same problem. The project treats this as a first-class concern rather than an afterthought, and ships a dedicated `ESCAPING.md` at the root of the repository.

The README describes the behaviour as smart escaping: special characters are escaped only when necessary, to avoid accidental Markdown rendering. That is a real design choice with a real cost. Over-escaping makes output that looks littered with backslashes, and under-escaping produces Markdown that renders differently from what the source said. Both failure modes are visible to whoever reads the result, which is why documenting the rule matters more here than in most converters.

There is a second, quieter escaping decision in the plugin table. In v2.5.0 the changelog records a change where the base plugin stopped escaping the ampersand character. Before that release an ampersand came out as an escape sequence, after it it does not. For a converter whose entire job is producing readable text, that is the kind of change worth knowing about before you blame your own pipeline.

## The published plugin list is short, and honest about the gaps

The README tabulates the plugins that live in the `plugin/` directory. Base implements shared behaviour such as removing nodes. Commonmark implements the Markdown output rules against the CommonMark specification. Strikethrough converts `<strike>`, `<s>` and `<del>` to the `~~` syntax. Table implements tables with alignment, rowspan and colspan support.

Then there are two rows that say planned. GitHubFlavored and TaskListItems are both listed that way, which is more informative than leaving them out. A reader can see that the author intends to cover GitHub's dialect and did not get there yet, rather than guessing from an absence.

The table plugin is the one with the most visible effect on real documents, since HTML tables are common in CMS output and in anything pasted from a word processor. rowspan and colspan are the details worth checking against your own input, because most quick converters flatten those and lose the structure. The v2.5.0 release added an option to remove padding from table cells, which suggests the plugin has been refined against real output rather than shipped once and left alone.

## What the dependency list reveals about the scope

The `go.mod` file is more informative about intent than the marketing copy. It requires Go 1.25.0, so this is not a library you can drop into an older toolchain without upgrading.

The dependency list mixes three kinds of work. `golang.org/x/net` provides the HTML parser itself. `andybalholm/cascadia` is a CSS selector parser, which is what makes the hosted demo and REST API at html-to-markdown.com able to accept a URL and select part of a page rather than converting everything blindly. `yuin/goldmark` is a second Markdown implementation, sitting in the dependency tree as the reference the commonmark plugin is written against and tested against.

The rest explain the CLI. `muesli/termenv` handles colour in terminal output, `bmatcuk/doublestar` handles glob patterns for file arguments, and `agnivade/levenshtein` computes string distance, which fits a converter that needs to judge how close two pieces of text are. `sebdah/goldie` is a golden file testing library, and its presence alongside `convert_test.go` at the repository root suggests the output format is pinned by fixtures rather than eyeballed.

The repository tree backs this up. Alongside `cli/`, `converter/`, `collapse/` and `marker/` there are three example directories, `basics`, `options` and `register`, which map almost exactly onto the three ways of using the library described above. There are also `WRITING_PLUGINS.md`, `SECURITY.md` and a `.goreleaser.yaml`, so plugin authoring and releases are both taken seriously enough to have documentation and automation.

## Where the project stands on activity and releases

The repository is not archived, and the last push was on 2026-08-03. The most recent release is v2.5.2, published on 2026-06-07, and the changelogs since v2.5.0 read like a project responding to specific reports rather than one doing periodic tidying.

v2.5.1, from 2026-05-07, fixed a panic on an empty text node and added a fallback version info path. A panic on empty text is exactly the class of bug that only shows up when you run against a large corpus of real pages rather than the examples in the README, and fixing it in a patch release rather than waiting for a minor one is the right call. v2.5.2 then added an FAQ entry about character encoding, which is a documentation change attached to a release, suggesting the maintainer uses releases as a container for whatever needs recording.

The homepage hosts a demo and a REST API alongside the library, which means the project has three delivery shapes: a Go dependency, a CLI, and a hosted service. That is a wider surface than most converters maintain, and it is worth knowing which one you are using before you judge the output, since each goes through the same converter but exposes different defaults. With 3,822 stars, 225 forks and 27 open issues, the ratio of open issues to forks suggests an active repository with a real user base rather than a wide but shallow one.

## Conclusion

html-to-markdown is at its best when the HTML comes from somewhere you do not control, a CMS export or a scraped page, and the output still has to look like something a person wrote. The plugin registry and the tag type system are the reason it handles messy input without turning the output into noise, and the ESCAPING.md document is unusually honest about the tradeoffs that creates. What it does not yet do is cover the whole GitHub flavour of Markdown: the plugin table lists GitHubFlavored and TaskListItems as planned rather than shipped. Start with ConvertString for a first look, then move to NewConverter once you need a specific tag removed, kept, or rendered as HTML.

## FAQ

### Can I convert HTML to Markdown?

That is what the library does, through either the `ConvertString` wrapper or a converter you build yourself with `converter.NewConverter`. It also runs as a CLI, and the maintainer hosts a demo and a REST API at html-to-markdown.com if you would rather not write any Go at all.

### Which Go module path do I import for v2?

Import `github.com/JohannesKaufmann/html-to-markdown/v2`, which is what the module path in `go.mod` declares. The v1 code is on a separate `v1` branch rather than in the main line, so the two versions are not interchangeable imports.

### How do I stop a tag such as nav from appearing in the output?

Register it with a tag type of remove, for example `conv.Register.TagType("nav", converter.TagTypeRemove, converter.PriorityStandard)`. The same registry accepts a renderer instead, which is how you keep a tag as raw HTML or keep a wrapper while converting its children.

### Does html-to-markdown support tables and strikethrough?

Tables are handled by a dedicated plugin with support for alignment, rowspan and colspan, and strikethrough converts `<strike>`, `<s>` and `<del>` to the `~~` syntax. The plugin table in the README lists GitHubFlavored and TaskListItems as planned rather than available.

## Sources

- [JohannesKaufmann/html-to-markdown on GitHub](https://github.com/JohannesKaufmann/html-to-markdown)
- [License: MIT](https://github.com/JohannesKaufmann/html-to-markdown/blob/main/LICENSE)
- [Project website](https://html-to-markdown.com)
- [README](https://github.com/JohannesKaufmann/html-to-markdown/blob/main/README.md)
- [Releases](https://github.com/JohannesKaufmann/html-to-markdown/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/johanneskaufmann-html-to-markdown
