# commonmark-java: a CommonMark parser and renderer for the JVM

> commonmark-java parses Markdown into an abstract syntax tree you can walk and rewrite, then renders it to HTML, Markdown or plain text. It suits JVM services that need spec-conformant Markdown without a browser or a Node dependency.

**commonmark/commonmark-java** — Java library for parsing and rendering CommonMark (Markdown)

- Repository: https://github.com/commonmark/commonmark-java
- Stars: 2,696 · Forks: 336
- Language: Java
- License: BSD-2-Clause
- Published: 2026-09-28 · Updated: 2026-09-28 · Language: en
- Canonical page: https://hysenlabs.com/projects/commonmark-commonmark-java

## What commonmark-java is for

Markdown in Java usually arrives as one of two problems: you have text from a user or a file and need HTML out, or you have a document model and need to inspect, transform or re-serialize it. commonmark-java addresses both by parsing input into an abstract syntax tree, exposing that tree, and then rendering it. The README describes the library as providing classes for parsing to an AST, visiting and manipulating nodes, and rendering to HTML or back to Markdown.

The README states the core has no dependencies and that extensions live in separate artifacts. It also states the library is supported on Java 11 and later, and that it works on Android on a best-effort basis with a minimum API level of 19, pointing at the commonmark-android-test directory. That combination points at server-side Java: a content pipeline, a comment renderer, a static site generator, a documentation tool. It is not a browser-side library, and there is no JavaScript build here.

## Parsing to an AST, then rendering

The pipeline is short. A Parser built from Parser.builder() turns a string into a Document node, a tree of node types such as Heading, Text and Image. A renderer walks that tree. Three renderers appear in the README: HtmlRenderer, MarkdownRenderer and TextContentRenderer for plain text with minimal markup.

The interesting part is that the tree is not sealed. You can visit nodes with a visitor derived from AbstractVisitor, count or collect things, and modify the tree before rendering. The README's word-count example overrides visit(Text) and calls visitChildren to descend. Because the same tree feeds every renderer, a transformation applied once shows up in HTML, Markdown and text output alike.

Source positions are opt-in. Parser.builder().includeSourceSpans(IncludeSourceSpans.BLOCKS_AND_INLINES) makes parsed nodes carry spans with getLineIndex, getColumnIndex, getInputIndex and getLength. The README shows a substring of the original source being recovered from a span. That is what you need for editor diagnostics or for mapping a rendered element back to a line in the input, and it is off by default, which keeps the common case cheap.

## Installing commonmark-java and rendering your first document

The README gives the Maven coordinates for the core artifact under the group id org.commonmark. Add this dependency to your build, substituting the version you intend to pin:

```xml
<dependency>
    <groupId>org.commonmark</groupId>
    <artifactId>commonmark</artifactId>
    <version>0.30.0</version>
</dependency>
```

Once the dependency resolves, the shortest useful program builds a parser and a renderer and runs a string through both. The README's example uses var, so compile against Java 11 or later:

```java
import org.commonmark.node.*;
import org.commonmark.parser.Parser;
import org.commonmark.renderer.html.HtmlRenderer;

var parser = Parser.builder().build();
var document = parser.parse("This is *Markdown*");
var renderer = HtmlRenderer.builder().build();
renderer.render(document);  // "<p>This is <em>Markdown</em></p>\n"
```

The comment on the last line is the output the README shows for that input, including the trailing newline. If you are on Java 9 or later and use the module path, the README states the module name is org.commonmark, with extension modules named after their packages, for example org.commonmark.ext.autolink.

If you clone the repository, the README mentions a DingusApp class for trying things interactively, and the repository ships mvnw, so the build runs through the Maven wrapper rather than a system Maven install.

## Sanitization is the caller's job

The README is direct about this: the library does not try to sanitize the resulting HTML with regard to which tags are allowed, that is the responsibility of the caller, and if you expose the resulting HTML you probably want to run a sanitizer after this. Two renderer options narrow the surface but do not replace a sanitizer. escapeHtml(true) escapes raw HTML tags and blocks, and sanitizeUrls(true) strips potentially unsafe URLs from a and img tags.

Treat that as a design boundary rather than a defect. A parser that silently dropped tags would be a worse fit for a documentation pipeline that legitimately embeds HTML. But it does mean that if your input is untrusted, commonmark-java is one stage of a pipeline, not the whole of it, and the sanitizer you pick is a separate decision the project does not make for you.

The other boundary is the version number. The README states that for 0.x releases the API is not considered stable and may break between minor releases, and that Semantic Versioning starts after 1.0. Packages containing beta are not subject to stable API guarantees. With releases at 0.28.0, 0.29.0 and 0.30.0 in 2026, that is a live constraint: read the CHANGELOG.md before moving a minor version, and expect to touch code, not just the version string.

## Extensions, and what they cost you

Core CommonMark stops short of tables, strikethrough and autolinking, so those ship as separate artifacts. The repository layout lists commonmark-ext-gfm-tables, commonmark-ext-gfm-strikethrough, commonmark-ext-autolink, commonmark-ext-footnotes, commonmark-ext-gfm-alerts, commonmark-ext-heading-anchor, commonmark-ext-image-attributes, commonmark-ext-ins, commonmark-ext-task-list-items and commonmark-ext-yaml-front-matter. Each is a distinct dependency, which keeps the core small but means an extension-heavy setup pulls in a list of artifacts rather than one.

There is a real trade-off in the extension model. Syntax you enable changes the meaning of documents that were previously valid, and the README does not document rollback for that. If you turn on tables and later remove the extension, previously parsed table text falls back to paragraphs, and any stored AST or rendered output reflects the older interpretation. Pin the extension set alongside the core version.

The repository also carries commonmark-test-util, which holds the spec.txt file the README points at when you want to know which version of the spec is currently implemented. That file is the authoritative answer to a conformance question, not the README prose.

## commonmark-java compared with flexmark-java

flexmark-java is the comparison that comes up, and the difference is scope. flexmark-java is a larger Markdown implementation with a much wider configuration surface, and it is the name people reach for when they want a parser with a long list of options and extension points exposed through a single dependency. commonmark-java takes the opposite approach: a dependency-free core, a small API, and extensions in separate artifacts, with the README describing the library as small, flexible and extensible.

The README also states the library is 10-20 times faster than pegdown, which used to be a popular Markdown library, and points at benchmarks in the repository. That is a claim about pegdown, not about flexmark-java, and it is a historical comparison against a library that is no longer the default choice.

The practical split: if you want a CommonMark-conformant parser whose AST you walk and rewrite, and you prefer a narrow API over a wide one, commonmark-java fits. If you need a very large set of Markdown dialects and rendering options behind one dependency, flexmark-java is the more likely fit. Neither choice is about raw speed at this point; it is about how much configuration surface you want to carry.

## Maintenance, licensing and upgrade cost

The repository is not archived, and the last push was on 2026-08-07. The most recent release listed is commonmark-java 0.30.0 on 2026-08-06, preceded by 0.29.0 on 2026-06-20 and 0.28.0 on 2026-03-31. That is a steady release cadence across 2026, with the parent artifact versioned as commonmark-parent.

The licence is BSD-2-Clause, a permissive licence that permits use in closed-source products. The repository ships LICENSE.txt at the top level, and that file is the text that governs, not the short identifier. Redistribution carries the usual obligation to keep the copyright notice and licence text with the source distribution. This is a description of what the repository contains, not legal advice; if your organisation has a licence review process, run the BSD-2-Clause text through it.

The upgrade cost is concentrated in the 0.x status. Because the API may break between minor releases, a version bump is a code change with a changelog review attached, and the extension artifacts move with the core. Budget for that on each minor release rather than treating the dependency as set-and-forget.

## Conclusion

Adopt commonmark-java if you need a dependency-free core parser whose AST you can inspect or modify before rendering, and you can pin a 0.x version because the API may break between minor releases. Do not adopt it if you expect the library to sanitize the HTML it produces, since the README states that is the caller's responsibility. Before wiring it into a release, check the spec.txt file for the spec version implemented and confirm on Maven Central that the artifact id and version you intend to use exists.

## FAQ

### What Java versions does commonmark-java support?

The README states the library is supported on Java 11 and later. It also states that it works on Android on a best-effort basis, with a minimum API level of 19.

### Does commonmark-java sanitize the HTML it renders?

No. The README states the library does not sanitize the resulting HTML with regard to which tags are allowed, that this is the responsibility of the caller, and that you probably want to run a sanitizer on the output if you expose it.

### How do I add commonmark-java with Maven?

The README gives the dependency as group id org.commonmark and artifact id commonmark, with the version you intend to pin. Extension artifacts such as commonmark-ext-autolink are added separately.

### Can commonmark-java render Markdown back to Markdown or to plain text?

Yes. The README documents MarkdownRenderer for rendering a document back to Markdown, and TextContentRenderer for plain text with minimal markup.

## Sources

- [commonmark/commonmark-java on GitHub](https://github.com/commonmark/commonmark-java)
- [Issues](https://github.com/commonmark/commonmark-java/issues)
- [License: BSD-2-Clause](https://github.com/commonmark/commonmark-java/blob/main/LICENSE)
- [README](https://github.com/commonmark/commonmark-java/blob/main/README.md)
- [Releases](https://github.com/commonmark/commonmark-java/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/commonmark-commonmark-java
