pulldown-cmark: a pull parser for CommonMark in Rust
An efficient, reliable parser for CommonMark, a standard dialect of Markdown
At a glance
- What is it?
- pulldown-cmark parses CommonMark into an iterator of events instead of a document tree. It suits Rust programs that need source offsets, low allocation, or a streaming renderer, and it is a poor fit if you want a ready-made Markdown data model.
- Who is it for?
- Adopt pulldown-cmark if you are writing Rust and want to render, filter or measure CommonMark while keeping source offsets and avoiding a full document tree; the examples directory and the published docs are the place to start. Do not adopt it if you need a mutable AST as the primary interface, or if you are not on Rust, since the crate and its CLI are Rust artifacts.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 21 days ago.
- What is it written in?
- Mainly Rust, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 24, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem pulldown-cmark solves, and for whom
Most Markdown libraries hand you a document tree. You get nodes, you walk them, you mutate them, and you pay for the tree in memory. pulldown-cmark takes the opposite position: it is a pull parser, so the caller drives the parse and receives events one at a time. The README frames this as the reason the project exists, arguing that pull parsing uses "dramatically less memory than constructing a document tree" while being easier to drive than a push parser with callbacks.
The audience is therefore narrow and specific. If you are writing a Rust program that renders Markdown to HTML, rewrites it, or measures it, and you care about allocation and source positions, this is aimed at you. If you are writing a Markdown editor that needs to mutate a tree and re-serialize it, the event stream is the wrong shape and you will end up building the tree yourself. The crate is also shipped with a small command-line tool for rendering to HTML, so the same parser is usable without writing any Rust at all.
How the event stream and source maps actually work
The parser type implements Rust's Iterator trait directly, and Event is an enum covering start and end tags, text, and the other pieces of the syntax. That means the whole iterator toolbox applies: for loops, map, and collecting events into a vector for later playback. The README gives the shape of it in one line, where Parser::new takes the input and html::push_html consumes the resulting iterator.
Source maps come from a second entry point. Calling into_offset_iter() produces an iterator that yields (Event, Range) pairs, where the range is the event's span in the source document. That is the feature that separates this crate from renderers that only ever emit a string. A tool that reports "this link is broken at line 42" needs exactly this mapping, and it is available without a separate parse pass.
Text handling deserves attention because it explains the performance story. The Text event is a small copy-on-write string, and the README states that the vast majority of text fragments are slices of the source document, so most text is copied once, from source to HTML buffer. The trade-off is that consecutive text events can occur, depending on how the parser evaluates the source. A TextMergeStream utility exists to smooth that over for callers who want merged runs of text rather than the raw stream.
Installing pulldown-cmark and rendering a first document
The crate builds with rustc 1.71.1 or newer, which the README states as the minimum. Adding it to a project is a single cargo command. The binary is built by default, so if you only want the library, disable default features as the README shows:
cargo add pulldown-cmark --no-default-featuresWith the dependency in place, the smallest useful program creates a parser and pushes HTML into a String buffer. The README's own example is this, and the assertion shows exactly what the output looks like:
let markdown_input = "hello world";
let parser = pulldown_cmark::Parser::new(markdown_input);
let mut html_output = String::new();
pulldown_cmark::html::push_html(&mut html_output, parser);
assert_eq!(&html_output, "<p>hello world</p>\n");One practical detail from the README: when using the built-in HTML renderer, write to a buffered target such as a Vec<u8> or a String. The renderer performs many very small writes, so writing straight to stdout, a file, or a socket hurts performance, and such writers should be wrapped in a BufWriter. For a build tuned for release, the workspace Cargo.toml already sets lto = true, codegen-units = 1 and panic = "abort" in the release profile, and the README repeats that configuration for downstream users. SIMD scanning on x64 is opt-in:
cargo build --release --features simdWhere the pull model gets in the way
The event stream is forward-only by design, and that is the limitation you will hit first. There is no document tree to hold on to, so any transformation that needs to look backwards or reorder blocks has to be buffered by the caller, usually by collecting events into a vector. The README acknowledges this by listing collection as one of the supported iterator uses, but collecting the whole stream into memory gives back much of what the pull model saves.
Consecutive text events are a second friction point. Code that matches on Event::Text and assumes one event per text run will be wrong on some inputs, which is why TextMergeStream exists. A renderer that ignores this will still produce correct HTML, because the HTML writer handles the fragments, but a custom consumer that does string work per event can behave subtly differently from what its author expected.
The extension set is another boundary. Footnotes, GitHub flavored tables, GitHub flavored task lists, strikethrough and highlight are all optional, and the README presents them as such. If your documents depend on an extension the crate does not implement, no amount of configuration will produce it, and you will be post-processing events yourself. Finally, SIMD acceleration is described as available for the x64 platform, so the fast path is not universal, and the crate's safety claim of no unsafe blocks carries the explicit exception of that opt-in feature.
pulldown-cmark compared with Comrak and other Rust Markdown parsers
The natural comparison is Comrak, which also parses CommonMark in Rust and appears alongside this project in search results. The difference is architectural rather than cosmetic. Comrak is built around an AST: you parse into a tree, walk and mutate it, and format it back out. pulldown-cmark gives you an iterator of events and, optionally, source ranges. If your task is to inspect and rewrite structure, the tree is more convenient; if your task is to render or measure a stream with minimal allocation, the iterator is lighter and the offsets come for free.
Against CommonMark implementations in other languages, such as commonmark-java, the dividing line is simply the host runtime. Those libraries solve the same specification problem, but they cannot be embedded in a Rust binary, and pulldown-cmark cannot be dropped into a JVM or Node project. The crate's stated goal is 100% compliance with the CommonMark spec, so the interesting differences between implementations are the extension sets and the API shape, not the core syntax. The repository is organized as a workspace with separate crates for the parser and for escaping, plus bench, fuzz and dos-fuzzer members, which tells you the maintainers treat fuzzing and benchmarking as part of the project rather than an afterthought.
Maintenance, licence and what an upgrade costs
The repository is not archived, and the last push was on 2026-09-11. The most recent release listed is v0.13.4 on 2026-05-20, preceded by v0.13.3 and v0.13.2 in March 2026. That pattern suggests patch releases land between minor bumps, but the README does not document a deprecation policy, a support window, or a migration guide for the 0.x line, so the upgrade cost is whatever the release notes say each time.
Because the crate is pre-1.0, a minor version bump can change the API, and the event enum is the surface most likely to move. Pinning an exact version in Cargo.toml and reading the release notes before bumping is the only reliable approach the project's documentation supports. The licence is MIT, which is permissive and permits use in closed-source products; the repository contains a LICENSE file at the top level, and that file, not this article, is the authoritative text. Nothing in the README describes a contributor licence agreement or a dual-licensing arrangement, so no further licence conclusion can be drawn here.
Editorial conclusion
Adopt pulldown-cmark if you are writing Rust and want to render, filter or measure CommonMark while keeping source offsets and avoiding a full document tree; the examples directory and the published docs are the place to start. Do not adopt it if you need a mutable AST as the primary interface, or if you are not on Rust, since the crate and its CLI are Rust artifacts. Before committing, verify the feature flags you need (simd, footnotes, tables, task lists, strikethrough, highlight) against your target and confirm your toolchain is at least rustc 1.71.1, the minimum the README states.
Frequently asked questions
What is pulldown-cmark used for?
It parses CommonMark, the standard Markdown dialect, into an iterator of events. The crate is designed to be used as a library in Rust, and it also ships a simple command-line tool for rendering Markdown to HTML.
How do I install pulldown-cmark in a Rust project?
Add it with cargo add pulldown-cmark, or use cargo add pulldown-cmark --no-default-features if you do not want the binary built. Building the crate requires rustc 1.71.1 or newer.
Does pulldown-cmark support tables, footnotes and task lists?
Yes, but as optional extensions. The README lists footnotes, GitHub flavored tables, GitHub flavored task lists, strikethrough and highlight as optionally supported features rather than defaults.
How do I get the source position of an event in pulldown-cmark?
Call into_offset_iter() on the parser to get an iterator that yields (Event, Range) pairs, where the range is the event's corresponding span in the source document. The README presents source maps as one of the crate's design goals.
Why does pulldown-cmark emit several Text events in a row?
The README states that consecutive text events can happen because of the manner in which the parser evaluates the source. A TextMergeStream utility exists to make iterating over the events more comfortable.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/pulldown-cmark-pulldown-cmark)