# Pandoc: more than fifty input formats and a README it generates itself

> Pandoc is a Haskell library and command line tool for converting between markup formats, released under GPL-2.0. The interesting parts of the repository are the edges: PDF is not an input, three format names are marked deprecated, the expected-output files for every conversion are committed, and the README you are reading was written by pandoc through a Lua filter.

**jgm/pandoc** — GitHub describes it as Universal markup converter. The repository metadata lists Haskell as its primary language. The metadata lists the GPL-2.0 license. This article stays within the project description and details documented in the GitHub repository README.

- Repository: https://github.com/jgm/pandoc
- Website: https://pandoc.org
- Stars: 46,471 · Forks: 5,341
- Language: Haskell
- License: GPL-2.0
- Published: 2026-08-13 · Updated: 2026-08-18 · Language: en
- Canonical page: https://hysenlabs.com/projects/jgm-pandoc

## The README is generated by pandoc, from a template and a Lua filter

The file describing pandoc was written by pandoc. Its opening lines are a warning not to edit it by hand, because it is generated from README.template and MANUAL.txt through a Lua filter:

```bash
pandoc --lua-filter tools/update-readme.lua README.template -o README.md
```

That one line settles three questions about the project. The format lists are not maintained by hand, they come out of the same source as the manual, so a format added to MANUAL.txt appears in the README on the next regeneration. The tools directory is real, with update-readme.lua inside it, and custom Lua filters are a first-class extension point rather than a community workaround. And the documentation pipeline is a demonstration of what the tool is for: read a markup format, transform it, write another. If you want the format list in editable form, open README.template rather than README.md.

## PDF is absent from the input list, so the reverse conversion needs a different tool

The input list runs from asciidoc and bibtex through docx, epub, fb2, html, ipynb, jats, json, latex, native, odt, opml, org, pptx, rtf, rst, typst and xlsx to xml, more than fifty names in all, and it finishes by pointing at a custom Lua reader. PDF is not in that list. Bibliographies are covered, with bibtex, biblatex, csljson, endnotexml and ris, and tables are covered, with csv and tsv.

So pulling a document out of a PDF is outside what this tool does, and it fails quietly rather than loudly: a wrong format name in a command produces a complaint about the format, not a hint about PDFs. If your source is a PDF you need an extraction step first, and only then does pandoc do the part you want, which is reformatting. On the writing side the same file lists beamer, context and latex, which are TeX formats an external engine compiles into a PDF, so a PDF stays a downstream step rather than a direct target.

The closest readers for something that used to be a PDF are docx, odt, epub, fb2, rtf, html and org, and which of those you have depends on whatever produced the original.

## markdown_github, asciidoctor and six bbcode dialects carry the deprecation warnings

Format names encode status, and the list marks three of them. markdown_github is described as deprecated and less accurate than gfm, with an instruction to reach for it only when you need extensions gfm does not support. asciidoctor is a deprecated synonym for asciidoc, so an old script keeps running while writing output from a different writer than you assumed.

The bbcode family shows the same problem at larger scale: bbcode on its own, then bbcode_fluxbb, bbcode_phpbb, bbcode_steam, bbcode_hubzilla and bbcode_xenforo. Six names for one forum markup, each a dialect. Pick the wrong one and the conversion finishes without complaint, because the tags are similar enough to parse, and the output is wrong in a way that only shows up after posting. asciidoc against asciidoc_legacy is the same trap with a wider gap, since those two are read by different implementations, Asciidoctor and asciidoc-py.

For a script, treat the deprecated names as work waiting to happen. The status is written next to the name, so you never need release notes to find it.

## native, json and xml expose one AST, which is how you debug a lossy conversion

Three entries in the format list are not markup languages. native is the native Haskell representation, json is the JSON version of that same AST, and xml is its XML version. The custom readers and writers section extends the idea by letting you give the path of a Lua reader as a format name.

This changes how a broken conversion gets diagnosed. When a table disappears, a citation is mangled or a nested list flattens, reading the output is guesswork, because you cannot tell whether the reader lost the structure or the writer dropped it. Converting to json and inspecting the tree answers that. If the tree holds the structure and the output does not, the writer or the template is at fault; if the tree is already wrong, the source document or the reader is. Either way you stop reading rendered text and start reading data.

A second point about the filter API: it is not an afterthought. The project builds its own README with one, which is the strongest evidence available that filters are a supported surface rather than a private trick.

## cabal, stack and nix coexist, and the compiler image is pinned in the Makefile

The top level carries cabal.project, stack.yaml, flake.nix, release.nix, shell.nix, hie.yaml, .hlint.yaml, .stylish-haskell.yaml and weeder.toml, next to a Makefile that drives cabal. Several build environments are supported at once, and the README does not say which one a contributor should use.

What the Makefile does pin is the compiler: the build image is quay.io/benz0li/ghc-musl:9.10, a GHC 9.10 build on musl, so the toolchain is fixed by a variable rather than by whatever your distribution ships. The default cabal options are written out too, including --disable-optimization along with the -fhttp and -f-export-dynamic flags, which is how you can tell the HTTP server is part of the build rather than an optional extra. The executable path comes from cabal list-bin instead of a hardcoded bin directory, and profiling is a separate target built with cabal build --enable-profiling all.

For a packager this matters: the compiler you build against is chosen by that image, not by the release tag.

## Golden tests carry an accept flag, so a behaviour change lands as a bulk file diff

Converting documents is work whose correctness is defined by example, and the repository treats it that way. A note above the test target in the Makefile tells you how to accept the current results as the new expected output:

```bash
make test TESTARGS='--accept'
```

The mechanism cuts both ways. It is how the project holds conversions honest across fifty formats, and it is also how a change to shared parsing code rewrites the whole expected-output set in a single commit. The tests stay green either way, and a reviewer is left with a diff of generated files in which the one conversion that actually regressed is hard to pick out.

So when you evaluate a change to this codebase, do not stop at the passing run. Find the files that moved under test/ and read the ones covering the formats you depend on. That directory is the record of what each format pair is supposed to produce, and for your specific conversion it is the most useful thing in the repository.

## Benchmarks diff against the newest CSV and abandon a case at six seconds

Performance is measured in the repository instead of claimed. The Makefile writes a timestamped bench file, takes the most recent bench_*.csv as the baseline, and passes --baseline with it when one exists. The run gets --timeout=6 plus the runtime options +RTS -T --nonmoving-gc -RTS, so every case has a time limit and the garbage collector is held in its non-moving mode to keep timings comparable. A PATTERN variable narrows a run, and a benchmark directory sits at the top level.

Two consequences for anyone reading the output. A figure is a difference against a stored file, not an absolute number, so a result without its baseline CSV beside it is close to meaningless, and a stale baseline quietly becomes the thing you are measured against. And the six second limit truncates a slow case rather than reporting it as slow, which is right for a test suite and wrong for a capacity question.

The single machine image and the RTS flags together also place a limit on the claim: this is a measurement of the machine running the make target.

## GPL-2.0, a COPYRIGHT file, and three releases in five weeks

The licence is GPL-2.0, with COPYING.md and COPYRIGHT committed at the top level next to AUTHORS.md, CITATION.cff and release-announcements.txt. The README does not discuss embedding the library in a product or redistributing a modified binary, so if that is your plan the licence file is the place to read, not this page.

The cadence is quick. Version 3.10.2 was released on 2026-08-12, 3.11 on 2026-08-29 and 3.12 on 2026-09-29, and the last push to the main branch was on 2026-09-29. For a tool sitting in a document pipeline, a monthly release means your expected-output files and your habits both have to keep pace, and 3.12 landing on the same day as the last push tells you the tag is fresh rather than abandoned.

The housekeeping files are the other signal: CONTRIBUTING.md, SECURITY.md, BUGS, changelog.md, MANUAL.txt as the manual's source, and RELEASE-CHECKLIST-TEMPLATE.org for cutting a release.

## Conclusion

Adopt pandoc if your documents already move between Markdown, LaTeX, docx and HTML and you want one deterministic converter plus a format list you can read. Do not adopt it expecting a PDF reader, a Python package, or fidelity guarantees for a specific office format, since none of those exist here. Before you commit, check two things: the one conversion you actually depend on, by reading its expected-output file in test/ rather than trusting a sample run, and the licence, because the project is GPL-2.0 with COPYING.md and COPYRIGHT at the root, which matters if you redistribute a binary you built rather than invoke a released one.

## FAQ

### What is pandoc used for?

It converts between markup formats. The input list covers more than fifty names, including Markdown variants, LaTeX, docx, epub, html, jats, typst and xlsx, and the writers include beamer, context, docbook and docx. The project is a Haskell library with a command line tool built on it.

### Is pandoc free?

Yes, under the GPL-2.0, with COPYING.md and COPYRIGHT at the repository root. It is also distributed through Hackage, Stackage, the Homebrew formula and GitHub releases, and the manual that the README is generated from is MANUAL.txt in the same tree.

### Is pandoc a Python library?

No. The primary language is Haskell, and the project describes itself as a Haskell library for converting between markup formats plus a command line tool that uses it, built with cabal and stack. The repository contains no Python package.

### How safe is pandoc?

The repository ships a SECURITY.md for reporting problems, and the distribution channels named in the README are GitHub releases, the Hackage package, the Homebrew formula and the Stackage LTS package, which is where you confirm what you install. The README does not describe a security model or any sandboxing of the conversion itself.

### How do I install pandoc?

The README prints no install command. Installation is handled by INSTALL.md at the repository root, and the header of the file points to the distribution channels: GitHub releases, the Hackage package, the Homebrew formula and the Stackage LTS package.

### How do I convert Markdown to a Word document with pandoc?

docx appears in both the input and the output list, so Word documents can be read and written. The command line syntax and the format flags are not shown in the README, which points to the manual instead, and that manual is generated from MANUAL.txt at the repository root.

## Sources

- [Official documentation](https://pandoc.org)
- [Official README](https://github.com/jgm/pandoc#readme)
- [Project repository](https://github.com/jgm/pandoc)
- [Release notes](https://github.com/jgm/pandoc/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/jgm-pandoc
