# markit: one Rust engine behind a CLI and SDK for turning documents into markdown

> A converter that reads the text layer of PDFs and handles DOCX, PPTX, XLSX, HTML, EPUB, notebooks, feeds and images through a single bundled engine, with no OCR and no model calls. The README publishes its own benchmark numbers against five competing parsers.

**shift-labs-ai/markit** — 🖍️ Convert anything to markdown. Mark it.

- Repository: https://github.com/shift-labs-ai/markit
- Stars: 1,325 · Forks: 53
- Language: Rust
- License: MIT
- Published: 2026-10-07 · Updated: 2026-10-07 · Language: en
- Canonical page: https://hysenlabs.com/projects/shift-labs-ai-markit

## One engine, nineteen ways in, and no model in the loop

The pitch is a single sentence: convert anything to markdown, and it means literally that. The supported formats table runs from PDF, Word, PowerPoint and Excel through EPUB, Jupyter notebooks, RSS and Atom feeds, CSV, JSON, YAML, XML and SVG, plus images and audio where the extraction is metadata only rather than content recognition.

The architectural commitment is stated plainly in the SDK section: every format uses the same bundled Rust engine, and on supported macOS and Linux platforms there is no second fallback implementation. That is the sentence to hold on to. Most converters in this space are a pile of per-format Python libraries with different failure modes, and the honest ones will tell you that a given file type is slower or lower fidelity. Markit has made the opposite bet, that one engine handles the range, and the consequence is that you cannot fall back to a better tool for the one format it handles badly.

The scope is deliberately bounded in one more direction: these are non-OCR parsers. They read the text layer already inside a PDF, with no OCR and no model calls. That is a limitation you should decide on early rather than discover late.

## Getting it installed and your first conversion

Installation is a single global npm package:

```bash
npm install -g @shiftlabs/markit
```

The quick start block then shows the shape of usage, one line per input kind:

```bash
# Documents
markit report.pdf
markit document.docx
markit slides.pptx

# Data
markit data.csv
markit config.json
markit schema.yaml
```

Web inputs are treated the same way as files, including a Wikipedia page, and media inputs return metadata, so `markit photo.jpg` gives you EXIF and `markit recording.mp3` gives you audio tags rather than a transcription.

The flags that matter in day-to-day use are `-o` to write to a file, `-i` to extract images into a directory, and `--no-images` to skip extraction entirely. Input also comes from stdin, which is how you would wire it into a pipeline:

```bash
markit report.pdf -o report.md
```

`markit formats` lists what is supported, and the CLI reference documents `-` for stdin and `--json` for structured output. Note the Node constraint in `package.json`: the engines field requires `^20.19.0 || >=22.12.0`, so if you are on Node 18 or on a Node 20 release below 20.19, the package is not for you without a newer runtime.

## Reading the benchmark tables without believing them blindly

The README publishes two comparison tables and is refreshingly specific about their construction. The primary one is non-OCR PDF to markdown, measured on one machine with image extraction off for every parser. On AllenAI's public olmOCR-bench, markit scores 45.0% against 38.8% for liteparse and 31.2% for anydoc. On the project's own shitty-pdf-bench, which it describes as 40 hash-pinned public PDFs spanning 50,884 pages of semiconductor manuals, standards, RTL documents, scans and malformed files, markit recovers 100.0% of text against 83.2% and 91.0%, and converts in 23.2 seconds against 55.4 and 322.2.

The wider field table adds pypdf, pdfplumber, markitdown and pymupdf4llm on a 16-document subset of 3,476 pages. Markit reports 100.0% text recovered and 3,612 pages per second, where the nearest competitor manages 486 pages per second and pymupdf4llm manages 11.

Read that carefully. Two of the five competitors in the second table are Python libraries and the third-party benchmark in the first is the project's own construction, so the framing favours the engine being measured. What the tables do establish is that the approach is not a model call in disguise, which is the honest question. The README is also straightforward that the method, provenance and per-category results live in `rust/BENCHMARKS.md`, with the reproducing harness in `benchmark/`, and both directories are in the tree. If you rely on this, run that harness on your own files rather than on the benchmark's.

## Markdown source discovery before HTML conversion

Version 0.5.3 changed how URL inputs are handled, and the change is more interesting than the patch version suggests. Instead of converting the HTML a docs site returns, markit now tries to fetch the raw markdown that site already publishes. Four detection methods run in priority order: an `Accept: text/markdown` header check on every request, which covers Cloudflare and Vercel deployments, then `<link rel="alternate" type="text/markdown">` tags found in the HTML, then VitePress marker detection that fetches `.md` source files, then `/llms.txt` for root URLs.

The design constraint stated in the notes is that no extra requests happen for normal sites. The second fetch only fires once a markdown source is confirmed in the response, so the cost of the feature on sites that do not offer markdown is one header and nothing more.

This is the difference between a parser and a converter for anyone who scrapes documentation. Markdown source is already structured, so you skip the entire HTML-to-prose problem, and heading levels and code fences survive without heuristics. The CLI reference also lists Wikipedia as a special case with main-content extraction, and the formats table calls the URL path a fetcher with markdown negotiation.

## Distribution as seven npm packages and a version-bump ritual

The distribution section explains something that trips people up with native tooling: this is not one npm package. Releases publish `@shiftlabs/markit` plus six scoped native binary packages covering macOS and Linux on x64 and ARM64, against glibc and musl, and npm selects the matching binary through `optionalDependencies`. The repository tree shows how the pieces are separated: `rust/` for the engine, `src/` for the TypeScript side, `npm/` for the per-platform packages, and `native.cjs` as the loader.

The previous `markit-ai` package is now a deprecated compatibility shim that re-exports the SDK, which matters if you inherited code depending on it. The shim still gets published on every release.

The release process itself is documented in the README as a numbered procedure, and step two is a verification command that runs a surprising amount:

```bash
git tag v0.6.0
git push origin v0.6.0
```

Tag and push is step three. Step one is updating the same version string in three places at once, `package.json`, `rust/Cargo.toml` and `npm/*/package.json`, and step two is `bun run verify`. That script, defined in `package.json`, chains the Biome check, the Rust fmt and clippy pass with warnings denied, `cargo test`, the native build, `bun test` and a TypeScript build. The workflow publishes platform packages first through `napi prepublish`, then the main package, then the shim, and finally creates the GitHub release. It needs an `NPM_TOKEN` automation token in the repository secrets.

## Two loose ends in the release history worth knowing

Version 0.6.1 and version 0.6.0 were both published on 2026-08-17, roughly half an hour apart, and the repository's last push carries the same date. That suggests a hurried follow-up rather than two planned releases.

The comparison link attached to 0.6.0 runs from v0.2.0 rather than from v0.5.3, so the GitHub-generated changelog for that release spans every intermediate version rather than just the last step. If you are auditing what changed, read the commits rather than that generated diff.

There is a clearer signal in the 0.5.3 notes. Their comparison link points at `github.com/Michaelliv/markit`, a different owner than the repository they now live in at `shift-labs-ai/markit`. Read together with the `markit-ai` deprecation shim, that points to a project that was renamed or transferred and repackaged under a new scope, while older links elsewhere still point at the previous owner. For anyone pinning a version or vendoring the source, that matters: confirm the package scope you actually depend on rather than trusting a changelog URL.

The tree also carries `AGENTS.md` and `SKILL.md`, which suggests the project is intended to be usable by coding agents as well as by people.

## Conclusion

markit is worth a look when your pipeline needs deterministic, local conversion of documents into markdown and you cannot send file contents to a hosted service. Its own published results put it well ahead of the Python parsers on text recovery and page throughput, but those numbers come from the project and the harness, so the useful test is your own document mix: mostly clean born-digital PDFs will confirm them, and scans, semiconductor manuals or malformed files will not. Check two things before adopting. The engines field requires Node `^20.19.0 || >=22.12.0`, which excludes several long-term-support Node lines people still run. And the conversion is strictly non-OCR, so any scanned document yields an empty or near-empty result no matter how fast the parser is. Start with `markit formats` to see the format list, then run `markit <file> --json` on a few representative files and check what came back.

## FAQ

### What formats can markit convert to markdown?

Nineteen kinds, including PDF, DOCX, PPTX, XLSX, HTML, EPUB, Jupyter notebooks, RSS and Atom feeds, CSV, TSV, JSON, YAML, XML, SVG, ZIP archives, source code and plain text. Images and audio are handled by metadata extraction rather than content recognition.

### Does markit use OCR or an AI model to read PDFs?

No. It is a non-OCR parser that reads the text layer already inside a PDF, with no OCR and no model calls. Scanned documents with no text layer are the case the project explicitly does not target.

### How do I use markit as a Node library?

Install `@shiftlabs/markit` and use the `Markit` class, which exposes `convertFile`, `convertUrl` and `convert`, each returning an object with a `markdown` property. Every format goes through the same bundled Rust engine on supported macOS and Linux platforms, with no separate fallback implementation.

### What Node.js version does markit require?

The engines field in `package.json` requires `^20.19.0 || >=22.12.0`. Node 18 and Node 20 releases below 20.19 fall outside the supported range.

## Sources

- [Issues](https://github.com/shift-labs-ai/markit/issues)
- [License: MIT](https://github.com/shift-labs-ai/markit/blob/main/LICENSE)
- [README](https://github.com/shift-labs-ai/markit/blob/main/README.md)
- [Releases](https://github.com/shift-labs-ai/markit/releases)
- [shift-labs-ai/markit on GitHub](https://github.com/shift-labs-ai/markit)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/shift-labs-ai-markit
