Open-source project
datalab-to/marker avatar
datalab-to/marker

marker-pdf: PDF to Markdown Conversion with a Hybrid VLM Path

GitHub describes it as Convert PDF to markdown + JSON quickly with high accuracy. The repository metadata lists Python as its primary language. The metadata lists the Apache-2.0 license. This article stays within the project description and details documented in the GitHub repository README.

40,098 stars2,903 forksPythonApache-2.0

At a glance

What is it?
Datalab's marker-pdf turns PDFs and office documents into markdown, JSON, chunks and HTML, using a local surya inference server and an optional LLM pass. It is Apache-2.0 code with separately licensed model weights, and the choice between its fast and balanced modes is the main decision an adopter has to make.
Who is it for?
Adopt marker-pdf if you need local, scriptable document conversion and can run either an NVIDIA GPU with Docker or llama.cpp on CPU and Apple Silicon. Do not adopt it if you need a supported commercial licence for the model weights without meeting the funding and revenue threshold, or if you cannot run a local inference server at all.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 16 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

DEEP OPEN-SOURCE ANALYSIS

What marker-pdf is for, and who actually needs it

The problem marker-pdf addresses is the gap between a PDF and anything a program can read. A PDF stores glyphs at coordinates, not paragraphs, so extracting structure means deciding what is a heading, what is a table, what is a two-column layout, and what is a running header that should be dropped. marker-pdf's README states that it converts documents to markdown, JSON, chunks and HTML, and that it formats tables, forms, equations, inline math, links, references and code blocks, extracts and saves images, and removes headers, footers and other artifacts.

The intended user is someone building a pipeline rather than reading a single file. The repository ships both a batch entry point and a single-file one, plus a Streamlit app and a server, which tells you the project expects to be embedded. The README also lists PDF, image, PPTX, DOCX, XLSX, HTML and EPUB as inputs, in all languages. That breadth is the reason to look at it: most converters pick one of those and stop.

It is not aimed at someone who wants a hosted API. Datalab runs a managed platform separately, and the README points at it, but the open source project is the local path.

How the conversion pipeline is put together

The mechanism the README describes is a single surya vision language model handling layout, OCR and table recognition, served by a local inference server that marker spawns automatically on first use. On NVIDIA GPUs that server runs under vLLM inside Docker; elsewhere it runs under llama.cpp. You can also point marker at a server you already have running by setting the environment variable SURYA_INFERENCE_URL to a host and port ending in /v1.

The mode setting changes how much of that model gets used. In balanced mode, which the README says is the default on GPU, surya's VLM does layout, OCRs inline math, and re-OCRs the whole page whenever any embedded text is judged bad. In fast mode, the default on CPU and MPS, a lightweight rf-detr layout detector runs instead, text comes from pdftext, and VLM use is kept to equations, surgical block-level repair of individual garbled or empty blocks, and a single full-page pass only for pages that are scanned or mostly bad. The README states that a clean digital document without equations never starts the VLM server at all in fast mode.

Tables are reconstructed from the PDF text layer in both modes, with scanned tables coming from full-page OCR, and low-confidence reconstructions falling back to the VLM under a stricter bar in balanced mode. There is also a --disable_ocr flag that turns off all VLM calls in either mode, leaving pure text-layer extraction.

Installing marker-pdf and converting your first document

The README requires Python 3.10 or newer and PyTorch. The package name on PyPI is marker-pdf, not marker. The base install handles PDFs; the full extra pulls in the libraries for the other formats.

bash
pip install marker-pdf

If you plan to feed it DOCX, PPTX, XLSX, HTML or EPUB, install the extra as well. The README is explicit that documents other than PDFs need these additional dependencies.

bash
pip install marker-pdf[full]

Before the first conversion you need an inference backend. On an NVIDIA GPU that means Docker plus the NVIDIA Container Toolkit. On CPU or Apple Silicon you need the llama-server binary from llama.cpp, which on macOS comes from Homebrew.

bash
brew install llama.cpp

After that, the installed console scripts are marker for batch conversion and marker_single for one file. Running marker_single on a PDF is the shortest path to seeing output; the README's examples directory links sample markdown and JSON produced from a textbook and two arXiv papers, which is the fastest way to judge whether the output shape suits your downstream code before you convert anything of your own.

The LLM pass, and what it costs you in dependencies

Passing --use_llm turns on hybrid mode, where an LLM runs alongside marker. The README says this is what merges tables across pages, handles inline math, formats tables properly, and extracts values from forms. The default model is gemini-3.5-flash, and the supported providers are Gemini, Claude, OpenAI-compatible endpoints, Azure, Vertex, OpenRouter and Ollama.

That list is also a dependency list. Looking at pyproject.toml, the base install already carries google-genai, anthropic and openai as hard dependencies, so the LLM client libraries are present whether or not you enable the flag. If you are packaging marker-pdf into a constrained image, that is weight you pay for before you use the feature.

Ollama in that provider list is the interesting entry for anyone who cannot send pages to a third party. It means hybrid mode can run against a local model, though the README does not state which local models reach comparable quality to the default, and it does not document what happens when the LLM returns malformed output for a table. Treat hybrid mode as a quality dial with an unquantified failure surface rather than a mode you enable blindly.

Where marker-pdf is the wrong tool

The clearest boundary is the model weights. The README states that the code is Apache 2.0 and free to use including commercially, but that the model weights use a modified AI Pubs Open Rail-M license, free for research, personal use, and startups under $5M in funding or revenue, with commercial use beyond that requiring the paid pricing page. The repository carries a separate MODEL_LICENSE file alongside LICENSE, which is the file to read before shipping anything. A team above that threshold cannot treat this as an Apache-2.0 project end to end, and no amount of reading the code licence changes that.

The second boundary is operational. marker-pdf is not a library you call in a request handler and forget. It spawns or connects to an inference server, and on GPU that server needs Docker and the NVIDIA Container Toolkit. A serverless environment with no GPU, no Docker and no llama.cpp binary has nothing to run the model on. In that setting a hosted conversion API is the honest answer, and the README itself points to Datalab's managed platform, noting it runs Chandra with higher accuracy than Marker, zero data retention by default, SOC 2 Type 2 and custom BAAs.

The third is scope creep. If your documents are all clean, born-digital, single-column PDFs with no equations, the balanced mode machinery is mostly idle and --disable_ocr gets you text-layer extraction without any VLM calls. That is a legitimate configuration, but at that point you should check whether pdftext or a simpler extractor already covers your case, because you are paying the install cost of torch and surya for a path that does not use them.

marker-pdf against docling and MinerU

The README names two alternatives directly: it states that balanced mode scores 76.0% overall on olmocr-bench, 83.5% on born-digital PDFs, ahead of MinerU and docling and within range of much larger VLMs. olmocr-bench is a third-party benchmark from Allen AI of 1,403 PDFs with tests covering math, tables, multi-column layout, scans and hard edge cases, and the README says the scores are the macro-average across its eight categories.

The difference in approach matters more than the number. MinerU and docling are separate projects with their own pipelines, and the README does not describe their internals, so the concrete distinction it draws is about marker's own design: a single surya VLM serving layout, OCR and table recognition through one local server, with a fast path that avoids the VLM entirely for clean digital pages. The README explicitly invites you to run your own benchmarks and links per-category scores below its performance section.

That is the right posture for a comparison like this. Benchmark scores move with model versions, and a macro-average across eight categories can hide a category where your documents live. If your corpus is heavy on scanned forms, the overall figure tells you less than the per-category breakdown the README points to.

Licensing, upgrades and what to check before you commit

The split licence is the first thing to put in front of whoever signs off on dependencies. Apache-2.0 on the code is permissive and the README says so plainly; the modified AI Pubs Open Rail-M terms on the weights are a separate question with a funding and revenue threshold attached. This is not legal advice, and the threshold is defined by the licence text in MODEL_LICENSE rather than by the README's summary, so read the file.

On upgrades, the version history is worth noting. v2.0.0 landed on 2026-07-20, the same date as the last push to master, and the previous release v1.10.2 was on 2026-01-31. A major version bump after roughly six months of patch releases means the 1.x to 2.x step is the one to test carefully; the README does not document a migration path from 1.x, so budget for reading release notes and diffing output on a fixed set of documents rather than assuming drop-in compatibility.

Dependency pinning is the other cost. pyproject.toml pins torch to >=2.7.0,<3 and surya-ocr to >=0.22.1,<0.23.0, with a long list of bounded ranges around them. In an environment that already has its own torch version, that upper bound is where conflicts will surface first.

Editorial conclusion

Adopt marker-pdf if you need local, scriptable document conversion and can run either an NVIDIA GPU with Docker or llama.cpp on CPU and Apple Silicon. Do not adopt it if you need a supported commercial licence for the model weights without meeting the funding and revenue threshold, or if you cannot run a local inference server at all. Verify first that your Python is 3.10 or newer, that marker-pdf installs cleanly against your existing torch pin, and that your documents fall inside the formats you actually installed for: pip install marker-pdf covers PDFs, while other formats need the full extra.

Frequently asked questions

How do I install marker-pdf?

Install it with pip install marker-pdf after making sure you have Python 3.10 or newer and PyTorch. If you need to convert formats other than PDF, the README says to install the extra dependencies with pip install marker-pdf[full].

What is marker-pdf?

marker-pdf is the package name for Marker, a tool that converts documents to markdown, JSON, chunks and HTML. The README lists PDF, image, PPTX, DOCX, XLSX, HTML and EPUB as supported inputs in all languages.

What is the difference between balanced and fast mode in marker-pdf?

Balanced mode is the default on GPU and uses the surya VLM for layout, OCRs inline math, and re-OCRs a whole page when its embedded text is bad. Fast mode is the default on CPU and MPS, uses the rf-detr layout detector with pdftext for text, and keeps VLM use to equations, block-level repair and scanned or mostly bad pages.

Can I use marker-pdf commercially?

The README states the code is Apache 2.0 and free to use including commercially, but the model weights use a modified AI Pubs Open Rail-M license that is free for research, personal use, and startups under $5M in funding or revenue. Commercial use of the weights beyond that points to Datalab's pricing page.

Official sources

  1. Official documentation
  2. Official README
  3. Project repository
  4. Release notes
For maintainers

Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/datalab-to-marker.svg)](https://hysenlabs.com/projects/datalab-to-marker)
Community notes

Community notes