marker
GitHub describes it as Convert PDF to markdown + JSON quickly with high accuracy. The repository metadata lists Python as its primary language. The metadata lists the Apache-2.0 license. This article stays within the project description and details documented in the GitHub repository README.
datalab-to/marker: Marker
GitHub describes it as Convert PDF to markdown + JSON quickly with high accuracy. The repository metadata lists Python as its primary language. The metadata lists the Apache-2.0 license. This article stays within the project description and details documented in the GitHub repository README.
Repository scope
GitHub describes it as Convert PDF to markdown + JSON quickly with high accuracy. The repository metadata lists Python as its primary language. The metadata lists the Apache-2.0 license. The README describes the project this way: - Converts PDF, image, PPTX, DOCX, XLSX, HTML, EPUB files in all languages - Formats tables, forms, equations, inline math, links, references, and code blocks - Extracts and saves images - Removes headers/footers/other artifacts - Extensible with your own formatting and logic - Optionally boost accuracy with LLMs (and your own prompt) - Works on GPU, CPU, or MPS
Try Datalab's Managed Platform
The README section "Try Datalab's Managed Platform" states: Our managed platform runs a version of our latest open source model, Chandra , higher accuracy than Marker, with zero data retention by default, SOC 2 Type 2, and custom BAAs.
Try Datalab's Managed Platform
The README section "Try Datalab's Managed Platform" states: If you have high volume workloads, we offer a batch processing service that has processed 1B+ pages per week , we manage the infrastructure so your workloads finish on time.
Performance
The README section "Performance" states: We measure marker on olmocr-bench, a third-party benchmark of 1,403 PDFs with tests covering math, tables, multi-column layout, scans, and hard edge cases. Balanced mode scores 76.0% overall , 83.5% on born-digital PDFs , ahead of MinerU and docling and within range of much larger VLMs, while fast mode runs the layout + text-layer path far cheaper (and a no-OCR mode goes faster still). Scores are the olmocr-bench overall (macro-average across the 8 categories).
Editorial conclusion
The repository README is the source for this review. It does not replace a local installation or an independent test.
Community notes