Hysen Labs
Open-source project
datalab-to/marker avatar
datalab-to

marker

GitHub describes it as Convert PDF to markdown + JSON quickly with high accuracy. The repository metadata lists Python as its primary language. The metadata lists the Apache-2.0 license. This article stays within the project description and details documented in the GitHub repository README.

38,712 stars2,764 forksPythonApache-2.0
01
DEEP OPEN-SOURCE ANALYSIS

datalab-to/marker: Marker

GitHub describes it as Convert PDF to markdown + JSON quickly with high accuracy. The repository metadata lists Python as its primary language. The metadata lists the Apache-2.0 license. This article stays within the project description and details documented in the GitHub repository README.

02
DEEP OPEN-SOURCE ANALYSIS

Repository scope

GitHub describes it as Convert PDF to markdown + JSON quickly with high accuracy. The repository metadata lists Python as its primary language. The metadata lists the Apache-2.0 license. The README describes the project this way: - Converts PDF, image, PPTX, DOCX, XLSX, HTML, EPUB files in all languages - Formats tables, forms, equations, inline math, links, references, and code blocks - Extracts and saves images - Removes headers/footers/other artifacts - Extensible with your own formatting and logic - Optionally boost accuracy with LLMs (and your own prompt) - Works on GPU, CPU, or MPS

03
DEEP OPEN-SOURCE ANALYSIS

Try Datalab's Managed Platform

The README section "Try Datalab's Managed Platform" states: Our managed platform runs a version of our latest open source model, Chandra , higher accuracy than Marker, with zero data retention by default, SOC 2 Type 2, and custom BAAs.

04
DEEP OPEN-SOURCE ANALYSIS

Try Datalab's Managed Platform

The README section "Try Datalab's Managed Platform" states: If you have high volume workloads, we offer a batch processing service that has processed 1B+ pages per week , we manage the infrastructure so your workloads finish on time.

05
DEEP OPEN-SOURCE ANALYSIS

Performance

The README section "Performance" states: We measure marker on olmocr-bench, a third-party benchmark of 1,403 PDFs with tests covering math, tables, multi-column layout, scans, and hard edge cases. Balanced mode scores 76.0% overall , 83.5% on born-digital PDFs , ahead of MinerU and docling and within range of much larger VLMs, while fast mode runs the layout + text-layer path far cheaper (and a no-OCR mode goes faster still). Scores are the olmocr-bench overall (macro-average across the 8 categories).

06
DEEP OPEN-SOURCE ANALYSIS

Editorial conclusion

The repository README is the source for this review. It does not replace a local installation or an independent test.

07
DEEP OPEN-SOURCE ANALYSIS

Official sources

08
Community notes

Community notes