docling vs MinerU: one unified library against a VLM and OCR parsing engine
Docling is a Python library and CLI that parses many formats into a single document representation for the gen AI stack. MinerU is a document parsing engine that turns PDFs, Office files and images into Markdown or JSON, with pipeline, vlm and hybrid backends. They overlap on PDF parsing, but they are built for different jobs and can be combined.
At a glance
| Project | docling-project/docling | opendatalab/MinerU |
|---|---|---|
| Licence | MITPermissive: commercial use allowed | Custom licenceCustom licence: read the LICENSE file |
| Maintenance | Commits in the last dayLast push September 29, 2026 | Commits in the last dayLast push September 29, 2026 |
| Language | Python | Python |
| GitHub stars | 68,180 | 80,836 |
| Read more | Our analysisGitHub | Our analysisGitHub |
Which one to choose
Choose docling if you need one Python representation for PDF, DOCX, PPTX, XLSX, HTML, EPUB, email, audio and video, with exports to Markdown, HTML, DocTags or lossless JSON, and native integrations with LangChain, LlamaIndex, Haystack or an MCP-connected agent.
Choose MinerU if your main problem is PDF and Office parsing with formulas as LaTeX, tables as HTML and reading-order output, and you want a choice between a CPU pipeline backend and GPU VLM or hybrid backends, or you need the SDKs, REST API and Docker deployment it documents.
Different centres of gravity: a unified document model versus a parsing engine
Docling's README describes a library that parses many formats and produces a unified DoclingDocument representation, then exports Markdown, HTML, WebVTT, DocLang, DocTags or lossless JSON. The supported list is wide: PDF, DOCX, PPTX, XLSX, HTML, EPUB, Apple Pages, WAV, MP3, WebVTT, Box Notes, email formats, images, LaTeX, DocLang and plain text. The README also lists advanced PDF understanding, page layout, reading order, table structure, code, formulas and image classification. The design goal is one representation that feeds the rest of a pipeline, and the README names integrations with LangChain, LlamaIndex, Crew AI and Haystack, plus an MCP server and a docling-serve API server. MinerU's README describes a parsing engine for LLM, RAG and Agent workflows that converts PDF, DOCX, PPTX, XLSX, images and web pages into structured Markdown or JSON. It highlights formulas to LaTeX, tables to HTML, layout reconstruction, scanned docs, handwriting, multi-column layouts, cross-page table merging, reading order and automatic header and footer removal. The centre of gravity is the parse itself, and the README lists MCP Server, LangChain, Dify and FastGPT integration, Python, Go and TypeScript SDKs, a CLI, a REST API and Docker. In short, docling is a library with a document model at the centre; MinerU is an engine with parsing quality and backend choice at the centre.
Architecture: one pipeline against three inference backends
Docling's README presents a single conversion path with options. The CLI command docling <url> generates a Markdown file, and the Python example constructs a DocumentConverter and calls convert. A VLM pipeline is available through --pipeline vlm --vlm-model granite_docling, and the README says several Visual Language Models are supported, including GraniteDocling. Audio is handled with ASR models, and OCR support covers scanned PDFs and images. The architecture is a library with pluggable models, not a set of named engines. MinerU's README names three inference backends: pipeline, described as fast and stable, no hallucination, running on CPU or GPU; vlm-engine, described as high accuracy and supporting the vLLM, LMDeploy and mlx ecosystems; and hybrid-engine, described as high accuracy with native text extraction and low hallucination. The changelog adds an effort parameter for the hybrid backend with medium and high levels. On OmniDocBench v1.6, the changelog states medium reduces overall accuracy by 0.13 points compared with high while giving 35% to 220% parsing speed improvements across devices and scenarios. The default hybrid backend uses effort=medium, and the changelog notes that medium does not support image analysis. That is a concrete architectural trade: docling exposes one pipeline with model choices, while MinerU exposes named backends with their own accuracy, speed and hardware assumptions.
Getting each one running
Docling installs with pip install docling. The README notes Python 3.9 support was dropped in version 2.70.0, so Python 3.10 or higher is required, and it says the project works on macOS, Linux and Windows for x86_64 and arm64. The quickstart is a CLI call or a short Python snippet, and the README points to detailed installation instructions for more control. MinerU's README lists Python, Go and TypeScript SDKs, a CLI, a REST API and Docker, and a Gradio WebUI and desktop client for no-code use. The earlier analysis notes pyproject.toml declares Python >=3.10,<3.14, so the interpreter range is bounded on both ends. The README also documents model source configuration, automatic source selection and local model usage, and the 3.4 changelog says first-time installs can choose a better model source based on the network environment and reuse locally cached model files. For a first run, docling is the shorter path: one pip install and a CLI command. MinerU asks you to pick a backend and understand model downloads, which is more setup but also more explicit control over where weights come from.
Operations, scaling and hardware
Docling runs locally, which the README frames as a fit for sensitive data and air-gapped environments. It can also run as a service through docling-serve, and the MCP server connects agents. The earlier analysis warns that the layout models need enough memory, and that you should check which model weights your environment can legally redistribute if you ship a container. There is no documented queue, worker pool or batching layer in the README; scaling is something you build around the library or the API server. MinerU's README is more explicit about deployment. The pipeline backend runs on CPU or GPU and is described as fast and stable with no hallucination. The vlm-engine targets GPU and supports vLLM, LMDeploy and mlx. The hybrid-engine combines native text extraction with high accuracy and low hallucination. The README also lists support for ten or more domestic AI chips, including Ascend, Cambricon, Enflame, MetaX, Moore Threads, Kunlunxin, Iluvatar, Hygon, Biren and T-Head. The 3.4 changelog reports OCR processing speed up about 100% and an OCR model upgrade to PP-OCRv6 with about 11% higher accuracy on OmniDocBench v1.6. Those are vendor-reported numbers, not independent measurements, but they show where MinerU invests: throughput and OCR quality on the pipeline backend, and backend-specific performance on the hybrid path.
Where each one falls short
Docling's README does not document a queue, worker pool, autoscaling or rollback procedure, so operational behaviour beyond running the library or the API server is unspecified. The earlier analysis says to skip it if you only need plain text extraction from clean, text-layer PDFs, because a smaller extractor costs less to install and run, and to skip it if you cannot give the process enough memory for the layout models. The README also lists metadata extraction and complex chemistry understanding as coming soon, so those are not available today. MinerU's licence is the first limitation. GitHub classifies it as Other, and the earlier analysis says it is not Apache or MIT, so a permissive OSI licence is not on offer. The README does not document rollback either, and the earlier analysis says to read LICENSE.md and MinerU_CLA.md before committing. The Python range is bounded to >=3.10,<3.14 in pyproject.toml, which can rule out newer interpreters. The hybrid backend's default effort=medium does not support image analysis, so maximum accuracy or image analysis requires switching to high. Both projects are silent on independent accuracy comparisons against each other; the numbers in MinerU's changelog come from its own release notes.
Licence and maintenance implications
Docling is MIT licensed, which the README links directly. That is a permissive licence and it matters if you plan to redistribute a container or embed the library in a product. The earlier analysis still flags one step: check which model weights your environment can legally redistribute, because the library licence does not automatically cover every model you download. MinerU's licence is classified by GitHub as Other, and the earlier analysis says it is its own licence, not Apache or MIT. The README does not summarise the terms; the earlier analysis points to LICENSE.md and MinerU_CLA.md. For a commercial deployment, that is a review task, not a checkbox. On maintenance, both repositories are unarchived and both were last pushed on 2026-09-17, so neither is stale by the six-month test. Docling's recent releases include v2.123.1 on 2026-08-28, v2.123.0 on 2026-08-26 and v2.122.0 on 2026-08-25, which shows a steady release cadence. MinerU's recent releases include v4.0.0a6 and mineru-3.4.5-released on 2026-08-14, and v4.0.0a5 on 2026-07-30, with an alpha line running alongside stable releases. Star counts are not evidence of quality, so the decision should rest on licence fit, release cadence and whether the documented features match your corpus.
Choosing for concrete scenarios
If you are building a RAG or agent pipeline that ingests many formats and wants one object to pass around, docling fits: the DoclingDocument representation and the LangChain, LlamaIndex, Haystack and MCP integrations are the point. If your corpus is mostly PDFs and Office files and you care about formulas as LaTeX, tables as HTML, reading order and header and footer removal, MinerU is the more direct match, and its backend choice lets you start on CPU with the pipeline backend and move to a GPU VLM or hybrid backend later. If you need a permissive OSI licence for redistribution, docling's MIT licence is the safer default, and MinerU requires reading its own licence first. If you need offline deployment on non-NVIDIA accelerators, MinerU's README lists support for many domestic AI chips, while docling's README does not document that. If your documents are clean, text-layer PDFs, both are heavier than a small extractor, and the earlier analysis says to skip docling in that case. The two are not mutually exclusive: docling can own the document representation and integration layer while MinerU handles difficult PDFs, but that means running two model stacks and two sets of dependencies, so it only pays off when the parsing quality difference is measurable on your own corpus.
Bottom line
For most teams that want one library feeding a gen AI stack and need a permissive licence, docling is the default choice. For teams whose hard problem is PDF and Office parsing quality, with formulas, tables and reading order, and who can accept a custom licence and a bounded Python range, MinerU is the stronger fit. Verify first on your own worst documents: how docling's default pipeline handles your scans and where its CLI output lands, and for MinerU which backend meets your accuracy target, whether effort=high is needed for image analysis, and what LICENSE.md and MinerU_CLA.md require of your deployment.