MinerU: turning PDFs, Office files and scans into LLM-ready Markdown
MinerU converts PDFs, Office documents, and images into LLM-ready Markdown or JSON, using a VLM+OCR dual engine that covers 109 languages, formulas, and complex layouts.
At a glance
- What is it?
- MinerU is an open source Python document parser from OpenDataLab that converts PDF, DOCX, PPTX, XLSX, images and web pages into Markdown or JSON. Its three inference backends trade speed against accuracy, and the licensing is its own, not Apache or MIT.
- Who is it for?
- Adopt MinerU if you need offline document conversion with formulas as LaTeX, tables as HTML and reading-order output, and you can accept the pipeline backend on CPU or a VLM backend on GPU. Do not adopt it if you need a permissive OSI license, or if your documents are plain text PDFs that pdftext or pypdf already handle.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Yes. The repository received new commits within the last day.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
DEEP OPEN-SOURCE ANALYSIS
What MinerU converts, and who ends up using it
MinerU's job is the step between a document and a language model. The README describes it as a high-accuracy document parsing engine for LLM, RAG and Agent workflows, and the package description in pyproject.toml is narrower: a practical document parsing tool for converting PDF, images, DOCX, PPTX, and XLSX into Markdown and JSON. The output is meant to be read by something downstream, not by a person browsing a folder.
The people who feel this problem first are RAG builders. A retrieval pipeline that chunks a PDF by raw text extraction gets a stream where a two-column layout interleaves, headers repeat on every page, and a table arrives as a column of disconnected numbers. MinerU targets that: the README states that output follows human reading order with automatic header and footer removal, formulas become LaTeX, and tables become HTML, including cross-page table merging. It also claims support for scanned documents, handwriting and multi-column layouts, with OCR across 109 languages.
The second audience is agent and tooling developers. The README lists an MCP Server for Cursor, Claude Desktop and Windsurf, plus native integration paths for LangChain, LlamaIndex, RAGFlow, RAG-Anything, Flowise, Dify and FastGPT. That is a different use case from batch conversion: here the document parse happens inside a conversation or an agent loop, which puts pressure on latency rather than throughput.
A third group is anyone with private documents. The README's deployment table is explicitly labelled private and fully offline, and lists domestic AI chips including Ascend, Cambricon, Enflame, MetaX, Moore Threads, Kunlunxin, Iluvatar, Hygon, Biren and T-Head. If your documents cannot leave your network, that list is the reason to look at this project rather than a hosted parser.
Three inference backends and what each one costs you
MinerU does not have one parsing path. The README's deployment table names three inference backends, and the choice between them is the main decision a new user makes.
The pipeline backend is described as fast and stable, with no hallucination, running on CPU or GPU. That last phrase matters. A vision-language model can invent content that is not in the document; a classical OCR and layout pipeline cannot, because it only transcribes what it detects. If you are parsing contracts or filings where a fabricated sentence is worse than a missed one, the pipeline backend is the conservative choice.
The vlm-engine backend is described as higher accuracy and supports the vLLM, LMDeploy and mlx ecosystems. That means a GPU and a serving stack, not just a pip install. The hybrid-engine backend is described as high accuracy with native text extraction and low hallucination, which reads as an attempt to get the accuracy of a VLM while leaning on the PDF's own text layer where one exists.
The changelog adds a control that the README table does not show. Version 3.3 introduced an effort parameter for the Hybrid backend with two levels, medium and high. The release notes state that on OmniDocBench v1.6, medium loses only 0.13 points of overall accuracy against high while running 35% to 220% faster depending on device and scenario, and that medium is now the default. The same notes say medium does not support image analysis, and that you need effort=high for image analysis or maximum accuracy. So the default configuration quietly excludes a capability. If your pipeline depends on describing figures, check this before you benchmark anything.
Installing MinerU and parsing your first PDF
The distribution name on PyPI is mineru, matching the project name in pyproject.toml. Python support is declared as >=3.10,<3.14, so a 3.14 interpreter will not resolve. The base dependency set is large and includes pypdfium2, pdftext, opencv-python, python-docx, openpyxl, fastapi and uvicorn, which tells you the install pulls in both parsing and serving machinery. The vlm extra exists as a separate optional dependency group and pulls torch, transformers and related packages, so the VLM backend is not installed by default.
Install the base package from PyPI:
pip install -U mineruThe repository ships demo/pdfs/ and demo/office_docs/ with sample inputs and demo/demo.py as a Python entry point, and the README lists a CLI among the development interfaces. The README does not print a CLI invocation in the text available here, so check the CLI's own help output before scripting a conversion.
The documentation states that the first run downloads models, and the 3.4 changelog describes automatic model source selection for first-time installations plus a local cache check before downloading. If your machine has no outbound access, that first run is where it fails, and the model source documentation linked from the changelog is the page to read before you try.
For programmatic use, the README points to Python, Go and TypeScript SDKs as well as a REST API. The REST API is served by FastAPI and uvicorn, both of which appear in the dependency list, which is consistent with a local service you start rather than a remote endpoint. The README does not document the API's routes, request schema or port in the text available here, so treat the SDK as the documented path and the HTTP surface as something to inspect in the source.
The licence is not an OSI licence, and that decides some evaluations
This is the part that ends adoption conversations. The pyproject.toml declares license = "LicenseRef-MinerU-Open-Source-License" and points license-files at LICENSE.md. That is a custom licence reference, not Apache-2.0, MIT or BSD. The repository also carries a MinerU_CLA.md, a contributor licence agreement, which is typical of projects that want to keep relicensing options open.
The changelog for version 3.1.0 is titled around licensing openness, and the visible text begins "License upgrade" before the excerpt cuts off. So the licence changed at some point in the 3.x line, and the current terms live in LICENSE.md rather than in the package metadata. Nothing here states the actual grant, the conditions, or whether commercial use is restricted.
That is a real constraint, not a formality. If your organisation runs a licence scanner, LicenseRef-MinerU-Open-Source-License will come back as an unknown identifier and land in a manual review queue. If you are shipping a product that embeds a parser, the terms in LICENSE.md determine whether you can, and no summary in a README substitutes for reading it. I am not going to characterise the terms, because this page does not contain them. Read the file.
Where MinerU is the wrong tool
The first failure mode is the one the project itself advertises around. A VLM backend can produce plausible text that is not in the source document. The README's own framing of the pipeline backend as having no hallucination is an admission that the accuracy-first paths carry that risk. For a legal or medical corpus where a wrong sentence is worse than a missing one, the higher-accuracy backend may be the worse choice, and the effort=medium default excludes image analysis on top of that.
Second, this is a heavy dependency tree for a narrow job. A clean, single-column, text-layer PDF does not need OCR, layout reconstruction or a vision model. The dependency list already includes pypdf and pdftext, which are the tools you would reach for in that case. If your documents are boring, MinerU is an expensive way to extract text.
Third, model downloads make the first run a network operation. The 3.4 changelog describes cache checking and automatic source selection as improvements to this, which confirms it was a friction point. In an air-gapped environment you need the model source documentation and a pre-populated cache; the README does not walk through that setup.
Fourth, the 4.0 line is in alpha. The most recent release listed is v4.0.0a6, published the same day as a 3.4.5 release. Running an alpha parser in a production ingestion pipeline means accepting that output format and CLI behaviour may move. The stable line and the alpha line are both live, and picking the wrong one is an easy mistake.
MinerU compared with Docling, and what the difference actually is
The obvious alternative is Docling, and people search for that comparison directly. Both convert documents to Markdown for downstream models, and both are Python. The difference is where the intelligence sits.
MinerU's architecture is a set of selectable engines. You choose pipeline for deterministic OCR and layout analysis, vlm-engine for a vision-language model served through vLLM, LMDeploy or mlx, or hybrid-engine for a mix that leans on the PDF's native text layer. The effort parameter then tunes how much work the hybrid path does. That is a knob-heavy design, and it assumes you have opinions about accuracy versus latency versus hallucination risk.
Docling's approach is a single document-conversion pipeline with its own model set, and it is not covered by the sources for this article, so I will not describe its internals. What can be said is structural: MinerU exposes backend selection as a first-class deployment decision, with a table mapping each backend to a use case, plus domestic AI chip support and an MCP server. If your requirement is choosing between CPU-only determinism and GPU-accelerated accuracy, or running on Ascend or Cambricon hardware, MinerU's backend split is the feature you are buying. If you want one pipeline with fewer decisions, the extra configuration is overhead you will not use.
The other alternative is not a parser at all. pypdf and pdftext are already in MinerU's dependency list, and for text-layer PDFs without complex layout they are the cheaper answer. The honest test is to run one representative document through a plain extractor first and see whether the output is actually broken.
Maintenance, upgrades and what to verify before you commit
The repository is not archived. The last push was on 2026-08-14, which is recent, and that same date carries both v4.0.0a6 and the 3.4.5 release. Two lines are being maintained at once: a stable 3.x series and a 4.0 alpha series. The changelog shows a steady cadence through 2026, with 3.1.0 in April, 3.3 in June and 3.4 in June, each carrying model upgrades and pipeline changes rather than only bug fixes.
Upgrade cost is dominated by models, not by code. The 3.3 release upgraded the VLM to MinerU2.5-Pro-2605-1.2B and the 3.4 release upgraded the pipeline OCR model to PP-OCRv6, claiming about 11% better OCR accuracy on OmniDocBench v1.6 and about 100% faster OCR processing. Every model swap changes your output on the same input. If you have regression tests over parsed documents, they will fail on upgrade for reasons that are improvements. If you do not, you will not notice until downstream retrieval quality shifts.
The 3.4 changelog also records a breaking change in language selection: Japanese, Traditional Chinese, English and Latin options were removed from OCR language selection and routed to the ch OCR model. Any script passing those language values will need updating. That is the kind of change a pinned version protects you from.
On licensing, the only safe statement is procedural. The package metadata declares LicenseRef-MinerU-Open-Source-License with LICENSE.md as the licence file, and the repository includes MinerU_CLA.md. Check both files against your own distribution model before you build on this, and do not rely on the phrase open source in the project's own name as a description of the terms.
Editorial conclusion
Adopt MinerU if you need offline document conversion with formulas as LaTeX, tables as HTML and reading-order output, and you can accept the pipeline backend on CPU or a VLM backend on GPU. Do not adopt it if you need a permissive OSI license, or if your documents are plain text PDFs that pdftext or pypdf already handle. Before committing, read LICENSE.md and MinerU_CLA.md, check the model source documentation for offline deployment, and confirm the Python version your environment provides falls inside the >=3.10,<3.14 range declared in pyproject.toml.
Frequently asked questions
What does MinerU do?
It converts PDFs, images, DOCX, PPTX and XLSX files into Markdown or JSON intended for LLM, RAG and agent workflows. The README states that formulas become LaTeX, tables become HTML, and output follows human reading order with headers and footers removed.
How to install MinerU?
The PyPI distribution is named mineru, so the base install is pip install -U mineru. Python support is declared as >=3.10,<3.14 in pyproject.toml, and the VLM backend lives in a separate optional dependency group that pulls torch and transformers.
How to use MinerU in Python?
The README lists Python, Go and TypeScript SDKs alongside a CLI and a REST API, and the repository ships demo/demo.py as a Python entry point with sample inputs under demo/pdfs/ and demo/office_docs/. The README does not document the individual SDK functions.
How to use MinerU?
The README lists several interfaces: a CLI, Python, Go and TypeScript SDKs, a REST API, Docker, and a Gradio WebUI, plus an MCP Server for tools such as Cursor and Claude Desktop. It does not print a CLI invocation in the text available here.
What is MinerU?
MinerU is an open source document parsing project from OpenDataLab, written in Python, that turns PDFs, Office documents, images and web pages into structured Markdown or JSON. It offers pipeline, vlm-engine and hybrid-engine inference backends and can run fully offline.
Official sources
Where this project is recommended
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/opendatalab-mineru)
Community notes