# PaddleOCR: PDF and image to Markdown/JSON with PP-OCRv6 and PaddleOCR-VL

> PaddleOCR is an Apache-2.0 Python toolkit that turns PDFs and images into Markdown or JSON. It ships a 0.9B document VLM, a PP-OCRv6 family from 1.5M to 34.5M parameters, and a CLI installed from PyPI.

**PaddlePaddle/PaddleOCR** — Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.

- Repository: https://github.com/PaddlePaddle/PaddleOCR
- Website: https://www.paddleocr.com
- Stars: 90,415 · Forks: 11,428
- Language: Python
- License: Apache-2.0
- Published: 2026-08-08 · Updated: 2026-08-18 · Language: en
- Canonical page: https://hysenlabs.com/projects/paddlepaddle-paddleocr

## The gap PaddleOCR fills between a scanned page and an LLM prompt

Most OCR libraries stop at a list of text boxes. PaddleOCR's stated goal is one step further: converting PDF documents and images into structured, LLM-ready data in JSON or Markdown. That distinction matters if you are building retrieval or agent workflows, because a flat string of recognized characters is not something you can chunk by heading or table row.

The project splits the job across two model families. PP-StructureV3 handles structure-aware conversion and, per the README, provides fine-grained coordinates including table cell coordinates and text coordinates. The PaddleOCR-VL series is the vision-language route: a 0.9B model that outputs Markdown and JSON directly, with the README claiming 96.3% on OmniDocBench v1.6 and gains on ancient documents, rare characters, seals and charts.

Who it is for: Python developers wiring document ingestion into a RAG or agent stack. The README names Dify, RAGFlow, Pathway and Cherry Studio as integrations, and the repository ships a langchain-paddleocr/ directory plus an mcp_server/ directory, which tells you the intended audience is people building pipelines, not people doing one-off scans in a GUI.

## Two engines, one package: PP-OCRv6 versus PaddleOCR-VL

The choice between the two engines is the first real decision you make, and the README does not spell out a rule for it.

PP-OCRv6 is the classic detection-plus-recognition path. It comes in three tiers measured in parameters: tiny at 1.5M, small at 7.7M, medium at 34.5M. The medium tier is claimed to beat PP-OCRv5_server by +4.6% detection and +5.1% recognition, and the release notes state it surpasses Qwen3-VL-235B and GPT-5.5 despite having 34.5M parameters. A single model covers Chinese, English, Japanese and 46 Latin-script languages, which means multilingual documents do not force a model switch mid-batch. The README also lists a 5.2x CPU speedup with OpenVINO, 6.1x on Apple M4 for the tiny tier, and 0.13s on an A100 GPU.

PaddleOCR-VL is a different shape of system. It is a 0.9B vision-language model that emits Markdown or JSON, and version 1.6 keeps the 1.5 architecture, so the release notes describe migration as a swap. If your output is a rendered document rather than a list of boxes, this is the path. If your output is coordinates you will post-process yourself, PP-StructureV3 is the one that exposes cell-level geometry.

One caveat worth flagging: the README's comparison numbers come from the project's own benchmarks against named models. Treat them as a starting hypothesis for your own documents, not a settled result.

## Installing PaddleOCR and running a first recognition pass

The package is on PyPI as paddleocr. The build metadata requires Python 3.8 or newer and depends on paddlex[ocr-core]>=3.7.0,<3.8.0, so a pip install pulls the PaddleX OCR core alongside it. The README does not reproduce an install transcript, so the only command the repository files give you is the package name itself.

```bash
pip install paddleocr
```

The pyproject.toml file declares a console script under [project.scripts], so the package installs a command-line entry point as well as the importable module. The repository files here do not show the full argument list for that command, and the README points to the documentation site for usage rather than listing flags inline.

For a first real use, the README's own framing is document parsing into Markdown or JSON, and the PaddleOCR-VL line is the one that produces those formats. The repository files available here do not include a runnable inference snippet, so the call to make is the one the documentation site describes for the PaddleOCR-VL pipeline. What you should see is structured output per input page in Markdown or JSON. Inspect the returned object before you write parsing code against it, because the README does not print a sample response body. For PDFs and multi-page documents, the PaddleOCR-VL and PP-StructureV3 pipelines are the documented route rather than a single-image predictor.

## Where PaddleOCR is the wrong tool

The dependency chain is the first hard limit. pyproject.toml pins paddlex to the 3.7.x range. If another package in your environment requires a different PaddleX line, pip will either refuse to resolve or you will be forced into a virtual environment just for OCR. That is a real cost in a monorepo with a shared lockfile.

Second, the project is a moving target. Three releases landed between 2026-04-21 and 2026-06-11: 3.5.0, 3.6.0 and 3.7.0. The 3.5.0 notes describe switching inference backends between Paddle static graph, Paddle dynamic graph and Transformers, plus DOCX export and a browser SDK. Each of those is a new surface. If your deployment freezes dependencies for a year at a time, you are adopting a library whose documented feature set changes roughly every six weeks.

Third, the README does not document a rollback procedure for model upgrades. The 3.6.0 notes say the PaddleOCR-VL-1.6 architecture is consistent with 1.5 and migration is a swap, which is reassuring for that pair, but there is no stated path for reverting a pipeline that regresses on your corpus after an upgrade. Plan to keep your own pinned model artifacts.

Finally, if your documents are clean, born-digital text with a simple layout, a lighter library will do. The structure-aware and VLM paths earn their weight on tables, formulas, seals and mixed layouts; on a plain English invoice you are paying for capability you will not exercise.

## PaddleOCR versus Tesseract: different assumptions about the page

The most common comparison people search for is PaddleOCR against Tesseract, and the difference is architectural rather than a matter of accuracy percentages.

Tesseract is a recognition engine with a long lineage and a line-based model of a page. PaddleOCR ships detection and recognition as separate trained stages, adds a structure-aware pipeline on top, and adds a vision-language model that produces Markdown or JSON. The output units differ: Tesseract gives you text, PaddleOCR's structure pipeline gives you text plus table cell coordinates and text coordinates, and PaddleOCR-VL gives you a document rendering.

The language story also diverges. PaddleOCR claims 100+ languages overall, with PP-OCRv6 covering 50 languages in one unified model covering Chinese, English, Japanese and 46 Latin-script languages. With Tesseract, language coverage is a matter of which traineddata files you install and which you select at runtime.

The honest trade-off: PaddleOCR's accuracy claims are self-reported against PP-OCRv5 and against named VLMs, and the README does not publish a head-to-head against Tesseract. If your workload is a single Latin-script language and a fixed layout, the smaller dependency footprint of a classic engine may win on operational simplicity regardless of which model scores higher.

## Licence, maintenance and what an upgrade actually costs

PaddleOCR is Apache-2.0, declared in both pyproject.toml and the LICENSE file at the repository root. For most commercial use that is a permissive arrangement, but the licence covers the code, not the model weights or any third-party component you bundle, and the README does not enumerate weight licensing. Check that separately before shipping a product; this is not legal advice.

On maintenance, the last push to the default branch was on 2026-06-11, the same date as the v3.7.0 release. The repository is not archived. Releases in the visible window run 2026-04-21, 2026-05-28 and 2026-06-11, so the cadence is roughly monthly, and the README's recent-updates list continues past that with a 2026-07-22 entry for HPD-Parsing.

The upgrade cost is not the pip install. It is the model artifacts. PP-OCRv6 ships in three tiers hosted on HuggingFace and ModelScope, and the PaddleOCR-VL line moves on its own schedule. A version bump can change which weights are downloaded and which output schema you receive. The pyproject pin on paddlex means the library and the OCR core move together, so a partial upgrade is not a supported configuration. Budget for a staging run over a representative document sample on every minor version, and keep the previous model directory so you can point the pipeline back at it.

## Conclusion

Adopt PaddleOCR if you need multilingual OCR or PDF-to-Markdown output inside a Python service and you are willing to pin paddlex to the 3.7.x line. Do not adopt it if you cannot install PaddlePaddle on the target machine, or if you need a documented rollback path for model upgrades, since the README does not describe one. Verify first that your Python version is 3.8 or newer, that your hardware backend is one of the supported ones, and whether your documents need PaddleOCR-VL or PP-OCRv6 by running both on a sample page before committing to a pipeline.

## FAQ

### What is PaddleOCR used for?

It converts PDF documents and images into structured, LLM-ready data in JSON or Markdown, and also performs general text recognition across 100+ languages. The README positions it as the ingestion layer for RAG and agent applications.

### Is PaddleOCR better than Tesseract?

The README does not publish a head-to-head comparison with Tesseract. It reports PP-OCRv6 gains of +4.6% detection and +5.1% recognition over PP-OCRv5 and states that the medium tier surpasses Qwen3-VL-235B and GPT-5.5, so any Tesseract comparison has to come from your own documents.

### Is PaddleOCR free?

The code is released under Apache-2.0, as declared in pyproject.toml and the LICENSE file. The licence covers the source code; the README does not state terms for the model weights.

### How do I install PaddleOCR in Python?

Install the paddleocr package from PyPI with pip. The build metadata requires Python 3.8 or newer and pulls paddlex[ocr-core] in the 3.7.x range as a dependency.

### How do I use PaddleOCR locally?

Install the package, then use the paddleocr command-line entry point declared in pyproject.toml, or the PaddleOCR-VL pipeline the documentation site describes. The README points to the documentation for pipeline-level setup rather than reproducing it.

### How do I use PaddleOCR with a GPU?

The README states that one-click deployment supports hardware backends including NVIDIA GPU, Intel CPU, Kunlunxin XPU and diverse AI accelerators, and cites 0.13s inference on an A100 GPU. It does not list GPU-specific install flags in the repository files, so the setup steps live in the linked documentation.

## Sources

- [Official documentation](https://www.paddleocr.com)
- [Official README](https://github.com/PaddlePaddle/PaddleOCR#readme)
- [Project repository](https://github.com/PaddlePaddle/PaddleOCR)
- [Release notes](https://github.com/PaddlePaddle/PaddleOCR/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/paddlepaddle-paddleocr
