Open-source project
frotms/PaddleOCR2Pytorch avatar
frotms/PaddleOCR2Pytorch

PaddleOCR2Pytorch: running PaddleOCR weights without installing PaddlePaddle

PaddleOCR inference in PyTorch. Converted from [PaddleOCR](https://github.com/PaddlePaddle/PaddleOCR)

1,210 stars224 forksPythonApache-2.0

At a glance

What is it?
A community port that reimplements PaddleOCR's dynamic graph models in PyTorch, including the PP-OCR series and a full PP-StructureV3 document parsing pipeline, with weights converted rather than retrained.
Who is it for?
This port exists for one concrete reason: to load a model PaddleOCR trained without installing PaddlePaddle, which matters when your deployment target already has a PyTorch runtime and adding a second framework is not an option.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 88 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 20, 2026, and from our analysis. They are not legal advice.

Editorial analysis

A conversion project first, an OCR library second

The stated aim of frotms/PaddleOCR2Pytorch is threefold, and the order matters. It is a way to study how PaddleOCR is built, a way to use models trained by PaddleOCR inside PyTorch, and a reference for anyone else porting something from Paddle to PyTorch. The name in the code is `PytorchOCR`, and the README states that it was ported from the dynamic graph version of PaddleOCR v2.0 and later.

That last qualifier does real work. PaddleOCR has both a dynamic graph and an older static graph lineage, and the conversion work only covers the former. Anyone arriving expecting the older inference path will find the port stops at a version boundary.

The framing also sets expectations about quality. Nothing here is retrained: the value is in faithfully reproducing the forward pass so existing published weights produce the same output in a different framework. Where a Paddle-specific operator had no direct equivalent, the porting guides under `skills/` are the place to look, since those are written per model generation.

The repository is licensed Apache-2.0, is written in Python, is not archived, and the most recent push is dated 2026-07-10. It has 1,210 stars and 224 forks.

What the directory layout tells you about the work

The tree is small and named for function rather than for Python packaging conventions. `pytorchocr/` is the runtime library, the part you import. `converter/` holds the code that turns Paddle checkpoints into PyTorch state dicts, which is the piece that has to know the name mapping between the two frameworks.

`ptstructure/` is a parallel runtime for document parsing, matching the upstream PP-Structure package rather than the OCR pipeline. That separation matters if you only want text detection and recognition, since the document layout machinery is a separate import surface.

`skills/` is the unusual directory and probably the most valuable one to read first. It contains per-version porting guides, including one for PP-StructureV3 and another for PP-OCRv6, which document what had to be rewritten for that particular generation rather than describing usage. For anyone extending the port to a new PaddleOCR release, these guides are the actual specification.

The rest is conventional: `configs/` for model definitions, `tools/` for scripts, `doc/` for documentation including installation and inference pages, and `misc/`. There is also a top-level `__init__.py`, so the root is importable rather than being only a container.

Model coverage tracks PaddleOCR's own release history

The README keeps a dated log of what has been ported, running from January 2021 to June 2026, and reading it tells you how closely the project tracks upstream. The earliest entries are the foundation: DB text detection, SAST, EAST, ROSETTA and CRNN in April 2021, then multilingual models covering more than 27 languages later that year, and PP-OCRv2 in September 2021 with the claim of a 220 percent CPU inference improvement over PaddleOCR server.

From 2022 the log shifts to individual algorithms as they appear upstream: PSENet for detection, NRTR and then SAR and SVTR for recognition, FCENET, and DB++ for detection. ViTSTR arrives in October 2022. In 2023 the project picks up the CAN formula recognition model and Text Telescope for text super-resolution, both of which are post-processing steps rather than core OCR.

PP-OCRv3 in May 2022 and PP-OCRv4 in February 2024 follow, the latter with both mobile and server variants and the claim that Chinese detection accuracy rose 4.5 percent over v3 while English rose 10 percent. PP-OCRv5 landed in May 2025 as the first single model to handle Simplified Chinese, Traditional Chinese, Chinese pinyin, English and Japanese together, with a claimed 13 point recognition gain over v4 and specific attention to heavily cursive handwriting.

The named production model families in the README are the ultralight PP-OCR series covering detection, a direction classifier and recognition, plus a `ptocr_mobile` set for phones and a `ptocr_server` set for general use.

PP-OCRv6 in three sizes with published accuracy claims

The June 2026 entry covers PP-OCRv6, which is the newest detection and recognition work in the port. Three model sizes are described: Tiny at 1.5 million parameters, Small at 7.8 million, and Medium at 35 million, spanning device-side through server-side compute. The Tiny model supports 49 languages and the small and medium models 50.

The architecture names are worth carrying, because they tell you which Paddle modules had to be reimplemented: a PPLCNetV4 backbone, RepLKFPN and RepLKPAN detection necks, and an EncoderWithLightSVTR recognition neck. Structural reparameterisation and lightweight attention are exactly the kinds of features that need careful translation, so their presence is a fair signal of how much of the port is not mechanical.

The accuracy and speed figures quoted are all relative to PP-OCRv5 server: detection precision up 4.6 percent, recognition precision up 5.1 percent, and GPU inference 2.37 times faster. Those are upstream PaddleOCR numbers restated in this README, so they describe the model's design targets rather than a measurement of the port. The README's TODO list marks PP-OCRv6 conversion as finished, with a pointer to a dedicated porting guide.

PP-StructureV3 ported as a complete document pipeline

The largest single piece of work is the PP-StructureV3 port, dated June 2026 and marked complete. Its own guide lives at `skills/ppstructurev3_porting_guide.md`, and the README states the goal plainly: pure PyTorch inference with zero Paddle dependency, across multi-scenario and multi-layout PDF parsing.

Seven components are listed. Layout detection classifies 23 document region types including titles, body text, tables, images, formulas and seals. Table structure recognition uses the SLANeXt ViT plus GRU architecture and emits HTML for bordered and borderless tables. Formula recognition uses PP-FormulaNet, built on PPHGNetV2 with MBart, in small and medium variants, producing LaTeX. Seal detection combines DB text detection with OCR on the seal region, falling back to a global OCR pass. Document preprocessing covers image orientation classification, UVDoc dewarping and text-line orientation. Global OCR runs detection and recognition once over the whole image, aligned to the PaddleX architecture. Output formats are Markdown, JSON and visualisations.

Layout detection itself ships in three sizes, S at 1.2 million parameters, M at 5.8 million and L at 20 million. One inconsistency is worth noting: the TODO list still has PP-Structurev2 unchecked even though the v3 port is called complete, so the two generations are at different states and should not be assumed equivalent.

Weights come from a Netdisk link, and the TODO list is long

Practical deployment detail first: the converted PyTorch weights are not on a package registry or a model hub. The README gives a Baidu Netdisk link with an extraction code for the PyTorch models, and a second link with its own code for the original PaddleOCR weights. Multilingual and additional models are pointed at a documentation page in `doc/`. Two links with two extraction codes is a real friction point for CI pipelines, and it means the weights cannot be fetched programmatically from the repository itself.

The TODO list is the best indicator of scope remaining. Checked items are PP-OCRv6, PP-OCRv5 including the document orientation, unwarp and text-line orientation modules, and PP-StructureV3. Unchecked items include PP-ChatOCRv4, which upstream pairs with a large language model, the DRRG and RFL detection and recognition algorithms, ABINet, VisionLAN, SPIN and RobustScanner for recognition, TableMaster for tables, three DocVQA models based on LayoutLM variants, and a set of optimisation tasks for the older v2 structure package covering layout model size, table accuracy and key information extraction.

Two smaller signals worth noting. The issue count of 85 against 1,210 stars is high relative to the repository's popularity, which is consistent with a port where each upstream release opens new conversion questions. And the README is Chinese-first with `README_en.md` alongside it, so English readers get the same content but not the first draft of it.

Editorial conclusion

This port exists for one concrete reason: to load a model PaddleOCR trained without installing PaddlePaddle, which matters when your deployment target already has a PyTorch runtime and adding a second framework is not an option. The conversion work is substantial rather than a thin wrapper, with PP-OCRv6 detection and recognition models in three sizes and a PP-StructureV3 pipeline including layout detection, table structure, formula recognition and seal detection all running in pure PyTorch. What to weigh before adopting it: the model list depends on a Baidu Netdisk link rather than a registry, the README is Chinese-first with an English companion file, and the issue tracker is deep at 85 open entries against 1,210 stars. Read the porting guides under `skills/` for the versions already marked complete, and check the weight link before planning an offline deployment.

Frequently asked questions

Do I need PaddlePaddle installed to use PaddleOCR2Pytorch?

No. That is the point of the port. It reimplements PaddleOCR's dynamic graph models, from PaddleOCR v2.0 onward, in PyTorch so you can run published PaddleOCR weights inside a PyTorch runtime. The README describes the PP-StructureV3 pipeline as pure PyTorch inference with zero Paddle dependency.

Which PaddleOCR models have been converted to PyTorch?

The README logs conversions from April 2021 onward, including DB, SAST, EAST, ROSETTA and CRNN, PSENet and DB++ for detection, and NRTR, SAR, SVTR and ViTSTR for recognition, plus the CAN formula model and Text Telescope for super-resolution. Whole generations are covered too: PP-OCRv2 through PP-OCRv6, and the PP-StructureV3 document parsing pipeline.

Where do I download the converted model weights?

The README links to a Baidu Netdisk share with an extraction code for the PyTorch models, and a separate share for the original PaddleOCR weights. There is no package registry or model hub involved, so an automated pipeline will need to handle the download manually. A documentation page under doc/ covers multilingual and additional models.

How does this port compare with running PaddleOCR directly?

Running upstream PaddleOCR gives you the reference implementation and everything it has ported. This repository gives you the same model weights in a framework that avoids a second runtime dependency, and it is actively maintained against new PaddleOCR releases, with PP-OCRv6 added in June 2026. The trade is that accuracy and speed figures quoted in the README are inherited from upstream rather than measured for the port.

Official sources

  1. frotms/PaddleOCR2Pytorch on GitHub
  2. Issues
  3. License: Apache-2.0
  4. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/frotms-paddleocr2pytorch.svg)](https://hysenlabs.com/projects/frotms-paddleocr2pytorch)