Model or dataset
Yuliang-Liu/MonkeyOCR avatar
Yuliang-Liu/MonkeyOCR

MonkeyOCR: a Structure-Recognition-Relation parser you install as magic_pdf

A lightweight LMM-based Document Parsing Model

6,652 stars462 forksPythonApache-2.0

At a glance

What is it?
MonkeyOCR is a Python document parsing model that splits page layout, text recognition and reading order into a three-stage triplet. The package installs as magic_pdf, and the same tree also ships an OpenAI-compatible server and a Gradio demo.
Who is it for?
Adopt MonkeyOCR if you need English and Chinese page parsing on a single GPU and can accept the pinned transformers==4.51.0 and PyMuPDF<=1.24.14 constraints. Do not adopt it if you need a release history to track, since the repository lists no releases, or if you need a documented rollback path.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 73 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What MonkeyOCR parses, and who the triplet paradigm is for

MonkeyOCR targets document parsing, the job of turning a PDF or image page into text plus structure. The README states the model handles English and Chinese documents, and the introduction frames the design as a Structure-Recognition-Relation triplet paradigm. The stated goal is to simplify what the README calls the multi-tool pipeline of modular approaches while avoiding the inefficiency of running a large multimodal model over a full page. That is the trade the project is built around: you get a pipeline of cooperating stages instead of one prompt over a whole page image.

The audience is engineers who already have a GPU and a corpus of pages to process, and who want the layout stage, the text stage and the ordering stage to be separable. If your pages are clean single-column English text, a plain OCR engine is cheaper. MonkeyOCR earns its place on pages where structure matters: multi-column layouts, tables, and mixed Chinese and English content. The README's own comparison table puts MonkeyOCR-pro-3B against Gemini 2.0-Flash, Gemini 2.5-Pro, Qwen2.5-VL-72B, GPT-4o and InternVL3-78B on OmniDocBench, and claims the best overall performance on both languages there. Those are the project's numbers, not an independent measurement.

How the SRR triplet splits a page: layout, recognition, relation

The paradigm name maps onto three stages. Structure covers where things are on the page. Recognition covers what the text says. Relation covers how the pieces read in order. The repository layout reflects that split: model_configs.yaml sits at the top level, and the parsing code lives under magic_pdf/, alongside parse.py as the entry script. A separate tools/ directory and a docs/ directory are present, and api/ holds the server side.

The package name is the detail most people miss. setup.py declares name="magic_pdf", so pip installs the distribution under that name even though the project is called MonkeyOCR. The same setup.py reads requirements.txt line by line, and for any line containing a URL it keeps only the part before the @ sign. That means git-sourced dependencies are installed by package name rather than by their pinned URL, which is worth knowing when you try to reproduce an exact environment.

Dependencies pin more tightly than usual. requirements.txt fixes transformers==4.51.0, qwen_vl_utils==0.0.10, pdfminer.six==20231228, doclayout_yolo==0.0.2b1 and gradio==5.23.3, and constrains numpy to >=1.21.6,<2.0.0 and PyMuPDF to >=1.24.9,<=1.24.14. Those ceilings are the practical boundary of the project. A newer transformers or a numpy 2.x install elsewhere in the same environment is a conflict, not a warning.

On throughput, the README publishes pages-per-second tables rather than a single headline figure. For MonkeyOCR-pro-1.2B the table lists 1.194 pages per second on a 4090 at 50 pages and 1.434 at 1000 pages; MonkeyOCR-pro-3B lists 0.972 and 1.006 on the same GPU. The 3090 and A6000 columns are lower throughout. Treat these as the project's own measurements on its own PDFs.

Installing magic_pdf and parsing your first PDF

The README does not print a pip install line, so the traceable path is the repository itself. The setup.py declares python_requires=">=3.9" and reads requirements.txt for install_requires. Clone the repository and install from the checkout so the pinned versions in requirements.txt are the ones that land in your environment.

bash
git clone https://github.com/Yuliang-Liu/MonkeyOCR.git
cd MonkeyOCR
pip install -r requirements.txt
pip install -e .

The last command installs the local tree under the distribution name magic_pdf, which is what setup.py sets. After that, the repository ships two ready entry points. parse.py is the script form of the parser, and demo/demo_gradio.py is a Gradio interface. The demo directory also carries demo1.pdf and demo2.pdf, so you can run the pipeline before pointing it at your own files. The README does not print a launch command for the demo, so run the script directly and read the console output for the address it serves on.

Gradio serves a local web UI by default; the README does not state the port, so read the console output rather than assuming one. If you prefer the server form, api/ contains a FastAPI application, and requirements.txt lists fastapi>=0.104.1 and uvicorn[standard]>=0.24.0 alongside openai==2.6.1, which is the client library used for the OpenAI-compatible surface. The model weights are not in the repository. The README links them on Hugging Face under echo840/MonkeyOCR-pro-3B and echo840/MonkeyOCR-pro-1.2B, and on ModelScope, so plan the download separately from the code install.

For a first real run, use the bundled demo PDFs and compare the output against the pages you can see. That tells you whether the layout stage is finding your columns before you spend GPU time on a full corpus.

Where MonkeyOCR is the wrong tool

The dependency pins are the first failure mode. transformers==4.51.0 is exact, not a floor, and numpy is held below 2.0.0. If your application already runs a newer transformers for another model, you are choosing between environments. There is no documented path in the README for running MonkeyOCR beside a second VLM stack in one interpreter.

The second limit is hardware. Every speed figure in the README is measured on a GPU: 3090, A6000, H800, 4090. There is no CPU throughput table, and no statement that CPU inference is supported. If your deployment target has no GPU, the published numbers do not describe your case at all.

The third is scope. The README describes English and Chinese documents for this repository. The 17-language claim belongs to MonkeyOCRv2, which the README announces as a separate repository with a 0.7B parser and a document-native vision backbone. If your corpus is multilingual beyond English and Chinese, this repository's own news section points you elsewhere.

The fourth is release tracking. The repository lists no releases, so there is no versioned artifact to pin against and no changelog to read. You pin a commit hash or you take main as it moves. The README also does not document rollback. If a model update changes your output, the recovery path is your own environment management, not anything the project describes.

MonkeyOCRv2, and what a modular pipeline does differently

The nearest alternative is the project's own successor. The README's news section announces MonkeyOCRv2 as a document-native vision backbone plus a 0.7B parser, released under Apache-2.0, and links to Yuliang-Liu/MonkeyOCRv2 with a separate paper. The difference in approach is architectural: MonkeyOCRv2 replaces the triplet split with a document-native backbone, and the README attributes the 17-language coverage to it. If you are starting today, comparing the two repositories before committing is the cheaper move, because the migration cost is paid at the start rather than after your pipeline is built.

Against a modular pipeline of separate tools, the difference is the other direction. A classic stack runs a layout detector, then a text recognizer, then a reading-order heuristic as independent components you can swap. MonkeyOCR keeps those three roles but couples them through one model family and one config file, model_configs.yaml. You lose the freedom to replace just the layout detector with a different project. You gain a single dependency set and a single set of published accuracy numbers instead of three sets that were never measured together.

Against a general-purpose VLM prompted over a full page, the README's own framing is efficiency: the introduction says the triplet avoids what it calls the inefficiency of using large multimodal models for full-page document processing. The pages-per-second tables are the evidence offered for that claim. They are the project's measurements on its own PDF set, so treat them as an upper bound until you reproduce them on your own documents.

Licence, maintenance and the cost of upgrading

MonkeyOCR is Apache-2.0, and LICENSE.txt sits at the top of the repository. The README states that MonkeyOCRv2 is also released under Apache-2.0. Apache-2.0 is a permissive licence with an explicit patent grant and a requirement to preserve notices; it is not a statement about the model weights, which are hosted separately on Hugging Face and ModelScope and may carry their own terms. Check the model card for the weights you download rather than assuming the repository licence covers them. This is a description of what the files say, not legal advice.

The last push to the repository was on 2026-07-20. That is recent enough that the code is not stale, but the repository has no releases, so there is no version number to reason about. Upgrading means moving to a newer commit and re-reading requirements.txt, because the pins are where breakage will show up first. The exact transformers==4.51.0 pin means a routine dependency refresh in your own project can silently break MonkeyOCR, and the reverse is also true.

The README's news list shows the project moving quickly: MonkeyOCR in June 2025, MonkeyOCR-pro-1.2B in July 2025, a v1.5 technical report in November 2025, the MonkeyDoc dataset in January 2026, dots.mocr in March 2026, acceptance by SCIENCE CHINA Information Sciences in July 2026, and MonkeyOCRv2 in July 2026. That cadence is a cost. Each release is a new set of weights and a new accuracy table, and nothing in the repository describes how to keep two versions side by side.

Editorial conclusion

Adopt MonkeyOCR if you need English and Chinese page parsing on a single GPU and can accept the pinned transformers==4.51.0 and PyMuPDF<=1.24.14 constraints. Do not adopt it if you need a release history to track, since the repository lists no releases, or if you need a documented rollback path. Verify first that your GPU throughput matches the pages-per-second table in the README and that MonkeyOCRv2, which the README points to, is not the better starting point for your workload.

Frequently asked questions

What is MonkeyOCR used for?

It parses documents, turning pages into text plus structure. The README states it handles English and Chinese documents using a Structure-Recognition-Relation triplet paradigm, with the layout, recognition and reading-order stages treated as separate roles.

Is MonkeyOCR open source?

Yes. The repository is licensed Apache-2.0, with LICENSE.txt at the top level, and the README states that MonkeyOCRv2 is also released under Apache-2.0. The model weights are hosted separately on Hugging Face and ModelScope, so check their terms as well.

Is MonkeyOCR considered AI?

The README describes MonkeyOCR as an LMM-based document parsing model, built on a large multimodal model rather than on a rule-based OCR engine, and it publishes comparisons against VLM baselines such as Gemini 2.0-Flash and GPT-4o.

Which AI OCR is best?

The README claims MonkeyOCR-pro-3B achieves the best overall performance on OmniDocBench for both English and Chinese, outperforming Gemini 2.0-Flash, Gemini 2.5-Pro, Qwen2.5-VL-72B, GPT-4o and InternVL3-78B. Those are the project's own results, not an independent comparison.

Official sources

  1. Issues
  2. License: Apache-2.0
  3. README
  4. Yuliang-Liu/MonkeyOCR on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/yuliang-liu-monkeyocr.svg)](https://hysenlabs.com/projects/yuliang-liu-monkeyocr)