# OCRFlux: the benchmark table where the complex-table row disagrees with the headline

> OCRFlux converts PDFs and images to Markdown with a 3B parameter vision language model, and its page carries two results tables against three baselines. The single-page table is a clear win. The table benchmark is not, and the page does not draw that distinction.

**chatdoc-com/OCRFlux** — OCRFlux is a lightweight yet powerful multimodal toolkit that significantly advances PDF-to-Markdown conversion, excelling in complex layout handling, complicated table parsing and cross-page content merging.

- Repository: https://github.com/chatdoc-com/OCRFlux
- Stars: 2,533 · Forks: 157
- Language: Python
- License: Apache-2.0
- Published: 2026-09-28 · Updated: 2026-09-28 · Language: en
- Canonical page: https://hysenlabs.com/projects/chatdoc-com-ocrflux

## The headline EDS gain is arithmetic on the total row

The key features section states that OCRFlux beats three baselines on its released single-page benchmark, and gives the gaps as 0.095 higher, from 0.872 to 0.967; 0.109 higher, from 0.858 to 0.967; and 0.187 higher, from 0.780 to 0.967. Those are the Total row of the first results table, and the subtraction checks out against olmOCR-7B-0225-preview at 0.872, Nanonets-OCR-s at 0.858 and MonkeyOCR at 0.780, with OCRFlux-3B at 0.967.

The metric is Edit Distance Similarity against ground-truth Markdown, measured on OCRFlux-bench-single, which holds 2000 PDF pages, 1000 English and 1000 Chinese, with the Markdown labelled by hand over several rounds. The Total figure is the plain average of the two language rows, which is why OCRFlux lands at 0.967 between an English 0.971 and a Chinese 0.962.

The per-language split is where the comparison gets interesting. MonkeyOCR drops from 0.828 on English to 0.731 on Chinese, a gap of nearly a tenth, while OCRFlux differs by nine thousandths between the two. Most of the 0.187 headline margin is MonkeyOCR's weakness on Chinese, not OCRFlux being far ahead on English.

## On complex tables the same page puts OCRFlux behind MonkeyOCR

The second table uses Tree Edit Distance-based Similarity on OCRFlux-pubtabnet-single, which is derived from the public PubTabNet benchmark with some format transformation and holds 9064 HTML table samples, split into simple and complex according to whether cells carry rowspan or colspan. Splitting by rowspan and colspan is what makes the two rows genuinely different problems.

On simple tables OCRFlux-3B scores 0.912 against MonkeyOCR's 0.880, Nanonets-OCR-s 0.882 and olmOCR 0.810. On complex tables the order changes: MonkeyOCR reaches 0.826, OCRFlux-3B reaches 0.807, Nanonets-OCR-s 0.772 and olmOCR 0.676. OCRFlux still wins the Total at 0.861 against MonkeyOCR's 0.853, but that total is carried by the simple category, where the margin is 0.032. On the complex category it is behind by 0.019.

The page never mentions this. The feature list claims superior parsing quality and cites only the Edit Distance numbers, so the table benchmark where the model loses a row goes unremarked in the prose above it.

## The result tables carry inconsistent markup between rows

Both tables are hand-written HTML inside the Markdown, with rowspan cells grouping four models per language or per table type. The markup is not consistent across rows.

In the first table every OCRFlux-3B cell is wrapped in a strong tag and a link to the model on Hugging Face, which is how the project's own row is marked as the subject. In the complex-table row of the second table the emphasis goes the other way: MonkeyOCR's cell is wrapped in strong tags and OCRFlux-3B's cell is a bare link. The MonkeyOCR tag in that row is also written as an opening strong element with no closing slash, where the surrounding rows close theirs.

None of this changes a number, and all of it is the kind of thing that makes a results table harder to read than a generated one. It is worth knowing if you plan to cite these figures, because the emphasis on the page and the emphasis in the markup do not always point at the same model.

## Two dependencies are pinned to exact versions

The manifest requires Python 3.11 or later and lists a long dependency set: pypdf at 5.2.0 or newer, pypdfium2, torch at 2.5.1 or newer, plus cached-path, smart_open, cryptography, lingua-language-detector, Pillow, ftfy, bleach, markdown2, filelock, orjson, requests, zstandard, boto3, httpx, img2pdf, nltk, bs4, distance, apted, gradio and gradio_pdf.

Two of those are pinned with an equality sign rather than a floor: transformers at 4.50.0 and vllm at 0.7.3. Those are the two libraries most likely to move in step with a new model release, so a fixed version is a deliberate constraint on anyone who wants a newer serving stack or a newer checkpoint format.

The rest of the list explains what the pipeline is doing. The language detector matches the English and Chinese split of the benchmark, ftfy and bleach clean text, markdown2 writes it, orjson handles structured output, and distance with apted compute the tree edit distance that the table benchmark reports.

## The container installs Microsoft core fonts and pulls torch from a third-party index

The Dockerfile is the only installation procedure on the page, and it is unusually specific about the environment. It starts from ubuntu:24.04, installs a font set that includes fonts-crosextra-caladea, fonts-crosextra-carlito, gsfonts, msttcorefonts and ttf-mscorefonts-installer alongside poppler-utils and poppler-data, then fetches pip from bootstrap.pypa.io and installs the package with an extra wheel index:

```
python3.12 -m pip install . --find-links https://flashinfer.ai/whl/cu124/torch2.5/flashinfer/
```

Two things follow. The document parsing quality this project claims depends on having a particular set of fonts present, and two of those entries pull in Microsoft's core fonts, which carry their own licence terms and are not a neutral build dependency. And the torch wheels are not taken from PyPI alone but from a third-party host, pinned by a CUDA version in the path, so the image's torch build is decided by a URL rather than by the manifest.

## The entrypoint runs the pipeline, so the demo is not the container's job

The image finishes with a single entrypoint:

```
ENTRYPOINT ["python3.12", "-m", "ocrflux.pipeline"]
```

Everything in the container runs the pipeline module. That matters because the dependency list includes gradio and gradio_pdf, which are interface libraries, so the demo exists in the package while the image's default command is the conversion pipeline rather than the interface.

The build also cleans aggressively: pip runs with PIP_ROOT_USER_ACTION=ignore and PIP_BREAK_SYSTEM_PACKAGES=true so it can write into the system Python as root, PIP_NO_CACHE_DIR and PIP_DISABLE_PIP_VERSION_CHECK keep the layer small, and after the install the working directory is emptied with `rm -rf ./*` before the image is committed. The source is therefore present during the build and absent afterwards, which is the right default for an image this size.

## Version 0.1.0, an alpha status, and no releases at all

The version is 0.1.0 in two places: the project manifest, which also classifies the package as Development Status 3 - Alpha and Intended Audience Science/Research, and the news section, whose only entry reads Jun 17, 2025, v0.1.0, initial public launch and demo. The repository has no GitHub releases, so that dated line is the entire release history the page offers.

The last push was on 2026-04-14. So the version number has not moved since the middle of 2025 while commits have continued into 2026, and there is no tag to point at, no changelog entry to read and no release note describing what changed after the launch commit.

The repository layout is small and shows where the evaluation lives: a package directory, an eval directory, an images directory, the Dockerfile, the manifest and the license. Apache-2.0 is the declared license, and the manifest points its Homepage field at the repository itself rather than at a separate site.

## The cross-page section introduces the problem and stops before the numbers

Cross-page merging is the feature the page leans on hardest, and the phrasing is careful about it: native support for cross-page table and paragraph merging, and then a parenthetical saying that to the authors' best this is the first project among open source ones to support the feature.

The list of released artifacts names four datasets, OCRFlux-bench-single, OCRFlux-pubtabnet-single, OCRFlux-bench-cross and OCRFlux-pubtabnet-cross, so the cross-page benchmarks exist as published datasets. What the page does not carry is a result. The cross-page section opens by explaining that PDF documents are typically paginated, which often splits tables or paragraphs across consecutive pages, and there the visible text ends.

So the one capability no open source alternative is said to have is also the one with no figure on the page. A reader has to go to the dataset repositories or the linked blog to see whether the merging works, and the case studies are pointed at by a link rather than shown.

## Conclusion

Use OCRFlux when your documents are mostly single pages with ordinary tables, and treat the cross-page merging claim as unverified from this page. Before you adopt it, check four things. The complex-table numbers, where the project's own table puts OCRFlux-3B behind MonkeyOCR, and decide whether your tables are the simple kind that carries the win or the complex kind that does not. The cross-page benchmarks, which the page names as datasets but does not report results for, while claiming cross-page merging as a first in open source. The version, which is 0.1.0 with an alpha development status and no GitHub releases. And the runtime, because two dependencies are pinned to exact versions and the container installs a specific font set and pulls torch from a third-party wheel index, all of which have to line up with the hardware you run on.

## FAQ

### What model does OCRFlux use and what hardware does it need?

It is based on a 3B parameter vision language model released as OCRFlux-3B, and the page states it can run even on a GTX 3090 GPU. The manifest requires Python 3.11 or later and pins transformers at 4.50.0 and vllm at 0.7.3.

### How is OCRFlux evaluated, and are the benchmarks in its training data?

Two single-page benchmarks are reported: OCRFlux-bench-single with 2000 PDF pages, 1000 English and 1000 Chinese, scored by Edit Distance Similarity, and OCRFlux-pubtabnet-single with 9064 HTML table samples derived from PubTabNet, scored by Tree Edit Distance-based Similarity. The page states the released benchmarks are not included in the training or evaluation data.

### How do I install OCRFlux?

The page prints no install command. The Dockerfile is the only procedure shown: ubuntu:24.04, Python 3.12, a set of fonts including the Microsoft core fonts, and an entrypoint of python3.12 -m ocrflux.pipeline. The manifest itself requires Python 3.11 or later.

### Does OCRFlux merge tables and paragraphs across page boundaries?

The page claims native support for cross-page table and paragraph merging, described as the first among open source projects to the authors' knowledge. It names OCRFlux-bench-cross and OCRFlux-pubtabnet-cross as the cross-page benchmarks but reports no result for either on the page.

## Sources

- [chatdoc-com/OCRFlux on GitHub](https://github.com/chatdoc-com/OCRFlux)
- [Issues](https://github.com/chatdoc-com/OCRFlux/issues)
- [License: Apache-2.0](https://github.com/chatdoc-com/OCRFlux/blob/main/LICENSE)
- [README](https://github.com/chatdoc-com/OCRFlux/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/chatdoc-com-ocrflux
