Open-source project
baidu/Unlimited-OCR avatar
baidu/Unlimited-OCR

Unlimited-OCR: Baidu's One-Pass Long Document Parsing Model

Unlimited OCR parses long documents and images in one pass, producing structured text for document and AI workflows.

26,282 stars2,745 forksPythonMIT

At a glance

What is it?
Unlimited-OCR is Baidu's MIT-licensed vision-language model for parsing long documents and images in a single forward pass, producing structured text output. It supports inference via Hugging Face Transformers, vLLM, and SGLang, and targets document processing and AI data pipeline workflows that need high-quality OCR at scale on CUDA GPUs.
Who is it for?
Unlimited-OCR is the right tool for engineering teams building document-to-text pipelines that process lengthy PDFs or multi-page scans and need structured output with layout-aware parsing in a single model call. It requires NVIDIA GPU hardware (tested on CUDA 12.9 and CUDA 13.0) and a working Python 3.12 environment.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 63 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 25, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What Unlimited-OCR Solves and Who Uses It

Unlimited-OCR is a vision-language model that takes a document image (or a page extracted from a PDF) and produces structured text in one inference pass. The README describes the project goal as 'one-shot long-horizon parsing', meaning the model processes an entire document page, including headers, tables, figures, and body text, without requiring multiple calls or a preprocessing pipeline to split the document into regions first.

The primary target users are engineering teams building data pipelines that ingest large volumes of scanned documents, PDFs, or image-based content for use in AI training datasets, search indexes, or downstream language model workflows.

The model is available on Hugging Face at baidu/Unlimited-OCR and on ModelScope at PaddlePaddle/Unlimited-OCR. A live demo is at huggingface.co/spaces/baidu/Unlimited-OCR. The project paper is at arXiv:2606.23050, titled 'Unlimited OCR Works'.

Model Modes: Gundam versus Base Configuration

The model supports two inference configurations documented in the README as gundam and base. The gundam configuration uses base_size=1024, image_size=640, and crop_mode=True. The base configuration uses base_size=1024, image_size=1024, and crop_mode=False.

The distinction affects how the model handles high-resolution or large images. The gundam mode crops the input to the smaller image_size, which is useful for documents with dense small text. The base mode processes the full image at 1024 pixels without cropping, which is suited for images where spatial context across the full page matters.

The infer.py batch script uses gundam as the default image_mode. Users targeting dense academic papers or tables with small text should test both configurations to determine which produces better structured output for their document type.

Loading the Model with Hugging Face Transformers

The README documents inference using Hugging Face Transformers on NVIDIA GPUs. The tested environment is Python 3.12.3 with CUDA 12.9. The required packages include torch 2.10.0, torchvision 0.25.0, transformers 4.57.1, Pillow 12.1.1, einops 0.8.2, addict 2.4.0, easydict 1.13, pymupdf 1.27.2.2, and psutil 7.2.2.

Loading the model with bfloat16 precision:

python
import torch
from transformers import AutoModel, AutoTokenizer

model_name = 'baidu/Unlimited-OCR'
tokenizer = AutoTokenizer.from_pretrained(model_name, trust_remote_code=True)
model = AutoModel.from_pretrained(
    model_name,
    trust_remote_code=True,
    use_safetensors=True,
    torch_dtype=torch.bfloat16,
)
model = model.eval().cuda()

Note that `trust_remote_code=True` is required for both the tokenizer and model. The use_safetensors flag loads weights in the safetensors format rather than the older pickle-based format.

Serving with vLLM or SGLang

For production throughput, the README documents two serving backends. For vLLM, Docker images are provided for two GPU generations:

bash
docker pull vllm/vllm-openai:unlimited-ocr

A separate image targets Hopper GPUs (docker pull vllm/vllm-openai:unlimited-ocr-cu129). The full vLLM deployment recipe is at recipes.vllm.ai/baidu/Unlimited-OCR.

For SGLang, the README documents a specific wheel file (sglang-0.0.0.dev11416+g92e8bb79e) with pinned kernels==0.11.7. The server is started with:

shell
python -m sglang.launch_server \
    --model baidu/Unlimited-OCR \
    --served-model-name Unlimited-OCR \
    --attention-backend fa3 \
    --page-size 1 \
    --mem-fraction-static 0.8 \
    --context-length 32768 \
    --enable-custom-logit-processor \
    --disable-overlap-schedule \
    --skip-server-warmup \
    --host 0.0.0.0 \
    --port 10000

The server exposes an OpenAI-compatible API. The context-length flag is set to 32768 in the example. The SGLang server with the Unlimited-OCR model also requires a custom logit processor (DeepseekOCRNoRepeatNGramLogitProcessor) to function correctly, as shown in the client code in the README.

Batch Inference with infer.py

The repository includes an infer.py script for batch processing an image directory or a PDF file. The script starts the SGLang server automatically and sends concurrent requests:

shell
python infer.py \
    --image_dir ./examples/images \
    --output_dir ./outputs \
    --concurrency 8 \
    --image_mode gundam

For PDF input, the --image_dir flag is replaced with --pdf ./examples/document.pdf. Additional options documented in the README include --model_dir for specifying a local model path or Hugging Face model ID, --gpu for setting CUDA_VISIBLE_DEVICES, and --server_log for redirecting the SGLang server log.

The infer.py script handles PDF-to-image conversion internally. The SGLang serving dependency is required even for the batch script, since the script starts a local server rather than loading the model directly.

Limitations: GPU, Dependencies, and Output Post-Processing

Unlimited-OCR requires NVIDIA GPU hardware. The Transformers path is tested on CUDA 12.9; the vLLM Docker image targets CUDA 13.0 with a separate build for Hopper (CUDA 12.9). There is no CPU inference path documented and no mention of Apple Silicon or AMD GPU support.

The SGLang serving path requires a specific development wheel file (sglang-0.0.0.dev11416+g92e8bb79e) that is included in the repository's wheel/ directory. Using a different SGLang version may not work correctly with the model's custom logit processor.

The model produces output with structured markers including det tags that identify element types and bounding boxes. For downstream use in most NLP pipelines, post-processing is needed to strip these markers. The README includes a remove_det function example (partially shown) that strips the markers and groups lines by block type. The output format requires this post-processing step before the text can be fed into a language model or search index.

Unlimited-OCR versus DeepSeek-OCR and PaddleOCR

The README explicitly positions Unlimited-OCR as building on DeepSeek-OCR: 'aiming to push Deepseek-OCR one step further'. Both are vision-language model-based OCR systems aimed at structured document parsing. The acknowledged difference is that Unlimited-OCR targets long-document one-pass processing as its primary design goal.

PaddleOCR is another Baidu project (acknowledged in the README) and represents an older, pipeline-based OCR approach that uses separate detection and recognition models in sequence. PaddleOCR can run on CPU and lower-end hardware because its models are smaller. Unlimited-OCR is a unified model that requires a GPU with substantial VRAM, trading hardware accessibility for one-pass structured parsing.

Docling, mentioned in the search data as a comparison point, is a document parsing library from IBM that supports multiple backends and outputs structured formats including Markdown and JSON. Unlike Unlimited-OCR, Docling can run without a GPU using traditional OCR engines as a fallback. The choice between them depends on whether GPU hardware is available and whether the structured marker output format of Unlimited-OCR fits the downstream pipeline.

Editorial conclusion

Unlimited-OCR is the right tool for engineering teams building document-to-text pipelines that process lengthy PDFs or multi-page scans and need structured output with layout-aware parsing in a single model call. It requires NVIDIA GPU hardware (tested on CUDA 12.9 and CUDA 13.0) and a working Python 3.12 environment. The repository has no GitHub releases and the last push was on 2026-07-29; production deployments should pin to a specific commit rather than pulling main. Verify GPU memory requirements against the model size before deployment, as the README does not document VRAM minimums.

Frequently asked questions

What is Unlimited-OCR?

Unlimited-OCR is a vision-language model from Baidu that parses long documents and images in a single inference pass, producing structured text with layout-aware output. It is available on Hugging Face at baidu/Unlimited-OCR and described in a paper at arXiv:2606.23050.

How do you install and run Unlimited-OCR?

Install the required packages (torch 2.10.0, transformers 4.57.1, einops, pymupdf, and others) in a Python 3.12 environment with CUDA 12.9. Load the model from Hugging Face using AutoModel.from_pretrained with trust_remote_code=True. For serving, use the provided vLLM Docker image or the SGLang wheel in the repository's wheel/ directory.

How does Unlimited-OCR compare to DeepSeek-OCR?

The README describes Unlimited-OCR as aiming to push DeepSeek-OCR one step further, with one-pass long-horizon parsing as the stated improvement. Both are vision-language model-based OCR systems. The README acknowledges DeepSeek-OCR and DeepSeek-OCR-2 as related prior work.

How does Unlimited-OCR compare to Docling?

Unlimited-OCR is a vision-language model requiring NVIDIA GPU hardware and producing structured output with layout markers. The README does not describe Docling directly, but Docling is a document parsing library that supports multiple backends and can fall back to traditional OCR engines without a GPU, which makes it accessible on lower-end hardware.

Official sources

  1. Official README
  2. Project repository