Open-source project
zai-org/GLM-OCR avatar
zai-org/GLM-OCR

GLM-OCR: Optical character recognition using a multimodal language model

GLM-OCR: Accurate × Fast × Comprehensive

7,482 stars670 forksPythonApache-2.0

At a glance

What is it?
GLM-OCR is an optical character recognition tool that uses a multimodal language model to extract text from images. It handles complex layouts and multiple languages.
Who is it for?
GLM-OCR suits teams that need accurate OCR for complex documents and various languages. It is most useful when standard OCR tools produce poor results on scanned documents or images with unusual layouts.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 162 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

OCR with multimodal language models

GLM-OCR uses a multimodal language model for optical character recognition. Traditional OCR tools like Tesseract work by analyzing pixel patterns and recognizing character shapes in isolation. They perform well on clean, uniformly-formatted text, but struggle with complex layouts, poor image quality, or mixed-language documents. GLM-OCR takes a fundamentally different approach: it uses a language model that understands both images and text simultaneously. This allows it to reason about context, layout, and semantic meaning while extracting text. The model understands that text flows in certain directions, that tables have rows and columns with logical relationships, and that some text is labels while other text is content data. This contextual and semantic understanding leads to more accurate results on complex documents. The model can infer missing or unclear characters based on surrounding text and context, improving accuracy on challenging real-world documents.

Accuracy and layout understanding

The project name emphasizes three qualities: accurate, fast, and comprehensive. Accuracy means the extracted text matches what is in the image. The model achieves this by using contextual reasoning, not just pixel analysis. Layout understanding means the tool preserves structure: tables stay tables with correct cell organization, paragraphs stay paragraphs, and columns are recognized correctly. Comprehensive means it handles varied document types: forms, charts, handwriting, and printed text. Multiple languages are supported. The model scores 94.62 on OmniDocBench V1.5, ranking first overall.

Installation and deployment options

GLM-OCR is available from the repository at https://github.com/zai-org/GLM-OCR. The SDK supports multiple installation profiles for different scenarios:

bash
# Cloud / MaaS + local images / PDFs (fastest install)
pip install glmocr

# Self-hosted pipeline (layout detection)
pip install "glmocr[selfhosted]"

# Flask service support
pip install "glmocr[server]"

For development installations, clone the repository and set up a virtual environment:

bash
git clone https://github.com/zai-org/glm-ocr.git
cd glm-ocr
uv venv --python 3.12 --seed && source .venv/bin/activate
uv pip install -e .

The project requires Python 3.10 or higher. Once installed, you can process images through the command-line interface via the `glmocr` command or through Python code. Input can be a single image or a batch of images. The system accepts various image formats: PNG, JPEG, TIFF, and others. Output includes extracted text and layout information in structured format. The README documents the API with detailed examples of common usage patterns like batch processing and language configuration. You can customize the model checkpoint and parameters for your specific use case. A cloud API is available through Zhipu's MaaS at https://open.bigmodel.cn with documentation at https://docs.z.ai/guides/vlm/glm-ocr for teams preferring managed hosting. The technical report at https://arxiv.org/abs/2603.10910 provides detailed benchmarks and evaluation results.

Language models vs Tesseract's pixel analysis

Traditional OCR tools like Tesseract work by analyzing pixel patterns and recognizing characters independently. They often struggle with irregular layouts, poor image quality, or unusual fonts because they lack contextual understanding. GLM-OCR's multimodal approach adds semantic understanding: the model knows what text usually looks like in context, so it can correct errors and understand structure. This is particularly useful for complex documents like forms, receipts, or mixed-language documents. The model can infer missing or unclear characters based on surrounding text and context, and it preserves the logical structure of documents: tables stay tables with correct row and column relationships, paragraphs maintain their coherence, and the document's semantic organization is respected. Tesseract excels on clean, uniformly-formatted text, while GLM-OCR is more robust on scanned documents with low resolution, images with unusual layouts like diagrams or handwritten notes, and documents with varied formatting or visual complexity.

Model efficiency and deployment scenarios

GLM-OCR uses a lightweight 0.9B parameter model, making it significantly more efficient than larger language models. The model integrates the CogViT visual encoder with a GLM-0.5B language decoder, combined with a two-stage pipeline of layout analysis and parallel recognition based on PP-DocLayout-V3. This architecture achieves state-of-the-art performance with a score of 94.62 on OmniDocBench V1.5, ranking first overall. The SDK supports multiple installation profiles for different deployment scenarios: the basic profile installs cloud API support or local image/PDF processing (fastest); the selfhosted profile adds layout detection capabilities; and the server profile includes Flask service support for running your own OCR endpoint. The model supports deployment through multiple frameworks: vLLM for high-concurrency services, SGLang for advanced scheduling, and Ollama for simple local deployment. Apple Silicon Macs can use MLX via the mlx-vlm deployment guide. The model supports speculative decoding with vLLM using Multi-Token Prediction (MTP) loss. Development installations require Python 3.10 or higher and support various input formats including PNG, JPEG, TIFF and others. GLM-OCR prioritizes accuracy over speed. Model inference takes more time than traditional OCR methods because the language model must process the entire image and reason about context. For a single document, expect processing time in seconds rather than milliseconds. However, the accuracy improvement is substantial on complex documents. Batch processing is supported, allowing you to process multiple images in parallel. The choice between traditional and model-based OCR depends on your requirements: if you need high accuracy on complex documents and can tolerate longer processing times, GLM-OCR is appropriate.

When GLM-OCR is appropriate

Use GLM-OCR when standard OCR tools fail or produce poor results. It is particularly useful for scanned documents with low resolution, images with unusual layouts like diagrams or handwritten notes, mixed-language documents, and forms with complex structures. The technical report at https://arxiv.org/abs/2603.10910 provides detailed benchmarks on formula recognition, table recognition, and information extraction. The model handles code-heavy documents, seals, and other challenging real-world layouts, making it ideal for business use cases where document understanding is critical. The cloud API is available through Zhipu's MaaS at https://open.bigmodel.cn with documentation at https://docs.z.ai/guides/vlm/glm-ocr. GLM-OCR is less useful for simple printed text where traditional OCR already works well, since GLM-OCR may have higher latency due to model inference. Consider your accuracy and speed requirements when choosing between traditional and model-based OCR.

Editorial conclusion

GLM-OCR suits teams that need accurate OCR for complex documents and various languages. It is most useful when standard OCR tools produce poor results on scanned documents or images with unusual layouts. The project documentation provides installation and usage instructions.

Frequently asked questions

What is OCR?

OCR stands for optical character recognition. It is the process of converting images of text into machine-readable text.

How is GLM-OCR different from traditional OCR?

GLM-OCR uses a multimodal language model to understand context and layout, while traditional OCR like Tesseract analyzes pixel patterns. GLM-OCR produces better results on complex documents.

What languages does GLM-OCR support?

GLM-OCR supports multiple languages. The documentation lists supported languages and how to configure them.

Can GLM-OCR handle tables and forms?

Yes. GLM-OCR understands layout and structure, preserving tables and form fields in the output.

Is GLM-OCR fast?

GLM-OCR prioritizes accuracy over speed. Model inference takes more time than traditional OCR, but produces more accurate results on complex documents.

Official sources

  1. Issues
  2. License: Apache-2.0
  3. README
  4. Releases
  5. zai-org/GLM-OCR on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/zai-org-glm-ocr.svg)](https://hysenlabs.com/projects/zai-org-glm-ocr)