# chineseocr_lite: CPU Chinese OCR with a 4.7MB ONNX Model Bundle

> chineseocr_lite is a Python OCR library for Chinese text that runs entirely on CPU using ONNX Runtime, packaging detection, orientation classification, and sequence recognition into three models totaling 4.7MB. It is suited to scripts, local pipelines, and embedded builds where installing a GPU stack is not practical.

**DayBreak-u/chineseocr_lite** — 超轻量级中文ocr，支持竖排文字识别, 支持ncnn、mnn、tnn推理 ( dbnet(1.8M) + crnn(2.5M) + anglenet(378KB)) 总模型仅4.7M 

- Repository: https://github.com/DayBreak-u/chineseocr_lite
- Stars: 12,350 · Forks: 2,273
- Language: C++
- License: GPL-2.0
- Published: 2026-09-21 · Updated: 2026-09-21 · Language: en
- Canonical page: https://hysenlabs.com/projects/daybreak-u-chineseocr-lite

## CPU OCR for Chinese Text Without a GPU

chineseocr_lite was written for developers who need to extract Chinese text from images without installing CUDA or running a separate inference server. The README describes its target environment as Python 3.6, running on Windows, Linux, or macOS, with CPU inference via ONNX Runtime and no requirement for a CUDA-capable device. That constraint is an explicit design choice: the Python runtime drops GPU acceleration entirely to keep the deployment footprint small.

The use case is narrow but well-defined. A batch pipeline that processes scanned documents on a cloud function with no GPU allocation, a desktop tool that needs offline text extraction, or a mobile prototype that calls the web service locally all fit the project's design. The project does not claim general OCR coverage: it targets Chinese characters, including vertical text, and the three bundled models are optimised for that specific task.

## Three Models, One Pipeline: DBNet, AngleNet, and CRNN

Text recognition in chineseocr_lite passes through three stages. The first model, `dbnet.onnx`, is a text detector that locates regions containing Chinese text. It weighs 1.8MB. The second, `angle_net.onnx`, classifies the orientation of each detected region so that vertical text lines and rotated blocks are correctly upright before recognition. It weighs 378KB. The third, `crnn_lite_lstm.onnx`, performs sequence recognition on each oriented region using a CTC-based decoder. It weighs 2.5MB.

All three ONNX model files live in the `models/` directory. When the Python package is installed via pip, those files are bundled with the package because pyproject.toml includes a `[tool.setuptools.package-data]` entry for `models = ["*.onnx"]`. The CLI therefore works without a separate model download step after installation. The `models_ncnn/` directory contains `.bin` and `.param` files for the C++ NCNN demos and is not required for the Python path.

## Installing chineseocr_lite and Starting the Web Service

The web service is the fastest way to verify that the models work. Clone the repository, install the dependencies:

```bash
pip install -r requirements.txt
```

Then start the Tornado backend:

```bash
cd chineseocr_lite
python backend/main.py
```

The server listens on port 8089. The README states that the terminal prints a line such as `server is running: 192.168.x.x:8089` after startup. Opening that address in a browser shows a web page where you can upload an image and see the detected text blocks.

The requirements.txt pins exact dependency versions: tornado 5.1.1, numpy 1.19.1, opencv-python 4.3.0.36, onnxruntime 1.4.0, Shapely 1.7.0, pyclipper 1.2.0, and Pillow 7.2.0. These are old releases. Installing into a shared environment that uses newer versions of any of these packages will require an isolated virtual environment to avoid conflicts.

The Docker image in the repository uses CentOS 7.2.1511 as its base and exposes ports 5000 and 8000, not 8089. The Dockerfile installs from requirements.txt and runs `python3 backend/main.py` as its command.

## Using the CLI for Batch Processing

To use the command-line interface, install the package in development mode:

```bash
pip install -e .
```

Then pass an image path to the `chineseocr` command:

```bash
chineseocr test_imgs/res.jpg
```

The CLI prints a JSON object to stdout. The top-level `text` field holds the full recognized string. The `blocks` array contains one object per text region, each with `text`, `score`, and `box` fields. The `elapsed` field reports inference time in seconds. The README shows this output structure:

```json
{
  "text": "识别出的全文",
  "blocks": [
    {
      "text": "单个文本块",
      "score": 0.93,
      "box": [[12, 30], [210, 31], [209, 60], [11, 59]]
    }
  ],
  "elapsed": 1.24
}
```

You can write the result to a file and generate an annotated image in one step:

```bash
chineseocr test_imgs/res.jpg --output result.json --draw result.jpg
```

The `--compress` flag sets the short-side dimension before detection, which trades speed for accuracy on very large or very small images:

```bash
chineseocr test_imgs/res.jpg --compress 960
```

The package also exposes the same entry point via `python -m chineseocr_lite test_imgs/res.jpg`.

## Cross-Platform Demos: C++, Android, JVM, and .NET

Beyond the Python implementation, the repository ships reference demos for four other platforms. The C++ demos in `cpp_projects/` cover three backends: an ONNX Runtime demo that runs on CPU only for Windows, Linux, and macOS; an NCNN demo with both CPU and Vulkan GPU variants; and an MNN demo that is CPU-only. The JVM demos in `jvm_projects/` expose a Java and Kotlin API by compiling ONNX Runtime or NCNN as a JNI library. The Android demos in `android_projects/` mirror the C++ choices: ONNX Runtime, NCNN (with CPU and GPU variants), and MNN. The .NET demos in `dotnet_projects/` provide C# and VB.NET wrappers over ONNX Runtime.

The README describes each of these as an independent reference implementation, translated from the Python version. They are not published as versioned packages on any package registry. Developers who need them are expected to integrate the source code directly. The README also notes that a complete Android source archive with all dependency libraries is available in the project's QQ group rather than as a GitHub release.

## Limitations and Cases Where It Falls Short

The Python deployment has several concrete constraints. GPU acceleration in Python is not documented; the README and pyproject.toml both describe the path as CPU-only ONNX Runtime. The NCNN C++ demo supports Vulkan, but the Python web service and CLI do not.

The pinned dependencies are old. onnxruntime is pinned to 1.4.0 in requirements.txt, a release from 2020. The pyproject.toml specifies a looser bound of `>=1.4.0,<=1.20.1`, but the older requirements.txt file pins the exact version. Projects that already use current versions of numpy, Pillow, or OpenCV will need to isolate this package.

The project has no GitHub releases. The last push to the repository was on 2026-05-18. There is no versioned binary distribution to fetch.

There is also a licence inconsistency. The top-level repository licence is recorded as GPL-2.0, which requires derived works to be released under the same terms. The pyproject.toml declares Apache-2.0, which has different requirements. This mismatch is unresolved in the repository and needs to be addressed before any commercial or proprietary deployment.

For non-Chinese languages, chineseocr_lite does not provide coverage. The models are trained specifically for Chinese characters. Attempting to use them on English, Japanese, or other scripts will produce unreliable output.

## Alternatives and Wider OCR Context

EasyOCR is the most frequently compared alternative. It is a multi-language OCR library written in Python that supports over 80 languages, including Chinese, and can run on both CPU and GPU. It uses PyTorch as its inference engine, which means the dependency footprint is larger than chineseocr_lite's ONNX-only stack. Projects that already use PyTorch and need broad language coverage will find EasyOCR a better fit.

RapidOCR is another project visible in the related searches. It also targets CPU inference with ONNX Runtime and is structured around the same DBNet and CRNN model family. The key qualitative difference is that RapidOCR ships as a maintained Python package on PyPI, while chineseocr_lite is built from source. The README makes no comparison between the two projects.

For callers that need vertical Chinese text specifically, chineseocr_lite's angle classification step is a concrete capability that not all lightweight OCR projects include. The README lists vertical text recognition as a primary feature.

## Conclusion

chineseocr_lite suits developers building Chinese OCR into CPU-bound environments: serverless functions, local tools, or mobile apps where the Python packaging path is acceptable. The 4.7MB model bundle and stable CLI JSON output make it straightforward to embed in scripts. It is not the right choice when GPU acceleration is needed in Python, or when a project cannot tolerate pinned old dependencies. Before deploying, verify the licence: the top-level repository lists GPL-2.0 while pyproject.toml declares Apache-2.0, and both claims require reconciliation before a production deployment.

## FAQ

### Does chineseocr_lite require a GPU to run?

No. The Python runtime uses ONNX Runtime for CPU-only inference. No CUDA driver or GPU is required for the Python web service or CLI. The C++ NCNN demo optionally supports Vulkan GPU acceleration, but that requires compiling the C++ code separately.

### What does the chineseocr CLI output?

The command prints a JSON object with a top-level `text` field containing the full recognized string, a `blocks` array where each item includes `text`, `score`, and `box` coordinates, and an `elapsed` field reporting inference time in seconds.

### Does chineseocr_lite support vertical Chinese text?

Yes. The README lists vertical text recognition as a feature. The `angle_net.onnx` model classifies the orientation of each detected text region before passing it to the recognition model, so vertical lines are handled in the pipeline.

## Sources

- [DayBreak-u/chineseocr_lite on GitHub](https://github.com/DayBreak-u/chineseocr_lite)
- [Issues](https://github.com/DayBreak-u/chineseocr_lite/issues)
- [License: GPL-2.0](https://github.com/DayBreak-u/chineseocr_lite/blob/master/LICENSE)
- [README](https://github.com/DayBreak-u/chineseocr_lite/blob/master/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/daybreak-u-chineseocr-lite
