Open-source project
xiaofengShi/CHINESE-OCR avatar
xiaofengShi/CHINESE-OCR

CHINESE-OCR: Legacy End-to-End Chinese Scene-Text Detection with CTPN, CRNN, and CTC

End-to-end Chinese scene-text detection and recognition with CTPN, CRNN, and CTC (legacy project).

2,957 stars941 forksPythonLicense varies

At a glance

What is it?
xiaofengShi/CHINESE-OCR implements a three-network pipeline for detecting and recognizing Chinese text in natural scene images, using CTPN for text region detection and CRNN with CTC for variable-length recognition. The maintainer has marked it a historical project; the Python 3.6 and TensorFlow 1.x requirements make it unsuitable for new production deployments.
Who is it for?
Researchers who need to reproduce a 2019-era CTPN-CRNN-CTC pipeline for Chinese scene text, or who want to study its architecture, will find this repository a useful reference. The Python 3.6 and TensorFlow 1.x requirements make it unsuitable for any production deployment today.
Can I use it commercially?
Not without permission. GitHub finds no licence file in the repository, and without a licence all rights are reserved by default: you may read the code but not reuse it. Check the README, or ask the authors, before using it.
Is it still maintained?
Yes. The repository last received commits 41 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 25, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What CHINESE-OCR Does and Who It Is For

xiaofengShi/CHINESE-OCR is an end-to-end pipeline for detecting and recognizing Chinese text in natural scene images. It was built to address a specific problem: reading text that appears in photographs of real-world scenes, such as street signs, product labels, and printed documents, rather than cleanly scanned documents. The pipeline handles text at four orientations (0, 90, 180, and 270 degrees), detects the bounding regions of text lines, and then reads the content of those regions without needing to know in advance how long the text is.

The README explicitly marks the repository as a historical project in maintenance mode. The maintainer states that it is based on Python 3.6, TensorFlow 1.x, and early versions of Keras and PyTorch, and is kept for research reproduction and community reference rather than active development. Problem responses may be slow, and dependency versions and download links may be outdated. For researchers who want to reproduce results from a specific period of Chinese OCR development, or who are studying the CTPN-CRNN-CTC approach, this repository provides a working reference implementation. For anyone building a new system, newer frameworks have largely replaced this stack.

The Three-Network Architecture

The pipeline chains three separate networks, each responsible for a distinct stage of the recognition process.

The first network is a text orientation classifier based on VGG16. It classifies an input image into one of four orientation categories: 0, 90, 180, or 270 degrees. The README states that this model was trained on 8,000 images and achieved 88.23 percent accuracy on the orientation classification task.

The second network is CTPN (Connectionist Text Proposal Network), responsible for detecting horizontal text regions. CTPN uses a fixed-width anchor design where all anchors share the same width of 16 pixels, but vary in height across ten values: 11, 16, 23, 33, 48, 68, 97, 139, 198, and 283 pixels. The anchor generation code in the repository reflects this:

python
def generate_anchors(base_size=16, ratios=[0.5, 1, 2],
                     scales=2 ** np.arange(3, 6)):
    heights = [11, 16, 23, 33, 48, 68, 97, 139, 198, 283]
    widths = [16]
    sizes = []
    for h in heights:
        for w in widths:
            sizes.append((h, w))
    return generate_basic_anchors(sizes)

Because all anchors have fixed width, CTPN can only detect horizontal text. The README notes that vertical text detection would require modifying the anchor generation function.

The third network is CRNN (Convolutional Recurrent Neural Network) combined with CTC decoding. It takes a detected text region as input and outputs the recognized string without requiring per-character position labels. The architecture uses CNN layers for visual feature extraction, GRU or LSTM layers for sequential modeling, and CTC loss for training. Both Keras and PyTorch versions of the CRNN training code are included.

Setting Up the Environment

The repository provides three shell scripts for environment setup, covering GPU, CPU, and Python 3 CPU configurations:

bash
sh setup.sh

For CPU-only environments:

bash
sh setup-cpu.sh

For Python 3 on CPU:

bash
sh setup-python3.sh

The README specifies the target environment as Python 3.6 with TensorFlow 1.7. This combination is significantly out of date relative to current Python and TensorFlow releases. Setting it up on a modern system will typically require using a virtual environment or container to isolate the old dependency versions, and some packages may no longer be installable from their original sources.

The pretrained model weights for all three networks are hosted on Baidu Cloud. The orientation classifier weights, the CTPN checkpoint, the CTPN dataset, and both Keras and PyTorch CRNN weights each have separate Baidu Cloud links in the README. The README warns that download links may be outdated, and verifying whether those links are still active is the first practical step before investing time in the environment setup.

Once the environment is ready and weights are downloaded, running the pretrained model requires editing demo.py to specify the path to a test image. The CTPN detection results can be visualized by modifying the draw_boxes function in ./ctpn/ctpn/other.py.

Training on Your Own Data

The repository supports training all three networks on custom data. For the orientation classifier, the training is a standard image classification task using VGG16 as the backbone.

For the CTPN text detection network, training requires the pretrained VGG ImageNet weights from Baidu Cloud and a dataset in Pascal VOC format. The dataset path is set by pointing the self.devkit_path parameter in ./ctpn/lib/datasets/pascal_voc.py to the dataset directory. The repository includes a link to a prepared CTPN dataset on Baidu Cloud.

For the CRNN recognition network, both Keras and PyTorch training scripts are provided. The Keras version is at ./train/keras_train/train_batch.py and takes model_path (the pretrained weight location) and MODEL_PATH (the output save location) as parameters. The PyTorch version at ./train/pytorch-train/crnn_main.py uses argparse and expects the pretrained weight path and the output directory to be specified as command-line arguments.

The README recommends the PyTorch CRNN implementation as more stable than the Keras version. The README also notes that the current CRNN architecture is relatively shallow and suggests that using ResNet or DenseNet as the feature extractor with multi-layer bidirectional RNNs and attention would improve recognition accuracy, though this was a planned improvement at the time of writing, not a completed one.

Core Limitations and Why It Is a Research Reference

The limitations are substantial and the README does not obscure them. The Python 3.6 and TensorFlow 1.x requirements place this repository well outside the current mainstream Python ecosystem. TensorFlow 1.x reached end of life years ago. Installing it on a modern system alongside current packages is technically possible but requires careful dependency pinning and will not work with the default pip environment on any current operating system.

The Baidu Cloud download links for pretrained weights are a second risk. Baidu Cloud links expire or change, and the README itself warns that links may be outdated. A user who cannot retrieve the weights cannot run the pretrained demo at all.

The CTPN network's fixed-width anchor design limits it to horizontal text detection. Detecting rotated or curved text requires architectural changes not present in the current codebase. The training data used for the orientation classifier and the character set used for recognition are both restricted to Chinese and English letters, so symbols, mathematical notation, and other character sets will not be recognized.

The repository also carries no open-source license. The README states explicitly that the repository has no attached open-source license, and that default copyright is retained by the author. Commercial use, redistribution, or large-scale derivative development requires contacting the maintainer for permission.

Comparison with Modern Chinese OCR Tools

PaddleOCR is Baidu's open-source OCR toolkit for Python. It supports modern Python 3.x versions, is actively maintained, and handles both text detection and recognition in a more integrated package than CHINESE-OCR's manually chained three-network setup. PaddleOCR supports curved and rotated text through newer detection architectures, covers a wider character set, and ships with pretrained models that can be downloaded through its own model zoo rather than through external cloud links.

The key difference in approach is generation: CHINESE-OCR is a 2019-era research implementation that chains CTPN, a VGG classifier, and a CRNN into a manual pipeline. PaddleOCR is a current production-grade framework with a more unified API. For a new project requiring Chinese text recognition, PaddleOCR represents the current state of the field. For a researcher who specifically needs to reproduce the CTPN-CRNN-CTC combination as it was implemented in 2019, CHINESE-OCR provides the original code.

Editorial conclusion

Researchers who need to reproduce a 2019-era CTPN-CRNN-CTC pipeline for Chinese scene text, or who want to study its architecture, will find this repository a useful reference. The Python 3.6 and TensorFlow 1.x requirements make it unsuitable for any production deployment today. The README does not carry an open-source license, meaning copyright is retained by the author; any use beyond personal research should involve contacting the maintainer. Before setting up the environment, verify that the Baidu Cloud download links for the pretrained weights are still active.

Frequently asked questions

What is OCR and what does CHINESE-OCR specifically do?

OCR stands for optical character recognition. CHINESE-OCR specifically implements end-to-end recognition of Chinese text in natural scene images, using three networks: a VGG16 orientation classifier, a CTPN text detection network, and a CRNN with CTC for variable-length character sequence recognition.

What Python and framework versions does CHINESE-OCR require?

The README specifies Python 3.6 with TensorFlow 1.7. These are significantly outdated versions; running the project on a modern system requires an isolated environment with pinned dependencies.

Can CHINESE-OCR detect rotated or vertical text?

The CTPN network uses fixed-width anchors and is designed for horizontal text only. The README notes that vertical text detection would require modifying the anchor generation function. The orientation classifier handles 0, 90, 180, and 270 degree rotations at the image level, but CTPN itself cannot detect text that runs diagonally or vertically within the image.

Official sources

  1. Issues
  2. Project website
  3. README
  4. xiaofengShi/CHINESE-OCR on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/xiaofengshi-chinese-ocr.svg)](https://hysenlabs.com/projects/xiaofengshi-chinese-ocr)