# text-detection-ctpn: A Frozen-Graph TensorFlow Rework of CTPN for Horizontal Scene Text

> This TensorFlow reimplementation of the Connectionist Text Proposal Network targets horizontal scene text detection, with a frozen-graph demo path and an oriented text connector still marked as experimental. The project is a pragmatic port, not a maintained library, and its value depends on how closely you can match the original training setup.

**eragonruan/text-detection-ctpn** — GitHub describes it as text detection mainly based on ctpn model in tensorflow, id card detect, connectionist text proposal network. The repository metadata lists Python as its primary language. The metadata lists the MIT license. This article stays within the project description and details documented in the GitHub repository README.

- Repository: https://github.com/eragonruan/text-detection-ctpn
- Stars: 3,427 · Forks: 1,310
- Language: Python
- License: MIT
- Published: 2026-08-13 · Updated: 2026-08-18 · Language: en
- Canonical page: https://hysenlabs.com/projects/eragonruan-text-detection-ctpn

## What This Project Actually Solves

The repository is a TensorFlow implementation of the Connectionist Text Proposal Network (CTPN), originally published in the 2016 paper 'Detecting Text in Natural Image with Connectionist Text Proposal Network'. The README states that the model can be used for 'almost every horizontal scene text detection task', with ID card detection as the demonstration example. The intended user is a developer or researcher who wants to run CTPN without using the original Caffe codebase, and who prefers TensorFlow. The project explicitly targets horizontal text; the oriented text connector is present but flagged as 'still need further improvement'. This is a narrow scope: if your images contain rotated text, this project will likely not deliver reliable results. The value is in the frozen-graph inference path, which lets you run detection without building the custom NMS libraries, and in the training pipeline that follows the paper's loss function.

## Architecture and Data Flow

The project follows the CTPN architecture: a VGG16 backbone (pre-trained weights are required) followed by a BLSTM layer and a text proposal network that outputs vertical anchor proposals. The README lists the roadmap items: freezing the graph for inference, pure Python and Cython NMS, CUDA NMS, the loss function as referred in the paper, an oriented text connector, and BLSTM. The data flow for inference is straightforward: load a frozen .pb file, read images from data/demo, run the graph, and save results to data/results. For training, the pipeline is more involved: you prepare training data as text lines with character-level or proposal-level annotations, convert them to VOC format, and then train the network. The text.yml configuration file controls key parameters: USE_GPU_NMS toggles CUDA NMS, DETECT_MODE selects between horizontal (H) and oriented (O) modes, and checkpoints_path points to the model weights. The architecture is not novel, but the implementation choices matter: the frozen graph removes the need to rebuild the library for inference, which is a practical convenience for deployment.

## Getting It Running: Demo and Training Commands

The README gives two distinct paths. For a quick demo, you do not need to build any library: clone the repository, download the ctpn.pb file from the releases, place it in data/, put images in data/demo, and run 'python ./ctpn/demo_pb.py'. This is the simplest entry point and works with a pre-trained model. For training, you need Python 2.7, TensorFlow 1.3, Cython 0.24, OpenCV-Python, and easydict; the README recommends Anaconda. If you have a GPU, build the custom NMS libraries by running 'cd lib/utils', 'chmod +x make.sh', and './make.sh'. For CPU-only setups, the README points to an issue for setup instructions. Data preparation requires downloading a pre-trained VGG model and placing it at data/pretrain/VGG_imagenet.npy, then preparing training data either from the provided links or by following the paper. The commands are: 'python split_label.py' to generate prepared data, 'python ToVoc.py' to convert to VOC format, then symlink the TEXTVOC folder to VOCdevkit2007. Finally, run 'python ./ctpn/train_net.py'. The README states that the provided model was trained on a GTX1070 for 50k iterations, with roughly 0.2s per iteration when using CUDA NMS, totaling about 2.5 hours. These commands are concrete, but the data preparation steps are underspecified; you must read the paper to understand the label format.

## The Oriented Text Connector: A Known Weak Point

The README explicitly says the oriented text connector 'has been implemented, it's working, but still need further improvement.' This is the main limitation for anything beyond horizontal text. The DETECT_MODE parameter switches between H and O, but the project does not claim robustness for arbitrary orientations. In practice, scene text often appears at angles, and this project will likely fail on rotated text. The horizontal mode is the default and the primary use case. If your application involves ID cards or documents that are roughly upright, the horizontal mode is acceptable. For general scene text with varied orientations, this project is the wrong tool. The lack of a published evaluation dataset or quantitative results in the README makes it hard to assess accuracy; the only evidence is a few sample images. This is a genuine gap: you cannot know how well it performs on your data without running it yourself.

## Training Data Preparation: The Hidden Cost

The training pipeline is not plug-and-play. The README says to prepare training data 'as referred in paper', which means you must understand the CTPN labeling scheme, typically text line bounding boxes with vertical anchor annotations. The provided scripts split_label.py and ToVoc.py automate some conversion, but you must modify paths and ground truth paths according to your dataset. This is a significant time investment. The README offers pre-prepared data from Google Drive or Baidu Yun, but those links may not be stable, and the data format is not described in detail. If you plan to train on your own dataset, expect to spend days on data annotation and format conversion. The project does not include a data generator or augmentation pipeline, so you must rely on the paper's methodology. This is a maintenance burden: the code is from 2018, and the dependencies are old, so reproducing the training environment on a modern system may require fixing compatibility issues.

## Comparison with the Original Caffe Implementation and Alternatives

The README links to the original Caffe implementation by tianzhi0549/CTPN. The main difference is the framework: this project uses TensorFlow, while the original uses Caffe. For a user already in the TensorFlow ecosystem, this port is more convenient, but the original Caffe version may have more community testing and updates. Another alternative is the EAST (Efficient and Accurate Scene Text Detector) model, which handles arbitrary orientations directly, unlike CTPN's horizontal focus. EAST is a different architecture that uses a fully convolutional network to predict rotated boxes, and it is often faster. If your text is not horizontal, EAST is a better choice. The README does not mention EAST, but the trade-off is clear: CTPN is designed for horizontal text, while EAST covers more orientations. This project's value is its simplicity for horizontal cases, but it is not a general solution.

## Maintenance, License, and Upgrade Costs

The repository has not been pushed to since June 2018, and the only release is an untagged v0.99 from that date. This means no active maintenance, no bug fixes for newer TensorFlow versions, and no updates to address dependency drift. The license is MIT, which is permissive and allows commercial use, modification, and redistribution, but you are responsible for any legal compliance; this is not legal advice. The code depends on Python 2.7 and TensorFlow 1.3, both of which are obsolete. Upgrading to TensorFlow 2.x would require significant changes to the graph construction and session handling. The frozen-graph inference path might survive with minor tweaks, but training would need a rewrite. If you adopt this project, plan to fork it and maintain it yourself. The checkpoints_path parameter in text.yml gives you control over model loading, but the training script is not designed for distributed training or modern hardware. The cost of upgrading is high, and the lack of a community means you are on your own.

## Conclusion

Adopt text-detection-ctpn if you need a TensorFlow-based CTPN implementation for horizontal text detection, especially for ID card or document scans, and you are comfortable with Python 2.7, TensorFlow 1.3, and manual data preparation. Do not use it if you require production-grade robustness on arbitrary scene text, oriented text beyond the experimental connector, or long-term maintenance. Before adopting, verify that your environment still supports TensorFlow 1.3 and Python 2.7, download the provided frozen model and test it on your own images, and inspect the text.yml parameters to ensure DETECT_MODE and checkpoint paths match your use case. The repository has not been updated since June 2018, so expect to port code yourself if you need newer TensorFlow versions.

## FAQ

### What is eragonruan/text-detection-ctpn?

A TensorFlow implementation of the connectionist text proposal network for text detection, ported from the original Caffe repository. It is demonstrated with identity card detection and is MIT licensed.

### What does text-detection-ctpn require to run?

The README lists python2.7, tensorflow1.3, cython0.24, opencv-python and easydict. requirements.txt pins easydict 1.7, tensorflow_gpu 1.3.0, scipy 0.18.1, numpy 1.11.1, opencv_python 3.4.0.12, Cython 0.27.3, Pillow 5.0.0 and PyYAML 3.12.

### How do I run the text-detection-ctpn demo without building?

Clone the repository with --depth=1, download ctpn.pb from the releases page into data/, put images into data/demo, then run `python ./ctpn/demo_pb.py`. Results are written to data/results.

### How long does training text-detection-ctpn take?

The README states about 0.2 seconds per iteration with CUDA nms, so about 2.5 hours for 50,000 iterations, and says the provided checkpoint was trained on a GTX1070 for 50k iterations. No CPU timing is given.

### What does DETECT_MODE do in text-detection-ctpn?

It selects horizontal or oriented detection, with H as the default. The README says the oriented text connector has been implemented and works, but still needs further improvement.

## Sources

- [Official README](https://github.com/eragonruan/text-detection-ctpn#readme)
- [Project repository](https://github.com/eragonruan/text-detection-ctpn)
- [Release notes](https://github.com/eragonruan/text-detection-ctpn/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/eragonruan-text-detection-ctpn
