CLI tool
breezedeus/CnSTD avatar
breezedeus/CnSTD

CnSTD: Scene Text, Math Formula and Layout Detection in One Python Package

CnSTD: 基于 PyTorch/MXNet 的 中文/英文 场景文字检测(Scene Text Detection)、数学公式检测(Mathematical Formula Detection, MFD)、篇章分析(Layout Analysis)的Python3 包

794 stars114 forksPythonApache-2.0

At a glance

What is it?
CnSTD wraps DBNet, YOLOv7 and PaddleOCR-derived detectors behind a single Python API and CLI, covering scene text, math formula and layout analysis. It is a detection-only toolkit, so anyone expecting recognized text has to pair it with an OCR package.
Who is it for?
Adopt CnSTD if you need bounding boxes for Chinese or English text, inline and isolated math formulas, or ten layout element classes, and you are willing to run a second package for recognition. Do not adopt it if you want a single call that returns a string, or if you need a documented rollback path for model downloads, because the README describes model fetching but not what happens when a download or a new model version goes wrong.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 89 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What CnSTD detects, and what it leaves to other tools

CnSTD is a Python 3 package for scene text detection (STD), math formula detection (MFD) and layout analysis. The README describes it as supporting Chinese and English text detection with bundled trained models, usable right after installation. Since V1.2.1 it also ships MFD models that separate inline formulas (labeled embedding) from displayed formulas (labeled isolated).

The distinction that matters most is the one the README states plainly: if you need to recognize the text inside a detected box, you combine CnSTD with the OCR package cnocr. CnSTD produces geometry. It does not produce strings. That makes it a building block for a pipeline rather than a finished extraction tool, and it changes how you evaluate it. A bad box coordinate is a detection failure; a bad character is not CnSTD's problem.

Who it is for: teams processing Chinese documents where formulas and page structure matter, for example academic PDFs or scanned textbooks. The layout model recognizes ten element classes: Text, Title, Figure, Figure caption, Table, Table caption, Header, Footer, Reference and Equation. That is a coarse but useful partition for downstream reading-order work.

DBNet as the default detector, and the external model path

Since V1.1.0 the in-house text detector is DBNet, replacing PSENet from V0.1. The README claims DBNet cut detection time by nearly an order of magnitude relative to PSENet while also improving accuracy. Treat that as the author's comparison, not an independent measurement.

Six pretrained DBNet variants are listed with parameter counts, file sizes, an IoU figure and an average inference time per image. The table carries its own caveat: the timings were measured on a local Mac and the absolute numbers have little reference value, only the relative ordering does. The IoU computation was also adjusted, so it too is comparative rather than absolute. That kind of self-disclosure is rarer than it should be, and it means you should benchmark on your own hardware rather than quoting the table.

On accuracy the spread is narrow: db_resnet34 reports 0.7322 IoU, db_mobilenet_v3 0.7269, db_shufflenet_v2_small 0.7190. The size and speed spread is not narrow at all. db_resnet34 is 86 MB and 3.11 seconds per image; db_mobilenet_v3_small is 7.9 MB and 1.24 seconds. A 0.027 IoU difference in exchange for roughly 11x the file size and 2.5x the latency is a trade-off most batch pipelines should think about explicitly.

Beyond its own models, CnSTD supports detectors ported from other engines. V1.2.5 added PP-OCRv4 detection via RapidOCR, V1.2.6 added PP-OCRv5 with ch_PP-OCRv5_det and ch_PP-OCRv5_det_server, and V1.2.8 added PP-OCRv6 multilingual models: multi_PP-OCRv6_det_tiny, multi_PP-OCRv6_det_small and multi_PP-OCRv6_det_medium. V1.2.4 added an Ultralytics-based YOLO detector. These external models are converted to ONNX and served through CnSTD's interface, which is why the ONNX extras exist.

Installing CnSTD and running a first detection

The README's installation section is one line for the common path. Python 3.8 or higher is required, and opencv is a dependency that may need separate installation.

bash
pip install cnstd

If you are behind a slow connection, the README suggests a mirror, for example the Aliyun index:

bash
pip install cnstd -i https://mirrors.aliyun.com/pypi/simple

ONNX models need an extra. Pick the CPU or GPU variant depending on your machine, and note the README's warning that an existing onnxruntime installation must be removed first with pip uninstall onnxruntime.

bash
pip install cnstd[ort-cpu]

The Makefile in the repository shows the CLI shape for each task. This is the formula detection invocation, pointed at an example image shipped with the repo:

bash
cnstd analyze -m mfd --conf-thresh 0.25 --resized-shape 700 --img-fp examples/mfd/zh4.jpg

For layout analysis the same command takes a different model name and resize value:

bash
cnstd analyze -m layout --conf-thresh 0.25 --resized-shape 800 --img-fp examples/mfd/zh.jpg

What you should see is a set of boxes with labels rather than text. The --conf-thresh flag filters weak detections; lowering it surfaces more candidates at the cost of false positives. The repository also ships a Streamlit demo, invoked through the Makefile as streamlit run cnstd/app.py after installing streamlit, which is the fastest way to eyeball results before writing integration code.

Where the model downloads and the training path get thin

The models live on HuggingFace under cnstd-cnocr-models and the breezedeus/cnstd-ppocr-* repositories, and the README says they are free to download. What the README does not document is what happens when a download fails partway, whether a partially fetched file is retried, or how to pin a model version so that an upstream update does not silently change your outputs. For a pipeline that runs unattended, the absence of a documented rollback or pinning story is a real operational gap, not a nitpick.

Training is a second gap. The Makefile exposes cnstd train with a config file and an input directory, and examples/train_config.json and examples/train_config_gpu.json are present in the repository. But the README points elsewhere for the math formula model: the MFD training code lives in a separate yolov7 repository on the dev branch. So the package gives you inference for MFD, and training for it lives in another project. If your goal is to fine-tune formula detection on your own annotations, budget time for that split.

There is also a data-format consideration. The README mentions scripts for Label Studio, which generate a JSON file importable into Label Studio after running CnSTD's MFD model over a directory, and convert Label Studio exports back into the format the MFD training expects. That round trip is the intended annotation workflow, and it assumes you are comfortable with Label Studio as an external dependency.

How CnSTD compares with calling PaddleOCR or RapidOCR directly

The honest alternative is to use PaddleOCR or RapidOCR on their own rather than through CnSTD. CnSTD's external models are, by the README's own description, ported from PaddleOCR and converted to ONNX for use inside CnSTD. If text detection is all you need and you are happy with PaddleOCR's own Python API, going straight to the source removes a layer and one set of version constraints.

The difference in approach is what CnSTD adds on top. First, a uniform interface: the same cnstd analyze call serves a DBNet model trained by this project, a PaddleOCR-derived detector, and a YOLO detector. Second, tasks PaddleOCR's detection API does not cover in the same package: math formula detection with the embedding and isolated distinction, and ten-class layout analysis. Third, the CnSTD-trained DBNet family, which is independent of the PaddleOCR lineage entirely.

So the choice is not about which detector is better. It is about whether you want one package that spans text, formulas and layout, or three separate integrations. If your work is pure Chinese text detection and nothing else, the extra abstraction buys you little. If you are assembling a document understanding pipeline, the shared box format (four point coordinates, unified since V1.0.0) is the actual value.

Licence and the cost of staying current

CnSTD is licensed under Apache-2.0. The setup.py header carries the standard Apache notice, and the repository includes a LICENSE file. Apache-2.0 permits commercial use and modification and includes a patent grant, but it also carries notice and attribution obligations, so redistributing a modified CnSTD means preserving the licence and notices. This is a summary of what the licence identifier means, not legal advice; check the LICENSE file and your own counsel for anything consequential.

On maintenance: the repository is not archived, and the last push was on 2026-07-05. The release history shows a steady cadence of small releases rather than long silences: v1.2.7.2 on 2026-04-28 fixing rapidocr model_root_dir support, v1.2.7.3 on 2026-05-01 improving subprocess call handling in the YOLOv7 modules, and v1.2.8 on 2026-07-05 adding PP-OCRv6 multilingual support. The pattern is maintenance plus periodic model refreshes as upstream PaddleOCR ships new detectors.

That cadence is also the upgrade cost. Each PP-OCR generation (v4, v5, v6) arrived as a new set of model names, and v1.2.8 added a --lang-type CLI flag specifically for the RapidOCR v6 detectors. Upgrading is therefore not just a version bump in your requirements file; model names and flags change alongside it. Pinning your cnstd version and your model name together is the practical approach. The requirements.txt is pip-compile output pinned to exact versions, which tells you the project itself treats dependency pinning as normal.

Frequently asked questions

See the faq field.

Editorial conclusion

Adopt CnSTD if you need bounding boxes for Chinese or English text, inline and isolated math formulas, or ten layout element classes, and you are willing to run a second package for recognition. Do not adopt it if you want a single call that returns a string, or if you need a documented rollback path for model downloads, because the README describes model fetching but not what happens when a download or a new model version goes wrong. Before committing, verify that your Python is 3.8 or higher, that opencv is present, and that the model you want (db_mobilenet_v3 for speed, db_resnet34 for accuracy) is actually reachable from your network.

Frequently asked questions

Does CnSTD recognize the text inside the boxes it detects?

No. CnSTD performs detection and returns box coordinates; the README states that to recognize the text in a detected box you combine it with the OCR package cnocr. Detection and recognition are separate steps.

What Python version does CnSTD require?

The README says to use Python 3.8 or higher. It also notes that opencv is a dependency and may need to be installed separately.

How do I use an ONNX model with CnSTD?

Install the ONNX extra, either pip install cnstd[ort-cpu] for CPU or pip install cnstd[ort-gpu] for GPU. The README warns that if onnxruntime is already installed in the environment, you must uninstall it manually before running the install command.

Which CnSTD model should I pick for speed?

The README's comparison table lists db_mobilenet_v3_small as the smallest and fastest at 7.9 MB and 1.24 seconds per image, with an IoU of 0.7054, while db_resnet34 reports the highest IoU at 0.7322 but is 86 MB and 3.11 seconds. The README notes the timings were measured on a local Mac and only the relative values are meaningful.

Can CnSTD detect math formulas as well as text?

Yes. Since V1.2.1 CnSTD includes math formula detection models that label inline formulas as embedding and displayed formulas as isolated. The README states the MFD models were trained on the English IBEM dataset and the Chinese CnMFD_Dataset.

Official sources

  1. breezedeus/CnSTD on GitHub
  2. License: Apache-2.0
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/breezedeus-cnstd.svg)](https://hysenlabs.com/projects/breezedeus-cnstd)