# cnn_for_captcha: five captcha families solved with templates, YOLOv5, ResNet50 and a multimodal prompt

> anexplore/cnn_for_captcha is an Apache-2.0 Python repository with one script per captcha family, from fixed length text through slider, click text, rotation and similar object. It opens by telling readers to check whether they can avoid captchas altogether, and it vendors YOLOv5 and pins TensorFlow and PyTorch together.

**anexplore/cnn_for_captcha** — 图片类验证码识别(数字验证码/缺口验证码/文字验证码/旋转验证码/相似物体验证码)

- Repository: https://github.com/anexplore/cnn_for_captcha
- Stars: 338 · Forks: 82
- Language: Python
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/anexplore-cnn-for-captcha

## The README opens by telling you not to use the project

The first thing in this repository is a diff block, and it is reflective rather than promotional. It notes that multimodal large models can solve a great many captcha problems, and at the same time that those same models have produced more kinds of captcha. It then poses three open questions, marked as still to be observed: how much protective value captchas will retain now that automated AI applications such as openclaw are spreading, whether ordinary captchas get abandoned in favour of more advanced forms, and whether sites change their defences to adapt to model-driven traffic.

Two recommendations follow in highlighted text, and they are unusual enough to quote rather than paraphrase. First: before trying any method in this project, confirm whether the captcha can be avoided, and if it can be avoided in most cases, do not read further. Second: before trying, confirm whether brute-force enumeration already covers every case.

That ordering is the honest framing for a repository like this one. What follows is a set of working techniques, but the author is explicit that the first question is whether the problem exists, not how to solve it. Anyone evaluating this code should treat those two lines as the entry criteria rather than as a footnote.

The domain coverage is stated in the project description: numeric and fixed length text captchas, slider or gap captchas, text click captchas, rotation captchas, and similar object click captchas.

## requirements.txt pins TensorFlow and PyTorch in the same file

The dependency list is short and fully pinned for the deep learning parts: `tensorflow==2.9.1`, `torch==1.10.0`, `numpy==1.23.2`, `Pillow==9.1.1`, `requests`, and `opencv-python` constrained to `>=4.5.4, <4.6`. Two deep learning frameworks in one requirements file is not an accident, and it maps onto the split in the code. The rotation work uses ResNet50 for feature extraction, which is the TensorTorch path, while the detection work uses YOLOv5, which is the PyTorch path.

The opencv upper bound is worth noticing too, since anything newer than 4.6 will not satisfy it. The README describes `requirements.txt` as the fuller list and tells you to install what you need from it, which suggests the pinned set is a superset rather than a minimal requirement for any single script.

At the top level the layout is flat and readable. Each captcha family has its own module, `fixed_length_captcha.py`, `slide_captcha.py`, `rotate_captcha.py` and `sameobject_captcha.py`, with `image_utils.py` for shared image work, `split_data.py` for dataset splitting, and `labelme_json_to_yolov5_format.py` for converting annotation format. A `yolov5/` directory is vendored into the repository rather than installed as a package, which is why training commands refer to `train.py` and `yolov5s.yaml` directly. Configuration for the fixed length model lives in `fixed_length_captcha.json` next to the script.

## Fixed length text: one script, one config file, one naming convention

The fixed length text captcha is the simplest entry point and the only one with a stated accuracy figure. Training is a single command:

```python
python fixed_length_captcha.py
```

Input requirements are strict and worth restating, because they are what make the accuracy claim reproducible. The training set and the validation set go into the directories named in the config file, every image inside a directory has the same size, and the naming rule is the captcha text followed by `_` and a number, such as `abce_012312.jpg`. Uniform image size is doing real work here: a fixed input shape is what lets a single model serve every sample without a resize stage.

Prediction is exposed through a `Predictor` object with three entry points:

```python
predictor = Predictor()
# 预测本地磁盘文件
predictor.predict('xxx.jpg')
# 直接二进制内容预测
predictor.predict_single_image_content(b'PNGxxxxx')
# 预测远程图片
predictor.predict_remote_image('http://xxxxxx/xx.jpg', save_image_to_file='remote.jpg')
```

The three methods cover a local path, raw bytes and a remote URL, which is the range you need when the image arrives from a browser request rather than from a dataset on disk.

On results, the stated number is conditional and the condition is training set size: with a training set of around twenty thousand images, accuracy above 90 percent is reachable after training. No figure is given for smaller sets, and the general statement is that the outcome relates to the size of the training set.

## Slider captcha: template matching first, YOLOv5 when the rules stop holding

The slider, or gap, captcha has two documented approaches and the README is clear about their trade-off. The first is OpenCV template matching, described as simple and easy to verify, which reaches a satisfactory result when combined with some rules:

```python
import slide_captcha
slide_captcha.detect_displacement('image_slider.jpg', 'image_background.jpg')
```

Template matching compares a slider piece against a background image, which works when the piece is cut from the same rendering with the same distortion and noise. That is exactly why the second approach exists.

The second is YOLOv5 object detection, which needs labelled data and training but is described as a more stable and general solution than template probing. The repository ships 100 already labelled images to start from, and documents the parameters: batch size adjusted to available memory or VRAM and to the training result, epochs set by the training result, `yolov5s.yaml` from the models directory of the YOLOv5 project, `--img` as the image scaling baseline where the image width or height is suggested with a caution to lower it for large images, and `--weights` for a pretrained checkpoint, empty if you have none, with `yolov5s.pt` recommended.

```text
python train.py --batch-size 4 --epochs 200 --img 344 --data displacement.yaml --weights '' --cfg yolov5s.yaml
```

Detection is then done through a small wrapper, loading your trained checkpoint:

```python
import slide_captcha
detector = slide_captcha.DisplacementFinderByYolo()
detector.load_models('best.pt')
detector.detect_displacement('image.jpg', 344)
```

The demo result shown comes from a model trained on those 100 labelled images, so treat it as a starting point rather than a measurement.

## Click text captcha: localization is easy, matching is where it breaks

A click text captcha asks the user to select target characters from an image in a given order, and the target is supplied either as text or as an image of the text. The README splits the problem into localization and matching, and is candid that matching is the harder half.

Localization is straightforward. Use a YOLO-type detector to box the candidate character regions, and if the target itself arrives as an image, run detection on the target image too so both sides are in the same coordinate space.

Matching has three documented branches depending on how well OCR reads the image. If the characters survive detection and OCR, the task reduces to comparing recognised characters, with PaddleOCR, tesseract and cnocr all named as options. The interesting case is the one the README says is common: many captcha fonts are distorted and bolded specifically so OCR fails. Then it offers three workarounds in order. If the target is also an image, train a CNN to judge whether two input images are the same character, with a Siamese network cited as the reference architecture. If the number of candidate characters is limited to a few hundred or a few thousand, skip the pairing problem entirely and train YOLO to output the character class directly. If the target arrives as text, render that text into an image and reuse the first approach.

The second and third options are the interesting engineering judgement here, because a bounded label space converts a matching problem into a classification problem.

## Rotation captcha: the 0/1 formulation fails on class imbalance, regression does not

The rotation captcha requires turning a rotated image back to upright. Two approaches are documented, and the negative result is recorded as precisely as the positive one.

The first is brute force, and it depends on a specific property of the target site. If you can obtain the image database and it is not infinitely large, on the order of a hundred images for some sites, you can label the upright version of each by hand, then generate the rotated variants yourself, since rotation angles for this family are usually in the 10 to 30 degree range. For a target image you then find the most similar library image and read off its angle. Similarity is computed by extracting features with ResNet or a similar network and comparing cosine distance.

The second is training, framed as either regression or classification, and all three framings are reported with verdicts. Treating it as regression, where the model outputs an angle, is stated to be better than the binary approach. Treating it as a binary question where 1 means upright and 0 means not upright is reported as performing poorly, with the reason given precisely: for any one image there is exactly one upright version and N rotated versions, so the training data is badly imbalanced. Treating it as angle classification is marked as not tried.

`rotate_captcha.py` takes the regression route and uses ResNet50 for feature extraction. Two further notes: a text image orientation classification model in PaddleClas is cited as possibly instructive, and for matching the N rotations of one image, the duplicate finder in the imagededup module is suggested.

## Similar object clicking needs 200 or more labels, and h/r and C/G still collide

The last captcha family asks the user to click objects that are the same or similar. The reasoning for solving it with detection is that the range of object types is controllable, so a detector can classify what is present and the comparison happens after that. `sameobject_captcha.py` implements detection on YOLOv5.

The data requirement is stated plainly: at least 200 or more labelled images, with more images giving better accuracy. Annotating with labelme is supported, and `labelme_json_to_yolov5_format.py` converts the format.

The reported result is the most useful number in this section, because it is a failure case rather than an achievement. The demo shown is 300 labelled images trained for 100 rounds. The objects involved are partly letters, digits and triangular cones. In the output, h and r are confused with each other, C and G are confused with each other, U and cylinders are confused, and some objects are missed entirely. The stated conclusion is that increasing the training set size should reduce the error rate.

That is an honest limitation for a captcha solver to publish: the confusions are semantic, not positional, so more data helps but so does choosing a model large enough to separate visually similar classes. Nothing in the repository claims a fixed accuracy for this script, which is consistent with the rest of the project's habit of stating conditions alongside numbers.

## The multimodal results are a March 2026 snapshot, not current state

The last section asks what image-capable large models do with captchas, and the framing is honest about the ceiling. As of 2026.03 the note is that multimodal models do adequately on some simple captcha types and still fall short on distorted text. The suggestion is that a multimodal model or agent can be asked to recognise the characters and self-correct after an error rather than needing one pass.

The prompt given is worth reproducing, because it is a strict output contract rather than a request:

```text
you are an ocr tool. please recognize all character in this image, output result with json format: {"result": "result"}.
Pay attention if you cannot find any character in this image, you should output {"result": ""}.
Please just output json format, no explains.
```

Two dated entries follow. Google Gemini with gemini 3.1 pro is recorded on 2026.03.09: simple captchas have their own subsection, and for complex captchas the note is that the recognition rate improved while errors still remained. A later entry dated 2026-03-19 covers Nano Banana2 producing standard fonts for direct annotation. GPT 5.4 is recorded on 2026.03.09 with a complex captcha subsection.

Two papers are cited for general OCR capability rather than for captchas specifically: a quantitative evaluation of GPT-4V's OCR abilities and CogAgent, a visual language model for GUI agents.

The timestamps matter for how much weight to put on any of this. The repository's last push was on 2026-03-19, the same date as the newest entry, so nothing here has been revisited since, and the pinned dependencies date from the same period.

## Conclusion

cnn_for_captcha is useful as a map of which technique fits which captcha family, and as a set of runnable references for template matching, object detection and rotation regression. It is not a turnkey solver: the fixed length model needs roughly twenty thousand same-size training images you supply yourself, the rotation script needs a labelled image database from the site you are targeting, and the similar object detector is documented as needing 200 or more labels before h and r and C and G stop colliding. Read the two recommendations at the top before anything else, because the project itself says to check whether the captcha can be avoided and whether brute-force enumeration covers the case before looking at any of these methods. On maintenance, the repository is not archived but its last push was on 2026-03-19, and the dated model results in it stop at the same day, so the multimodal findings are a March 2026 snapshot rather than current state. The licence is Apache-2.0 and `requirements.txt` pins TensorFlow 2.9.1 alongside torch 1.10.0, so expect to reconcile those before training anything.

## FAQ

### What captcha types does cnn_for_captcha cover?

Five families, each with its own script: fixed length text captchas in fixed_length_captcha.py, slider or gap captchas in slide_captcha.py, rotation captchas in rotate_captcha.py, and similar object click captchas in sameobject_captcha.py. Click text captchas are covered as an approach writeup using detection for localization and OCR, Siamese CNNs or direct classification for matching.

### What accuracy does cnn_for_captcha report for text captchas?

For fixed length text, a training set of around twenty thousand images is reported to reach above 90 percent accuracy after training, with the result explicitly tied to training set size. No figure is given for the similar object detector, which is only described as needing at least 200 labelled images, with 300 images and 100 rounds used for the demo.

### How does cnn_for_captcha handle rotation captchas?

rotate_captcha.py treats rotation as a regression problem and uses ResNet50 for feature extraction. The README reports that a binary upright or not upright formulation performs poorly because there is one upright image per N rotated ones, and that regression is better, while angle classification is marked as not tried.

### What does cnn_for_captcha require to install and run?

requirements.txt pins tensorflow 2.9.1, torch 1.10.0, numpy 1.23.2, Pillow 9.1.1, requests and opencv-python between 4.5.4 and 4.6, and the README says to install from it as needed. YOLOv5 is vendored into the repository rather than installed, and training uses split_data.py to cut a prepared image directory into train and validation sets.

### When was cnn_for_captcha last updated?

Its last push was on 2026-03-19, and the repository is not archived. The dated multimodal model entries in the README stop at the same day, recording Gemini with gemini 3.1 pro and GPT 5.4 on 2026.03.09 and a Nano Banana2 note on 2026-03-19.

## Sources

- [anexplore/cnn_for_captcha on GitHub](https://github.com/anexplore/cnn_for_captcha)
- [Issues](https://github.com/anexplore/cnn_for_captcha/issues)
- [License: Apache-2.0](https://github.com/anexplore/cnn_for_captcha/blob/main/LICENSE)
- [README](https://github.com/anexplore/cnn_for_captcha/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/anexplore-cnn-for-captcha
