Open-source project
anexplore/cnn_for_captcha avatar
anexplore/cnn_for_captcha

cnn_for_captcha: A Python Toolkit for Five Kinds of Image CAPTCHA Recognition

图片类验证码识别(数字验证码/缺口验证码/文字验证码/旋转验证码/相似物体验证码)

338 stars82 forksPythonApache-2.0

At a glance

What is it?
The repository collects separate scripts for fixed-length text, slider, click-text, rotation and same-object CAPTCHAs, mixing classical OpenCV matching with Keras and YOLOv5 models. It is a set of worked examples, not a packaged library, and its own README tells you to check whether you can avoid the CAPTCHA first.
Who is it for?
Adopt it if you already have labelled images of one specific CAPTCHA and want a working reference for the training loop, the YOLOv5 annotation format or the rotation regression setup. Do not adopt it if you need a supported library with an API, since there is no setup.py, no CLI and no released version, and the scripts use pinned TensorFlow 2.9.1 and torch 1.10.0.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Activity is slowing. The repository last received commits 6 months ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What cnn_for_captcha actually covers, and who it is written for

The repository is a collection of standalone Python scripts, each aimed at one CAPTCHA family: fixed-length text CAPTCHAs, slider CAPTCHAs, click-the-text CAPTCHAs, rotation CAPTCHAs and click-the-same-object CAPTCHAs. The README describes the project as deep-learning-based image CAPTCHA recognition, and the file list matches that description: fixed_length_captcha.py, slide_captcha.py, rotate_captcha.py, sameobject_captcha.py, plus image_utils.py, split_data.py and a yolov5 directory. There is no package, no command line entry point and no released version.

The intended reader is someone who already has a pile of labelled images and wants to see how each problem was framed. That framing is the useful part. For fixed-length text the README states the input requirements plainly: training and validation images go in directories named in the config file, every image in a directory must be the same size, and filenames follow the pattern captcha_text_index.format, with the example abce_012312.jpg. That convention is what lets the script derive labels from filenames instead of a separate annotation file.

The README opens with two suggestions that are worth repeating because they set expectations. First, check whether the CAPTCHA can be avoided at all; if it usually can, the project says not to bother with it. Second, check whether the CAPTCHA space is small enough to enumerate by brute force. A toolkit whose own documentation starts by arguing you may not need it is unusual, and it is the most honest thing in the file.

How the fixed-length text pipeline turns filenames into training labels

The fixed-length script is the most complete path in the repository. Labels come from the filename, not from a manifest: abce_012312.jpg means the answer is abce. The config file fixed_length_captcha.json points at the training and validation directories, and the README says its fields are self-explanatory, which is a way of saying the JSON is not documented in prose. You have to open it to learn the keys.

Training is one command, python fixed_length_captcha.py. Prediction is exposed through a Predictor class with three entry points: predict for a local file, predict_single_image_content for raw bytes, and predict_remote_image for a URL, with a save_image_to_file argument. That split is practical. The bytes variant is what you would call from a service that already has the image in memory, and the remote variant avoids a separate download step.

The accuracy claim in the README is conditional, and the condition matters: it says results depend on training set size, and that for a particular style of image, roughly 20,000 training samples reached above 90 percent accuracy. That is one reported figure for one image style, not a general benchmark. If your CAPTCHA has more characters, more distortion or a different font, the number tells you nothing. The README also does not document rollback, checkpoint resumption or how to export the trained model, so plan on reading the script if you need those.

Installing the dependencies and running a first prediction

There is no install step in the usual sense. The README points at requirements.txt and says it is fairly complete and can be installed as needed. The file pins tensorflow==2.9.1, Pillow==9.1.1, numpy==1.23.2, torch==1.10.0, requests, and opencv-python with the range >=4.5.4, <4.6. Those pins are old relative to the last push on 2026-03-19, and TensorFlow 2.9.1 plus torch 1.10.0 is not a combination you can drop into an arbitrary modern environment without checking Python version support.

A minimal setup looks like this:

bash
pip install -r requirements.txt

Before training you need the directory layout the config expects. The split helper takes an existing image directory and divides it, and the README notes the destination directories must already exist:

bash
python split_data.py all_image_dir train_image_dir validation_image_dir 0.9

After training, prediction from a script uses the Predictor class the README shows:

python
predictor = Predictor()
predictor.predict('xxx.jpg')
predictor.predict_single_image_content(b'PNGxxxxx')
predictor.predict_remote_image('http://xxxxxx/xx.jpg', save_image_to_file='remote.jpg')

What you should see is the predicted text for the image. The README does not specify the return type or the output format, so treat the first run as a way to discover that.

Slider CAPTCHAs: template matching first, YOLOv5 when that stops working

The slider script offers two approaches and the README is explicit about the trade-off. The first is OpenCV template matching, called with slide_captcha.detect_displacement and a slider image plus a background image. The README calls it simple and easy to verify, and says it reaches a satisfactory result when combined with some rules. The phrase some rules is doing a lot of work there; template matching degrades when the background is noisy or the gap is rendered with transparency, and the README does not describe what those rules are.

The second approach trains a YOLOv5 model to detect the gap, which the README describes as more stable and general than template matching, at the cost of annotation and training. The repository includes 100 annotated images for this. The training command is given in full:

text
python train.py --batch-size 4 --epochs 200 --img 344 --data displacement.yaml --weights '' --cfg yolov5s.yaml

The README explains each argument: batch size by available memory, epochs by observed results, img as the resize baseline (use the image width or height, and lower it if images are large), and weights left empty or set to a pretrained yolov5s.pt. Detection then goes through DisplacementFinderByYolo, loading best.pt and calling detect_displacement with the image and the size 344.

The honest reading is that 100 annotated images is a starting point, not a dataset. If your slider CAPTCHA varies in background, the template path will break first and you will be annotating more images.

Click-text and rotation CAPTCHAs are where the repository gets thin

For click-the-text CAPTCHAs the README describes a two-stage plan rather than a finished script: locate candidate characters with a YOLO-style detector, then match the target text against them. The matching stage is where it admits difficulty. If the characters survive detection well enough for OCR, the README lists PaddleOCR, tesseract and cnocr as options and says the problem becomes easy. It then notes that many CAPTCHAs distort or embolden the text so OCR fails, and offers three fallbacks: train a Siamese network to judge whether two images are the same character, use YOLO to predict the class directly when the character set is only a few hundred or thousand, or render the target text as an image and reuse the Siamese approach. Those are ideas, not code paths in the repository.

Rotation is better specified. The README gives a brute-force route first: if the underlying image library is small enough to enumerate, label the upright images by hand, generate rotated variants at the 10 to 30 degree steps typical of these CAPTCHAs, then match a new image by similarity using CNN features and cosine distance. It suggests imagededup's find_duplicates for grouping the many rotated copies of one source image.

The learned route is framed as regression versus classification. The README reports that a 0/1 classification of upright versus not-upright performed poorly because of class imbalance (one upright image against N rotated ones), that regression on the angle worked better, and that angle classification was not attempted. rotate_captcha.py implements the regression approach with ResNet50 features. That is a rare thing in a README: a negative result stated plainly.

Same-object CAPTCHAs, the 300-image result, and where the model confuses classes

The same-object script uses YOLOv5 to detect object classes in the image, then compares them. The README asks for at least 200 annotated images and says more is better, and it ships a converter, labelme_json_to_yolov5_format.py, for annotations made in labelme.

The reported result is 300 annotated images trained for 100 epochs. The README shows the output and describes the failure mode directly: h and r, C and G, and U against a cylinder are confused or misdetected. It attributes this to dataset size and expects more data to reduce the error rate. That is a reasonable diagnosis, but it also means the shipped setup is a demonstration. Three hundred images across several classes is a few dozen examples per class, and the confusion pairs listed are exactly the shapes that differ by a stroke or a curve.

The larger point about this script is that object detection gives you class labels, not a similarity judgement. If the CAPTCHA asks you to click the object that matches a reference image, and the reference is itself a small image, you still need a matching step the repository does not implement. The click-text section's Siamese suggestion applies here too, and it is not wired into sameobject_captcha.py.

Multimodal models as an alternative, and when they are the better choice

The README's own alternative is a multimodal model rather than another Python library. It dates its observations to 2026.03 and reports that for simple CAPTCHA types the success rate of image-capable models is acceptable, while recognition of distorted text still falls short. It includes a short OCR-style prompt asking the model to return JSON with a result field, and advises outputting an empty result when no character is found. Screenshots are shown for Gemini and for GPT, with a note that a model was used to annotate standard fonts.

Compared with the scripts in this repository, the difference in approach is total. The scripts require you to collect labelled images, pick an architecture and train a model per CAPTCHA family; the model route requires an API call and a prompt, and the README notes that an agent can retry after a wrong answer. What you give up is determinism and cost control: every attempt is a request, latency is not yours to tune, and the README does not quantify accuracy for any specific CAPTCHA, only describes it qualitatively as improving but still error-prone on complex cases.

For a one-off job on a handful of images, calling a model is less work than annotating 300. For a high-volume, stable CAPTCHA where you control the training set, a trained model is cheaper per call and does not depend on an external service. The README's own two opening suggestions sit above both options: if the CAPTCHA can be avoided, or brute-forced, neither is the right tool.

Maintenance, licence and what a year of dependency drift costs

The repository is not archived. The last push was on 2026-03-19, which is roughly six months before the date this article was written, so there is recent activity but nothing that establishes a release cadence. There are no retrieved releases, so there is no versioned artifact to pin against and no changelog to read. Upgrades happen by pulling the branch.

That matters because of the pins. tensorflow==2.9.1 and torch==1.10.0 are exact versions, and opencv-python is constrained to >=4.5.4, <4.6. Installing this alongside a modern ML stack in the same environment will conflict. The practical cost is a dedicated virtual environment per script family, since the slider and same-object scripts lean on a vendored yolov5 directory while the fixed-length and rotation scripts lean on TensorFlow and torch respectively. The README does not describe a shared environment that satisfies all of them.

The licence is Apache-2.0 according to the repository metadata, and a LICENSE file is present at the top level. Apache-2.0 includes an explicit patent grant and requires that modifications be noted, which is friendlier for commercial use than a copyleft licence, but the vendored yolov5 directory carries its own upstream licence and the README links to the ultralytics project for training instructions. Check both files before shipping anything. This is a description of the terms, not legal advice.

There is also a boundary the README does not discuss: how you are permitted to use a CAPTCHA solver depends on the site and the jurisdiction, and the project says nothing about it.

Editorial conclusion

Adopt it if you already have labelled images of one specific CAPTCHA and want a working reference for the training loop, the YOLOv5 annotation format or the rotation regression setup. Do not adopt it if you need a supported library with an API, since there is no setup.py, no CLI and no released version, and the scripts use pinned TensorFlow 2.9.1 and torch 1.10.0. Before writing any code, read the README's first two suggestions and confirm the CAPTCHA cannot simply be avoided, and check whether the answer space is small enough to enumerate by brute force. Then verify the licence file matches the Apache-2.0 identifier shown on the repository, and confirm the CAPTCHA type you face is one of the five the scripts cover.

Frequently asked questions

How do I pass a CAPTCHA test?

The README's first suggestion is to confirm whether the CAPTCHA can be avoided at all, and its second is to check whether the answer space is small enough to enumerate exhaustively. Only after those does it recommend the recognition methods, which require labelled images of the specific CAPTCHA.

Is bypassing a CAPTCHA illegal?

The repository does not address legality or permitted use anywhere in the README. The licence file covers the terms for the code itself, not for how you apply it to a third-party site.

Why do I keep getting asked to do CAPTCHAs?

The README does not explain why sites present CAPTCHAs. It only notes that the spread of AI automation applications raises open questions about how much protection CAPTCHAs still provide and whether sites will move to other forms.

Can a robot pass the CAPTCHA test?

The repository's own README says multimodal models handle simple CAPTCHA types acceptably as of 2026.03 but still fall short on distorted text, and that trained models reach above 90 percent on one fixed-length style given roughly 20,000 training images.

Official sources

  1. anexplore/cnn_for_captcha on GitHub
  2. Issues
  3. License: Apache-2.0
  4. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/anexplore-cnn-for-captcha.svg)](https://hysenlabs.com/projects/anexplore-cnn-for-captcha)