Model or dataset
QIN2DIM/hcaptcha-challenger avatar
QIN2DIM/hcaptcha-challenger

hcaptcha-challenger: Solving hCaptcha with ONNX Models and Multimodal LLMs

🥂 Gracefully face hCaptcha challenge with multimodal large language model.

2,516 stars448 forksPythonGPL-3.0

At a glance

What is it?
A Python library that pairs local ONNX classifiers with multimodal LLM agents to answer hCaptcha challenges. It is GPL-3.0, alpha-status, and aimed at developers automating browser flows rather than at end users.
Who is it for?
Adopt it if you are automating your own browser flows in Python and can accept GPL-3.0 obligations plus Gemini API costs, and if you can live with a project whose own metadata marks it Development Status 3 - Alpha. Do not adopt it if you need a commercial licence without source disclosure, if you cannot send challenge screenshots to an external model provider, or if you want a maintained API with a stability guarantee.
Can I use it commercially?
Yes, with conditions. GPL-3.0 is a copyleft licence: if you distribute software that includes it, you must release that software's source code under the same licence. Running it internally without distributing it does not trigger that obligation.
Is it still maintained?
Yes. The repository last received commits 47 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What hcaptcha-challenger actually solves

hCaptcha sits between an automation script and the page it wants to reach. The usual answers are paid solving services, which mean handing a third party your session data, or a browser userscript, which means a human still has to click. hcaptcha-challenger takes a third route: it runs the challenge locally where it can, using ONNX models it downloads itself, and falls back to a multimodal language model for the cases the local models cannot decide.

The README is explicit about the two things it avoids: it does not rely on any Tampermonkey script, and it does not use any third-party anti-captcha service. That framing tells you who this is for. It is for developers writing Playwright-based automation in Python who want the solving step inside their own process, not behind someone else's API key. It is not a consumer tool, and the pyproject classifiers put it at Development Status 3 - Alpha, which is worth taking literally.

The five challenge types and what handles each one

The README's feature table maps each hCaptcha challenge type to a pluggable resource. image_label_binary uses a ResNet ONNX classifier. image_label_area_select in point mode uses YOLOv8 ONNX detection. The same challenge type in bounding box mode uses YOLOv8 segmentation, and image_drag_drop uses what the README calls Spatial Chain-of-Thought.

The Agent Capability column is the part most people will skim past. Three rows are marked with a check: image_label_binary, area select in point mode, and drag and drop. Two rows are marked with a dash: area select in bounding box mode, and image_label_multiple_choice, which uses a ViT ONNX zero-shot model. So the agentic workflow does not cover the full set. If your target site serves multiple-choice challenges, you are on the ViT model path, not the LLM path.

A second table lists advanced tasks: Rank.Strategy backed by nested-model-zoo, self-supervised challenge backed by CLIP-ViT, and the Agentic Workflow backed by an AIOps multimodal large language model. The dependency list confirms the LLM side is Google-specific: google-genai is a direct dependency, and the README's reference section links to the Gemini model docs.

Installing hcaptcha-challenger from PyPI

The package is published on PyPI as hcaptcha-challenger, so the install path is the ordinary one. Python 3.10 or newer is required according to requires-python in pyproject.toml.

bash
pip install hcaptcha-challenger

Playwright is a direct dependency, but the browser binaries are a separate step. The README does not spell this out in the section reproduced here, so check the docs directory for the current instruction before assuming the browser is present.

The repository ships an examples directory with an environment template. Copy it and fill in the values it defines rather than inventing your own variable names.

The examples folder contains demo_captcha_agent.py, demo_camoufox.py, demo_collector.py, demo_custom_agent_config.py and demo_statistic.py. For a first real run, demo_captcha_agent.py is the entry point that exercises the agent path; demo_custom_agent_config.py shows how the agent configuration is overridden. Read them before running, because the README does not document the environment variables these scripts expect. The docs directory is linked from the README as the documentation for English, Simplified Chinese, Russian and Vietnamese.

Where the LLM path breaks down

The agentic workflow sends challenge imagery to a multimodal model. That has a cost and a privacy consequence the README does not discuss. Every escalated challenge is an outbound request carrying a screenshot of the page the challenge appeared on. If you are automating a login flow, that screenshot may contain account identifiers or other page content. There is no documented local-only mode for the agent path, because the agent path is defined by the external model.

The second limitation is coverage. Two challenge types have no agent capability in the README table. The third is maturity: the project labels itself alpha, and the release history shows a rapid patch cadence, with v0.18.12 and v0.18.13 landing four days apart in October 2025 and v0.19.0 arriving in January 2026. Fast patch releases are normal for a target that changes its challenges, but they also mean the API surface moves.

The fourth is the one that decides most adoptions. hCaptcha is an adversarial target. Any solver that works today is working against a specific challenge distribution, and the collector and sentinel workflows in the repository exist precisely because the models need retraining as that distribution shifts. A project that solves hCaptcha is not a solved problem you install once.

How it compares to undetected-playwright and hosted solvers

The README points at two neighbouring projects. undetected-playwright, from the same author, attacks a different layer: it tries to hide the fingerprint of a Playwright-driven browser so the challenge is never issued in the first place. hcaptcha-challenger assumes the challenge appears and answers it. Those are complementary, not competing, and the README lists undetected-playwright under What's next.

The real alternative is a hosted solving service. The difference is where the work happens and who holds the data. A hosted service takes your sitekey and page URL, solves remotely, and returns a token; you pay per solve and you never run a model. hcaptcha-challenger runs the classifier in your process, downloads its own model artifacts from GitHub releases, and only calls out when the local model is not enough. You pay Google for the Gemini calls instead of paying a solver per challenge, and you keep the challenge imagery in your own pipeline up to the point of escalation. The trade is operational: you now own model downloads, browser binaries and a Python dependency tree that includes opencv-python, which is not a small install.

Licence, maintenance and upgrade cost

The licence is GPL-3.0-or-later, declared in pyproject.toml as the SPDX expression GPL-3.0-or-later and confirmed by the LICENSE file at the repository root. That is a copyleft licence. If you distribute software that links this library, the GPL's terms apply to the combined work. Whether your particular integration counts as distribution, and what obligations follow, is a question for your own counsel; the repository does not offer an alternative licence or a commercial exception. This is the single largest adoption constraint for anyone building a closed product.

On maintenance, the last push was on 2026-08-15, and the repository is not archived. The most recent tagged release in the list is v0.19.0 from 2026-01-05. The gap between the last release and the last push suggests work is landing on main that has not been tagged, so pinning to v0.19.0 gives you a stable point while main keeps moving.

Upgrade cost is driven by two things. The model artifacts are distributed through GitHub release assets, and the README's workflow table describes a model upload and upgrade path, so the models can change independently of the Python package. And the agent path depends on google-genai, whose model identifiers are referenced in the README's link to the Gemini 2.5 Pro preview docs. A preview model identifier can be retired on the provider's schedule, not yours.

Editorial conclusion

Adopt it if you are automating your own browser flows in Python and can accept GPL-3.0 obligations plus Gemini API costs, and if you can live with a project whose own metadata marks it Development Status 3 - Alpha. Do not adopt it if you need a commercial licence without source disclosure, if you cannot send challenge screenshots to an external model provider, or if you want a maintained API with a stability guarantee. Before writing any code, check the docs directory for the current agent configuration keys and read examples/demo_captcha_agent.py to see which challenge types the agent path actually handles, since the README table leaves two of the five types without agent support.

Frequently asked questions

How much does hCaptcha cost?

The repository does not state hCaptcha's pricing, and hcaptcha-challenger is a solver rather than a hCaptcha product, so it does not answer this. What the project does tell you is its own cost shape: the local ResNet, YOLOv8 and ViT ONNX models run in your process, while the agent path calls a Google Gemini model through the google-genai dependency.

How to solve CAPTCHA challenge?

hcaptcha-challenger maps each hCaptcha challenge type to a pluggable resource: image_label_binary to a ResNet ONNX classifier, area select in point mode to YOLOv8 detection, area select in bounding box mode to YOLOv8 segmentation, multiple choice to a ViT zero-shot model, and drag and drop to Spatial Chain-of-Thought. The README states it does not use Tampermonkey scripts or third-party anti-captcha services.

Can AI outsmart CAPTCHA?

This project is built on the premise that it can, at least for hCaptcha: it describes an AI versus AI approach and pairs local ONNX classifiers with a multimodal language model agent. The README's own capability table shows the agent path covers three of the five listed challenge types, so the answer the project gives is partial rather than total.

Is hCaptcha better than reCAPTCHA?

The repository does not compare the two, and hcaptcha-challenger targets hCaptcha only. Nothing in the README, pyproject.toml or repository layout addresses reCAPTCHA, so the project offers no basis for that comparison.

Official sources

  1. Issues
  2. License: GPL-3.0
  3. QIN2DIM/hcaptcha-challenger on GitHub
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/qin2dim-hcaptcha-challenger.svg)](https://hysenlabs.com/projects/qin2dim-hcaptcha-challenger)