Open-source project
sarperavci/CloudflareBypassForScraping avatar
sarperavci/CloudflareBypassForScraping

CloudflareBypassForScraping draws no boundary around its targets: any host named in a header, on a port published by default

A cloudflare verification bypass script for webscraping

2,610 stars387 forksPythonMIT

At a glance

What is it?
CloudflareBypassForScraping is a self-hosted HTTP service that fronts a patched stealth browser to obtain Cloudflare clearance for a host the caller names, then replays requests against it. The front page documents no per-target authorisation, no allowlist and no authentication, while the quick start publishes port 8000 on the host, so anyone who can reach it inherits a bypass for any site they care to name. This article covers the mechanism and the deployment facts and deliberately stops short of the request recipes.
Who is it for?
Do not expose this service to anything but your own machine. The project is a bypass for a protection the target's operator deliberately deployed, and the front page offers no per-target authorisation, no allowlist, no dry run and no authentication, while its own quick start publishes the port to the host.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 53 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 4, 2026, and from our analysis. They are not legal advice.

Editorial analysis

There is no boundary around which host this may be pointed at

This is the fact that should decide whether you run it at all. The target is not configured anywhere. It arrives per request, in a header called `x-hostname`, and the page documents no allowlist, no per-target credential, no rate limit keyed to a destination and no dry run. The quick start is one command that publishes the port to the host:

bash
docker run -p 8000:8000 ghcr.io/sarperavci/cloudflarebypassforscraping:latest

The compose file beside it adds workers and resource ceilings but nothing about access:

yaml
services:
  cloudflare-bypass:
    image: ghcr.io/sarperavci/cloudflarebypassforscraping:latest
    container_name: cloudflare-bypass
    ports:
      - "8000:8000"
    environment:
      - WORKERS=3
    deploy:
      resources:
        limits:
          memory: 4G
          cpus: '2'
    shm_size: 2g
    restart: unless-stopped 

So the scope is whatever the caller says it is. There is no step where you declare a site you own and the service refuses the rest, which is the single control that would turn this from an open bypass into a tool with a perimeter.

Four vendor blocks with referral codes take up more of the page than the documentation

The front page opens with four sponsored sections before any explanation of the software, each one a proxy vendor with a referral parameter in the link and, in two cases, a discount code. They advertise monthly plans, uptime figures, address counts by country, concurrent connection limits and rotation per request, and one of them offers a game on its landing page. None of that belongs to the project and none of it is verified by it. What it tells you is the funding model: the page is monetised through referral links to the same scraping infrastructure the tool depends on. That matters when you read the rest of the page, because the claims in those blocks are the vendor's claims, not the project's, and the project's own summary of itself is one sentence about bypassing protection with cookie generation and request mirroring.

Version 2.0 is announced, and there is no release and no version to pin

The first line announces Version 2.0 with enhanced request mirroring, improved caching and better reliability, and thanks readers for a stated number of stars. Behind that, the repository has no releases at all, and the manifest takes its version dynamically with no source for it declared in the visible configuration. The classifiers call the project Beta. So there is no artifact to pin, no changelog to read, and no way to tell from the outside whether the container you are running today is the one described in the announcement. The image reference in both the quick start and the compose file ends in a floating latest tag, which means every pull is whatever was built last.

Two layers do the work: a real browser and a fingerprint client

The design uses both a browser and a plain HTTP client, for different halves of the problem. The browser is a patched stealth Chromium, obtained through a dependency called cloakbrowser, which navigates to the named URL and lets the challenge resolve itself for non-interactive cases. Interactive checkbox challenges are handled by locating the widget inside its shadow DOM and clicking it natively, which is why the image runs a headed browser rather than a headless one. Replay is the second layer: once clearance exists, the request is reproduced to the origin through an HTTP client that impersonates Chrome's TLS handshake fingerprint, and the page names the dependency for it. So the tool is not one trick but a browser for the challenge and an impersonating client for the request, which is why the dependency list is longer than a browser automation project would suggest.

The clearance cookie is bound to the user agent that earned it

The response to a cookie request carries two values, not one: the clearance cookie itself and the exact user agent string that obtained it, in the same payload. The page is explicit that both must be sent together or the cookie is rejected, and it caches the pair so a repeat lookup is instant. That coupling is the operational detail most likely to bite a scraper written by hand: a cookie lifted into a different client, or paired with a different user agent, or kept past whatever validity window the configuration sets, stops being a credential. The cookie endpoint exists for the case where you drive your own client, and the mirror endpoint exists for the case where you would rather not, which is the difference between the two modes the project offers.

The browser binary is fetched at build time and frozen there

The image sets an environment variable that turns off automatic updates for the browser, then downloads the patched Chromium into the ubuntu user's cache as a build step and prints the path it resolved to. The effect is that the browser is pinned by whatever the build machine fetched that day, and a running container will not pick up a newer one. Two other details come from the same file: the process drops to an unprivileged user after the application directory is handed over, and the base image is Ubuntu rolling, so the operating system packages underneath are not pinned either. The shared memory size in the compose file is set to two gigabytes, which is a browser requirement rather than a service one.

Two dependency lists, and the image installs the other one

The project is both a Python package and a container image, and it declares its dependencies twice. The package metadata lists the browser wrapper, the fingerprinting HTTP client, a web framework, an ASGI server, a validation library and a virtual display helper, with Python 3.12 or later required and only 3.12 and 3.13 classified. The image does not install from that metadata. It installs from a separate requirements file, and the package configuration is set to include only the bypass package directory, with the server module sitting beside it at the top level of the repository. Two lists that nothing cross-checks is a slow drift, and a reader who reads only the metadata will not know which set the running image was built from.

The two example projects show the intended use

The page lists two projects built on this one: an automated book downloader for Calibre Web, and an unofficial API for a streaming service. Neither is described further, and neither has documentation on this page, so nothing can be said about how they work. What the listing does establish is direction of travel. Both take content from behind a protection and republish or redistribute it, which is a different activity from testing a site you administer, and it is the use the surrounding sponsor blocks are selling capacity for. The contribution section is one line, welcoming pull requests against the main codebase, and the licence is MIT.

Editorial conclusion

Do not expose this service to anything but your own machine. The project is a bypass for a protection the target's operator deliberately deployed, and the front page offers no per-target authorisation, no allowlist, no dry run and no authentication, while its own quick start publishes the port to the host. Anything that can reach it can name any site and have a stealth browser clear that site's challenge on its behalf, which means your infrastructure can end up acting as the requesting client for traffic you have no relationship with and no permission from. If you need this class of tooling for work you are authorised to do, the defensible shape is a single host on loopback or behind an authenticating reverse proxy, one target at a time, with the request log kept. Read two documentation facts before anything else: the clearance cookie is only valid alongside the exact user agent that earned it, so a mismatch is rejected, and the version announced on the page has no release behind it, so you are tracking a floating image tag. The front page is also mostly vendor advertising with referral codes, which tells you how the project is funded and nothing about how it works. Nothing here should be read as a method: the mechanism is described so you can assess the risk, not so you can point it at somebody else's site.

Frequently asked questions

What does CloudflareBypassForScraping actually do?

It runs as a self-hosted HTTP service that fronts a patched stealth Chromium. You name a host, a real browser visits that host and clears the challenge, and the service then either returns the clearance cookie with the user agent that earned it, or replays your request to the origin through an HTTP client that impersonates a Chrome TLS fingerprint.

Does the CloudflareBypassForScraping server require authentication?

Nothing on the front page describes one. The quick start publishes port 8000 on the host, and the compose file configures workers, a memory limit, a cpu limit and a shared memory size but no credential of any kind, so access control is left to whatever you put in front of the service.

Can I restrict CloudflareBypassForScraping to sites I own?

Not through anything the page describes. The target host is supplied per request in an x-hostname header, and there is no documented allowlist, no per-target authorisation and no dry run. Restricting it means putting the service behind your own reverse proxy with authentication, and only ever pointing it at a host you administer.

Why do the cookie and the user agent have to be sent together?

The clearance cookie is only valid for the user agent that earned it, so the service returns both values together and the page says Cloudflare rejects the cookie if they are separated. The pair is cached, which is why a repeated lookup for the same host returns immediately.

How does CloudflareBypassForScraping get past the JavaScript challenge?

With a patched stealth Chromium rather than by parsing responses. The browser navigates to the named URL and non-interactive challenges resolve on their own; interactive checkbox challenges are found inside their shadow DOM and clicked natively. The image therefore runs the browser headed under a virtual display, and the browser binary is downloaded during the image build with automatic updates disabled.

Official sources

  1. Issues
  2. License: MIT
  3. README
  4. sarperavci/CloudflareBypassForScraping on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/sarperavci-cloudflarebypassforscraping.svg)](https://hysenlabs.com/projects/sarperavci-cloudflarebypassforscraping)