# Botasaurus: Python Web Scraping Framework with Built-In Bot Detection Bypass

> Botasaurus is a Python web scraping framework that wraps Chrome automation and HTTP requests behind two decorators, @browser and @request, while handling anti-detection, parallelization, caching, and proxy management as built-in concerns. The README documents bypass of Cloudflare WAF, BrowserScan, Fingerprint, and Datadome detection systems.

**omkarcloud/botasaurus** — The All in One Framework to Build Undefeatable Scrapers

- Repository: https://github.com/omkarcloud/botasaurus
- Website: https://www.omkar.cloud/botasaurus/
- Stars: 5,747 · Forks: 506
- Language: Python
- License: MIT
- Published: 2026-09-22 · Updated: 2026-09-22 · Language: en
- Canonical page: https://hysenlabs.com/projects/omkarcloud-botasaurus

## What Problem Botasaurus Solves and Who It Is For

Building a web scraper that actually works against modern bot detection requires assembling multiple tools: a browser driver, a human-cursor simulation library, a proxy rotator, a request library that mimics browser TLS fingerprints, and a caching layer to avoid redundant fetches. Botasaurus packages all of those into a single Python framework so engineers can focus on the extraction logic rather than the plumbing.

The primary audience is Python developers who are building scrapers against sites that actively detect bots. The README describes passing Cloudflare Web Application Firewall, BrowserScan Bot Detection, Fingerprint Bot Detection, Datadome Bot Detection, and Cloudflare Turnstile CAPTCHA as verified outcomes. It also targets developers who want to distribute their scrapers beyond technical users, since the framework can convert a scraper into a desktop application for macOS, Windows, and Linux or into a web application.

## The @browser and @request Decorator Model

Botasaurus centers on two Python decorators that transform ordinary functions into scraping tasks. The `@browser` decorator provides a Chrome driver that the README calls a "Humane Driver," which simulates realistic mouse movements and behaves like a human-operated browser.

```python
from botasaurus.browser import browser, Driver

@browser
def scrape_heading_task(driver: Driver, data):
    driver.get("https://www.omkar.cloud/")
    heading = driver.get_text("h1")
    return {"heading": heading}

scrape_heading_task()
```

The framework automatically saves the returned dictionary to `output/scrape_heading_task.json`. The function name becomes the output filename, which provides a consistent convention across a project. The driver receives all human-like behavior configuration automatically; the developer writes only the site-specific logic.

The `@request` decorator provides an HTTP client that makes browser-like requests without launching a full browser:

```python
from botasaurus.request import request, Request
from botasaurus.soupify import soupify

@request
def scrape_heading_task(request: Request, data):
    response = request.get("https://www.omkar.cloud/")
    soup = soupify(response)
    heading = soup.find('h1').get_text()
    return {"heading": heading}

scrape_heading_task(
```

The `soupify()` helper converts the response into a BeautifulSoup object directly. The README describes this mode as making browser-like humane HTTP requests that produce the same output as the full browser mode while consuming fewer resources.

## Installing Botasaurus and Running a First Scraper

Installation requires Python and pip:

```shell
python -m pip install --upgrade botasaurus
```

The setup.py in the repository lists the runtime dependencies: psutil, javascript_fixes, requests, joblib 1.3.2 or later, beautifulsoup4 4.11.2 or later, lxml 5.2.2 or later, openpyxl, close_chrome, botasaurus-humancursor, botasaurus-api, botasaurus-driver, bota, botasaurus-proxy-authentication, and botasaurus-requests. These are installed automatically by pip.

After installing, create a directory, write a Python script with a decorated function, and run it:

```shell
mkdir my-botasaurus-project
cd my-botasaurus-project
```

The README's first example shows that after execution, the script launches Chrome, visits the target URL, extracts the heading, and saves the result to `output/scrape_heading_task.json` automatically. The output file path is derived from the function name, so naming functions descriptively produces a self-organizing output directory.

## Proxy Cost Reduction and Parallel Execution

The README claims up to 97 percent savings on browser proxy costs by using browser-based fetch requests instead of routing full Chrome sessions through a proxy for every page. The mechanism is that the `@request` decorator sends HTTP requests with browser-like headers and TLS fingerprints, which passes basic detection checks without the overhead of maintaining a full browser session through the proxy tunnel.

Parallelization is built in. The README describes passing a list of data items to a decorated function to distribute scraping across multiple workers. Configuration for profiles, browser extensions, and proxy rotation is handled at the decorator level rather than requiring manual thread management. The framework also includes a caching layer that stores results so re-running a scraper skips pages already fetched, which matters during development when the scraping logic is being refined.

For larger deployments, the repository includes a run-scraper-in-kubernetes.md document describing how to scale across multiple machines. The bot_detection_tests.py file at the repository root contains the code used to verify bypass of the detection systems listed in the README.

## Converting a Scraper to a Desktop or Web Application

Botasaurus includes a pathway to package a scraper as a desktop application. The repository contains a botasaurus-desktop-tutorial.md, and the README describes converting a scraper into a desktop app for macOS, Windows, and Linux in what it estimates as one day of work. This makes the scraper available to non-technical users who cannot run Python scripts.

The web application conversion turns the scraper into a site-hosted interface where customers can submit URLs or parameters and retrieve results. The botasaurus_server/ directory in the repository contains the server-side components for this path. The botasaurus-controls/ directory provides the UI control components.

These two packaging options are the main reason Botasaurus describes itself as an all-in-one framework rather than a scraping library. A pure scraping library ends at data extraction; Botasaurus extends to distribution of the resulting tool.

## Where Botasaurus Is the Wrong Tool

Botasaurus is not designed for scraping tasks that require very low latency at high volume without browser overhead. The human-like mouse movement simulation and Chrome automation that make it effective against bot detection also add overhead that a lightweight HTTP client does not have. If the target site has no bot detection and the goal is raw throughput, a simpler tool is more appropriate.

The framework is also not suited for teams operating under strict open-source license policies that prohibit MIT-licensed dependencies. While the MIT license is permissive, the dependency tree is large and includes multiple botasaurus-prefixed packages that all need review.

The project had its last push on 2026-07-26 and has no GitHub releases, meaning semantic versioning and release notes are not part of the current workflow. Teams who depend on documented releases for change management will need to track commits manually. Version 4.0.97 is the version shown in setup.py.

## Playwright as the Main Alternative

Microsoft's Playwright is the most comparable alternative for browser-based scraping. Like Botasaurus, it drives a real browser and can execute JavaScript. The difference is in focus: Playwright is a general-purpose browser automation library that does not ship anti-detection techniques, human cursor simulation, or parallel scraping orchestration as built-in features. Developers who use Playwright for scraping typically add separate libraries for fingerprint randomization and proxy rotation.

Botasaurus trades Playwright's flexibility and cross-language support (Playwright targets JavaScript, Python, Java, and .NET) for a more opinionated Python-only API where anti-detection and output management are already solved. Teams who are already invested in the Playwright ecosystem and prefer to compose their own anti-detection stack will likely stay with Playwright. Teams who want a single Python package that makes an opinionated tradeoff in favor of anti-detection from the start are the intended Botasaurus users.

## Conclusion

Engineers who need to scrape sites with aggressive bot detection and want a single framework to handle driver management, anti-detection, parallelism, and output formatting should evaluate Botasaurus. It is the wrong tool for scraping tasks that require rendering JavaScript on a strict budget without proxy costs, since full browser automation still has overhead despite the proxy savings. Teams who need Kubernetes-scale scraping should review the run-scraper-in-kubernetes.md documentation in the repository before committing to the architecture. The MIT license allows commercial use.

## FAQ

### Can bot detection identify Botasaurus scrapers?

The README documents bypassing Cloudflare WAF, BrowserScan, Fingerprint, and Datadome detection, and includes a bot_detection_tests.py file with the verification code. Detection systems continuously update, so any claim of permanent undetectability would be inaccurate; the README makes no such claim.

### What are the main dependencies Botasaurus requires?

The setup.py lists psutil, requests, joblib, beautifulsoup4, lxml, openpyxl, and several botasaurus-prefixed packages including botasaurus-driver, botasaurus-humancursor, and botasaurus-requests. All are installed automatically with pip install botasaurus.

### Does Botasaurus support scaling across multiple machines?

The README mentions Kubernetes-based scaling, and the repository includes a run-scraper-in-kubernetes.md file. The README also references PostgreSQL integration documents for cloud deployments, suggesting a distributed persistence layer is part of the scale-out architecture.

## Sources

- [Issues](https://github.com/omkarcloud/botasaurus/issues)
- [License: MIT](https://github.com/omkarcloud/botasaurus/blob/master/LICENSE)
- [omkarcloud/botasaurus on GitHub](https://github.com/omkarcloud/botasaurus)
- [Project website](https://www.omkar.cloud/botasaurus/)
- [README](https://github.com/omkarcloud/botasaurus/blob/master/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/omkarcloud-botasaurus
