CLI tool
D4Vinci/Scrapling avatar
D4Vinci/Scrapling

Scrapling: An Adaptive Python Web Scraping Framework with Anti-Bot Bypass

An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl!

84,037 stars8,596 forksPythonBSD-3-Clause

At a glance

What is it?
Scrapling is a Python web scraping library that handles everything from single HTTP requests to full-scale concurrent crawls. Its adaptive parser remembers element locations and relocates them when a page's structure changes, while its StealthyFetcher bypasses protections like Cloudflare Turnstile without additional configuration.
Who is it for?
Scrapling is worth evaluating for Python scraping projects where CSS selectors break frequently as target sites update their layouts, or where Cloudflare Turnstile blocks naive HTTP clients. The adaptive selector feature makes the most sense for long-running scrapers that need to survive site redesigns without manual intervention.
Can I use it commercially?
Yes. BSD-3-Clause is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 2 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 27, 2026, and from our analysis. They are not legal advice.

DEEP OPEN-SOURCE ANALYSIS

What Scrapling Solves: Layout Changes and Bot Detection

Two problems break most web scrapers in production. The first is structural: websites change their CSS classes, IDs, or DOM structure, and selectors that worked yesterday return nothing today. The second is access: websites behind Cloudflare Turnstile or other bot detection systems return challenge pages to automated HTTP clients. Scrapling addresses both. The adaptive parser stores information about how elements were located during an initial scrape and uses that memory to find the same elements again even after the page structure has changed, by matching on content patterns rather than exact CSS paths. The StealthyFetcher launches a real Chromium browser in a configuration designed to avoid detection. The library is aimed at developers building production scrapers that need to survive site updates without constant selector maintenance.

Adaptive CSS Selectors: How the Memory System Works

The adaptive feature uses two special parameters on CSS selection calls. Setting auto_save=True on a selection tells Scrapling to record the element's location in a database keyed to the URL and selector. On a later scrape of the same URL, passing adaptive=True on the same selector tells Scrapling to use the stored location data to find the element even if the CSS path has changed. From the README:

python
from scrapling.fetchers import Fetcher, AsyncFetcher, StealthyFetcher, DynamicFetcher
StealthyFetcher.adaptive = True
p = StealthyFetcher.fetch('https://example.com', headless=True, network_idle=True)
products = p.css('.product', auto_save=True)
products = p.css('.product', adaptive=True)

The auto_save call creates the memory. The adaptive call uses it on the next run. The pyproject.toml does not specify which database backs this storage, but the feature is described in the README as surviving website design changes.

Installing Scrapling and Running a First Scrape

Install from PyPI using the package name scrapling:

bash
pip install scrapling

Scrapling requires Python 3.10 or higher, as specified in pyproject.toml. The package is version 0.4.15. For full browser support through the DynamicFetcher or StealthyFetcher, Playwright and Chromium are required. The Dockerfile in the repository installs them automatically:

bash
uv run playwright install-deps chromium
uv run playwright install chromium

The pyproject.toml lists all dependencies. Key ones include playwright (via cdp-use==1.4.5 and browser-harness==0.1.13), pydantic==2.13.5 for data validation, and markdownify==1.2.2 for HTML-to-text conversion. A Docker image is available at pyd4vinci/scrapling on Docker Hub.

The Spider Framework for Full-Scale Crawls

Beyond single-page fetching, Scrapling provides a Spider class for concurrent, multi-session crawls with pause and resume support, automatic proxy rotation, and adaptive speed control. From the README:

python
from scrapling.spiders import Spider, Response

class MySpider(Spider):
  name = "demo"
  start_urls = ["https://example.com/"]

  async def parse(self, response: Response):
      for item in response.css('.product'):
          yield {"title": item.css('h2::text').get()}

MySpider().start()

The Spider class handles concurrency and session management. The crawl speed adapts to how fast each target website responds and backs off when the site starts blocking requests, according to the README. The proxy rotation integrates with external proxy providers; the README lists sponsor providers that work with Scrapling.

The MCP Server and CLI Integration

Scrapling ships an MCP (Model Context Protocol) server, exposing scraping capabilities as tools that AI agents can call directly. The Dockerfile exposes port 8000 for the MCP server HTTP transport. An agent-skill directory in the repository provides integration with coding assistants via the Agent Skills specification. The repository also includes a CLI, documented at scrapling.readthedocs.io/en/latest/cli/overview.html. The package.json in pyproject.toml lists server.json for MCP server metadata and zensical.toml for additional configuration. This positions Scrapling as a scraping backend that AI agents can delegate tasks to, rather than only a Python library for direct use.

Where Scrapling Is Not the Best Choice

Scrapling's adaptive parser and stealthy browser add overhead that simple scraping tasks do not need. For websites that return complete HTML without JavaScript rendering and do not use bot detection, the combination of requests and BeautifulSoup is faster, lighter, and easier to debug. The StealthyFetcher launches a real Chromium instance, which is memory-intensive and slower per request than an HTTP client. Scrapling's beta classification in pyproject.toml (Development Status 4 - Beta) signals that the API may change between releases. The BSD-3-Clause license permits commercial use but differs from MIT in attribution requirements. The ROADMAP.md in the repository documents planned features, indicating the project is still evolving.

Scrapling versus Scrapy and Playwright

Scrapy is a mature Python crawling framework built for high-throughput, distributed scraping with a middleware pipeline architecture. Scrapy provides no adaptive selectors and no built-in anti-bot bypass, but it is well-documented, production-tested, and has a large middleware ecosystem. Scrapling is more recent and adds adaptive selectors and StealthyFetcher as first-class features, but with a smaller ecosystem and a beta-stage API. Playwright is a browser automation library that controls Chromium, Firefox, or WebKit for end-to-end testing and scraping. It gives precise control over every browser interaction but does not provide a spider framework or adaptive selectors. Scrapling's DynamicFetcher wraps Playwright internally, adding the adaptive layer and crawling orchestration on top of Playwright's browser control.

Editorial conclusion

Scrapling is worth evaluating for Python scraping projects where CSS selectors break frequently as target sites update their layouts, or where Cloudflare Turnstile blocks naive HTTP clients. The adaptive selector feature makes the most sense for long-running scrapers that need to survive site redesigns without manual intervention. Engineers who need a fully deterministic script with no LLM costs and no browser launch overhead should use Scrapy or requests with BeautifulSoup instead. The library requires Python 3.10 or higher. The last push was on 2026-09-27 and the current version is 0.4.15.

Frequently asked questions

What is Scrapling?

Scrapling is a Python web scraping framework that combines adaptive CSS selectors, anti-bot bypass through a stealthy Chromium browser, a Spider class for full crawls with proxy rotation, and an MCP server for AI agent integration. It requires Python 3.10 or higher.

How do I install Scrapling?

Run pip install scrapling. For features that require a real browser (StealthyFetcher, DynamicFetcher), install Playwright and Chromium separately using uv run playwright install chromium, or use the Docker image pyd4vinci/scrapling.

How does Scrapling compare to Scrapy?

Scrapy is a mature, high-throughput crawling framework with a large middleware ecosystem and no built-in adaptive selectors or anti-bot bypass. Scrapling is newer, adds adaptive element location and a stealthy browser fetcher, but has a smaller ecosystem and a beta-stage API. Both support concurrent crawls; the choice depends on whether adaptive selectors or stealth fetching are requirements.

How do I use Scrapling?

Import StealthyFetcher or Fetcher from scrapling.fetchers, call the fetch method with a URL, then select elements using CSS selectors on the response. Use auto_save=True on the first run to store element locations, and adaptive=True on subsequent runs to relocate elements even if the page structure has changed.

Official sources

  1. Official documentation
  2. Official README
  3. Project repository
  4. Release notes
For maintainers

Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/d4vinci-scrapling.svg)](https://hysenlabs.com/projects/d4vinci-scrapling)
Community notes

Community notes