crawl4ai vs scrapy: Markdown for LLMs or structured data pipelines
Crawl4AI is a browser-driven crawler that turns pages into LLM-ready Markdown, while Scrapy is a mature framework for extracting structured data at scale. They solve adjacent problems and can be combined, but most teams should pick one as the primary tool.
At a glance
| Project | unclecode/crawl4ai | scrapy/scrapy |
|---|---|---|
| Licence | Apache-2.0Permissive: commercial use allowed | BSD-3-ClausePermissive: commercial use allowed |
| Maintenance | Commits in the last six monthsLast push September 25, 2026 | Commits in the last six monthsLast push September 25, 2026 |
| Language | Python | Python |
| GitHub stars | 84,474 | 64,483 |
| Read more | Our analysisGitHub | Our analysisGitHub |
Which one to choose
Choose crawl4ai if you need clean Markdown from JavaScript-heavy pages for RAG, agents, or LLM extraction, and you want a self-hosted crawler with browser control and a CLI.
Choose scrapy if you need a maintainable, multi-site structured data pipeline with built-in request handling, selectors, and export formats, and you do not need browser rendering by default.
Different jobs: Markdown for models versus structured records
The core difference is the output and the runtime. Crawl4AI describes itself as turning web pages into clean Markdown for retrieval systems, agents, and data pipelines, with browser control and structured extraction. Its README lists LLM-ready output with headings, tables, code, and citation hints, plus an async browser pool and caching. In practice this means it drives a real browser through Playwright, which the README confirms in its manual install step (python -m playwright install --with-deps chromium). That browser dependency is what lets it render JavaScript-heavy pages and emit Markdown that a model can consume.
Scrapy is a framework, not a Markdown converter. Its README states it is a web scraping framework to extract structured data from websites, cross-platform, requiring Python 3.10+, maintained by Zyte and many other contributors. The mechanism is spiders plus selectors plus pipelines: you define how to parse a response, and Scrapy handles requests, scheduling, retries, and export. It does not ship a browser by default, so pages that require JavaScript rendering need an extra integration that the README does not document.
So the two are complementary rather than direct competitors. Crawl4AI optimises for feeding text to models. Scrapy optimises for repeatable extraction of fields into a database or file. If your goal is a knowledge base, the first is closer. If your goal is a price table or catalogue, the second is closer.
Getting each one running
Crawl4AI is installed as a package and then set up. The README gives pip install -U crawl4ai, then crawl4ai-setup, then crawl4ai-doctor to verify the installation. Browser issues are handled with a manual Playwright install. A minimal crawl is a few lines of async Python using AsyncWebCrawler, and the README also documents a CLI called crwl with flags such as -o markdown, --deep-crawl bfs, and -q for LLM extraction. That is a low-friction start for a single script.
The catch is the browser layer. Because Crawl4AI depends on Playwright and a Chromium build, your environment needs the right system libraries, and the README notes that crawl4ai-doctor exists precisely because browser-related issues are common enough to need a check. Docker is offered as an alternative deployment, and the release notes for v0.9.0 describe a secure-by-default Docker API server where auth is on by default and the server binds loopback unless given a token.
Scrapy starts differently. The README says install with pip install scrapy and follow the documentation. There is no setup command and no browser install, because the default runtime is HTTP. The cost is structural: you create a project, define spiders, and write parsing logic with selectors. That structure is what makes multi-site work maintainable, but it is heavier than a single async function. For a one-page fetch, Scrapy's own analysis says requests or httpx are lighter.
Operations, scaling, and long-running crawls
Crawl4AI's recent releases show where its operational attention has gone. v0.8.0 introduced deep crawl crash recovery with resume_state and on_state_change callbacks for long-running crawls, plus a prefetch mode the notes describe as faster URL discovery. v0.9.2 is a maintenance patch that fixes a MemoryAdaptiveDispatcher task and page leak when a streaming crawl is closed, along with Docker Playground config, Monitor WebSocket auth, Playwright headless-shell packaging, and GPU Docker builds. The pattern is a fast-moving project that adds scaling features and then patches the edges of them.
Scrapy's scaling story is older and more conventional. It is a framework with built-in request handling, and its analysis points to export pipelines as a reason to adopt it. Scheduling, concurrency, and retries are part of the framework rather than something you assemble. For many sites with stable HTML, that is a smaller operational surface than running a fleet of browsers.
The trade-off is clear. Crawl4AI gives you browser rendering and Markdown but asks you to manage Playwright, possibly Docker, and a rapid release cycle. Scrapy gives you a predictable HTTP pipeline but you must add rendering yourself if a target needs it. Neither README documents a fully managed hosting option; Crawl4AI's README mentions a closed beta cloud API, described as launching soon and onboarding in phases, so it is not something to plan production around today.
Where each one falls short
Crawl4AI's weakness is security and operational care. Its v0.8.7 release notes describe a security-hardening release fixing Docker API vulnerabilities including RCE, SSRF, auth bypass, file write, XSS, and a hardcoded JWT secret. v0.9.0 then made the Docker API server secure by default. That history is not disqualifying, but it means self-hosting requires attention: verify your Docker deployment binds to loopback, uses a strong token, and applies the v0.9.0 secure-by-default settings, and confirm the v0.8.7 and v0.9.0 fixes are in your version. The README does not document rollback for a bad upgrade, and the release cadence is fast, so pinning versions and testing upgrades is part of the job.
Scrapy's weakness is the opposite: it is conservative. The README says it requires Python 3.10+, so older runtimes are out. Its analysis notes a real learning curve and project structure cost, and that it is not the right tool for a single page fetch. It also does not render JavaScript by default, so modern single-page applications need additional components that the README does not describe. If your targets are mostly client-rendered, Scrapy alone will return empty or partial responses.
Both are maintained. Crawl4AI's last push is 2026-09-16 and Scrapy's is 2026-09-14, both recent, and neither repository is archived. Crawl4AI's release notes show a faster cadence, with v0.9.0 in June 2026 and v0.9.2 in July 2026. Scrapy's releases are also regular, with 2.16.0 in May 2026, 2.17.0 in July 2026, and 2.18.0 in August 2026.
Licence and maintenance implications
Crawl4AI is Apache-2.0. Scrapy is BSD-3-Clause. Both are permissive and allow commercial use, modification, and redistribution. The practical difference is not the licence text but the dependency and governance picture. Crawl4AI is a single-maintainer project with a sponsorship programme and a commercial cloud beta in the README. Scrapy is maintained by Zyte and many other contributors, which spreads continuity risk across an organisation and a contributor base.
For a company with strict open-source review, both licences are usually acceptable, but Apache-2.0 includes an explicit patent grant while BSD-3-Clause does not. That matters to some legal teams and not to others. The README does not state any additional terms for either project beyond the licence.
Maintenance signals should be read from the last push and release history rather than popularity. Crawl4AI shipped v0.9.2 on 2026-07-15, v0.9.1 on 2026-07-08, and v0.9.0 on 2026-06-18. Scrapy shipped 2.18.0 on 2026-08-20, 2.17.0 on 2026-07-07, and 2.16.0 on 2026-05-19. Both are current. The difference is that Crawl4AI's recent releases are dominated by security and Docker fixes, while Scrapy's are routine framework releases. If your organisation cannot absorb frequent security-driven upgrades, that is a real factor against Crawl4AI for self-hosted Docker use.
Choosing for concrete scenarios
For a RAG pipeline that ingests documentation, news, or product pages into a vector store, Crawl4AI is the closer fit. Its Markdown output with headings, tables, and citation hints is designed for that, and its CLI (crwl with -o markdown) makes one-off ingestion easy. Pair it with version pinning and the v0.9.0 Docker settings if you deploy the API server.
For a recurring scrape of a catalogue across several sites, with fields that must land in a database, Scrapy is the closer fit. Spiders, selectors, and pipelines give you a structure that survives staff changes, and the framework handles requests and exports. If those sites render client-side, you will need to add a rendering component, which the README does not cover.
For a hybrid, run Scrapy as the scheduler and pipeline layer and call Crawl4AI for pages that need a browser and Markdown. That combination is not documented as an official integration in either README, so treat it as custom glue you own. For a single page fetch or a quick prototype, neither is necessary; Scrapy's own analysis points to requests or httpx as lighter. If you cannot keep up with security patches, Crawl4AI's self-hosted Docker path is the wrong choice regardless of features. If you need JavaScript rendering and only want one dependency, Crawl4AI is the more direct route; if you need long-term stability and structured exports, Scrapy is.
Bottom line
Pick Crawl4AI when the deliverable is LLM-ready Markdown and you can operate a browser-based stack with disciplined version pinning. Pick Scrapy when the deliverable is structured records from multiple sites and you want a framework with built-in request handling and exports. Before committing, verify two things: for Crawl4AI, that your deployment applies the v0.9.0 secure-by-default Docker settings and includes the v0.8.7 security fixes; for Scrapy, that your Python runtime is 3.10+ and that any JavaScript-rendered targets have a rendering component you are prepared to maintain.