Crawlee for Python: a crawler framework with HTTP and browser engines behind one API
Crawlee—A web scraping and browser automation library for Python to build reliable crawlers. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Parsel, BeautifulSoup, Playwright, and raw HTTP. Both headful and headless mode. With proxy rotation.
At a glance
- What is it?
- Crawlee for Python wraps request queues, storage and two crawling engines (Impit HTTP and Playwright) behind a single handler API. It is Apache-2.0, requires Python 3.10 or newer, and its last push was on 2026-09-10.
- Who is it for?
- Adopt Crawlee for Python if you want one handler API over both HTTP and browser crawling, with a request queue and dataset included, and you can live with a young library whose docs are still thinner than Scrapy's. Do not adopt it if you need mature distributed scheduling or a large third-party extension ecosystem.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository received new commits within the last day.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem Crawlee for Python solves for scraper authors
A hand-written scraper usually starts as a loop over URLs with requests and BeautifulSoup. It ends up needing a visited set, retry logic, concurrency limits, a place to put extracted records, and a second implementation for pages that only render in a browser. Crawlee for Python packages those parts: a request queue, a dataset, retry and concurrency handling, and two crawler classes that share one handler signature. The audience is Python developers building crawlers for data extraction, including pipelines that feed LLMs, RAG indexes or GPT-style applications, as the repository description puts it. The library targets Python 3.10 and newer and ships as the crawlee package on PyPI. It is not a hosted service. It is a library you import, and the README notes that a crawler run creates a storage/ directory in your current working directory.
BeautifulSoupCrawler and PlaywrightCrawler: one handler, two engines
The architecture splits along how a page is fetched. BeautifulSoupCrawler downloads pages with an HTTP library and hands you parsed HTML. By default it uses ImpitHttpClient for HTTP communication and BeautifulSoup for parsing, per the README. It does not run a browser, so it is the cheaper path, but client-side JavaScript will not execute. PlaywrightCrawler drives a headless browser through Playwright and gives you an API for data extraction on pages that need rendering. Both expose the same shape: you construct a crawler, register a handler on crawler.router.default_handler, and call await crawler.run([...]) with seed URLs. Inside the handler, context.request.url is the current URL, context.push_data writes a record into the default dataset, and context.enqueue_links follows links found on the page. Because the handler contract is identical, moving a site from the HTTP crawler to the browser crawler is mostly a class swap. The cost is that the browser path pulls in Playwright and its browser binaries, which the HTTP path does not.
Installing Crawlee for Python and running a first crawl
The README recommends installing everything with the all extra, then installing Playwright's browser dependencies separately. The two commands are distinct because the Python package and the browser binaries ship separately.
python -m pip install 'crawlee[all]'
playwright installTo confirm the install, the README gives a one-liner that prints the version. If it prints a version string, the package is importable from the active interpreter.
python -c 'import crawlee; print(crawlee.__version__)'There is also a CLI path. With uv installed, uvx runs the crawlee package with the cli extra and scaffolds a project from a template; if crawlee is already installed, the crawlee create command does the same.
uvx 'crawlee[cli]' create my-crawlerThe README's BeautifulSoupCrawler example is the shortest working crawl. It caps the run at ten requests, extracts the page title, pushes a record, and enqueues every link it finds. Set max_requests_per_crawl to a small number on the first run: without it, enqueue_links will keep following links.
import asyncio
from crawlee.crawlers import BeautifulSoupCrawler, BeautifulSoupCrawlingContext
async def main() -> None:
crawler = BeautifulSoupCrawler(max_requests_per_crawl=10)
@crawler.router.default_handler
async def request_handler(context: BeautifulSoupCrawlingContext) -> None:
context.log.info(f'Processing {context.request.url} ...')
data = {
'url': context.request.url,
'title': context.soup.title.string if context.soup.title else None,
}
await context.push_data(data)
await context.enqueue_links()
await crawler.run(['https://crawlee.dev'])
if __name__ == '__main__':
asyncio.run(main())After the run, look in storage/ in the working directory for the dataset written by push_data. The README notes that BeautifulSoupCrawler requires the beautifulsoup extra, so a bare pip install crawlee is not enough for that example.
Where Crawlee for Python is the wrong tool
The extras system is the first sharp edge. The base crawlee package carries core functionality only, and features arrive as optional extras to keep dependencies and package size down. Pick the wrong extra and the import fails at runtime, not at install time. The README states plainly that BeautifulSoupCrawler needs the beautifulsoup extra, so installing the base package and copying the example will not work. Second, PlaywrightCrawler is not a drop-in for every site: it carries a browser, which means memory and startup cost per run that the HTTP crawler avoids. If a page returns its content in the initial HTML, using the browser crawler is wasted work. Third, the README does not document rollback behaviour for the storage/ directory, so if you care about what happens to partially written datasets when a run is interrupted, that is something to check in the source or the website docs rather than the README. Finally, the version in pyproject.toml is 1.10.1 while the most recent release listed is v1.10.0, which is normal for a repository between a release and the next tag, but it means the version string you get from pip may not match the newest changelog entry.
Crawlee for Python versus Scrapy, and versus the JS implementation
Scrapy is the comparison most Python developers will make. Scrapy's model is a spider class with parse callbacks and its own engine; Crawlee's model is a crawler object with a router and a handler, and the same handler runs whether the fetch came from Impit or from Playwright. That is the practical difference: in Scrapy, JavaScript rendering is an add-on (scrapy-playwright) rather than one of two first-class engines. Scrapy has the longer track record and a wider set of third-party extensions; Crawlee for Python is younger and its README points to the project website for full documentation, guides and examples. The other alternative is Crawlee for JS/TS, which the README calls a TypeScript implementation of Crawlee and links to on GitHub. If your team is already on Node, the TS version is the one to look at; the Python package is a separate codebase with its own release cadence, so bug fixes do not move between them in lockstep. Crawl4AI is another tool people compare against, but the repository does not describe it, so the honest comparison stops at Scrapy and at the TS sibling.
Maintenance, licence and upgrade cost
The repository is not archived, and the last push was on 2026-09-10, days before the most recent release tag in the list, v1.10.0 on 2026-08-31. Releases have been frequent: v1.9.2 on 2026-08-17, v1.9.3 on 2026-08-24, v1.10.0 on 2026-08-31. That cadence cuts both ways. Fixes arrive quickly, but a minor bump can change behaviour, so pin the version in your requirements and read CHANGELOG.md before upgrading. The project is licensed Apache-2.0, with the licence file at the repository root and the same identifier declared in pyproject.toml; that is a permissive licence, and the usual obligations around notices and attribution apply. This is not legal advice, and if you redistribute the library inside a product, have someone check the LICENSE file rather than the classifier list. Dependency weight is the other ongoing cost: the all extra pulls in Playwright, scikit-learn, jaro-winkler and database drivers for the optional SQL and Redis backends, so installing everything to avoid import errors also installs a lot you will never call. Install the narrow extras you actually use.
Editorial conclusion
Adopt Crawlee for Python if you want one handler API over both HTTP and browser crawling, with a request queue and dataset included, and you can live with a young library whose docs are still thinner than Scrapy's. Do not adopt it if you need mature distributed scheduling or a large third-party extension ecosystem. Before committing, verify the extras you install (crawlee[beautifulsoup] versus crawlee[all]), run playwright install separately if you use the browser crawler, and confirm where the storage/ directory will live, because a crawler run creates it in the current working directory.
Frequently asked questions
Is Crawlee for Python free to use?
Yes. The package is published on PyPI as crawlee and the repository is licensed Apache-2.0, with the licence file at the repository root.
What are the key differences between Crawlee for Python and Scrapy?
Crawlee for Python puts an HTTP crawler and a Playwright browser crawler behind the same router and handler API, so switching engines is largely a class swap. Scrapy is not described in the repository, so a detailed feature comparison is not something the README documents.
What is Crawlee for Python?
It is a web scraping and browser automation library for Python, published as the crawlee package on PyPI, with BeautifulSoupCrawler for HTTP fetching and PlaywrightCrawler for headless browser rendering.
What is the difference between crawling and scraping?
The README does not define the two terms separately. What it shows is that Crawlee handles both halves in one run: enqueue_links follows links from a page, and push_data stores the extracted record in the default dataset.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/apify-crawlee-python)