Framework
scrapy/scrapy avatar
scrapy/scrapy

Scrapy: one install command, two async stacks, no way to make either optional

Scrapy, a fast high-level web crawling & scraping framework for Python.

64,483 stars11,973 forksPythonBSD-3-Clause

At a glance

What is it?
Scrapy is a BSD-3-Clause Python framework for extracting structured data from websites, maintained by Zyte and published on PyPI and conda-forge. Every dependency in pyproject.toml is mandatory, four of them swap depending on the interpreter, and the README documents one install command and nothing about how a crawl is configured.
Who is it for?
Adopt Scrapy when the job is a multi-page crawl with a queue, retries and item pipelines, and stay with a plain HTTP client plus a parser when you need one page. Check first that your interpreter and Python version resolve the dependency markers you expect, because pyproject.toml makes the installed set differ per Python, and read NEWS before scheduling an upgrade since the README never links it.
Can I use it commercially?
Yes. BSD-3-Clause is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 5 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 25, 2026, and from our analysis. They are not legal advice.

Editorial analysis

Twisted and aiohttp are both non-optional in pyproject.toml

The dependency list in pyproject.toml names Twisted>=21.7.0 and aiohttp>=3.13.3 side by side, and neither carries a marker that would let a resolver skip it. The markers that do exist are the implementation swaps listed further down, plus typing-extensions>=4.5.0, which applies only when python_version is below 3.13. So the install pulls two asynchronous downloader stacks into every environment, and a reader who wants only one of them has no supported way to drop the other. Editing a lock file to remove a package changes a set the project publishes as a unit, and the manifest gives you no alternative list to switch to. The TLS side is equally unconditional: pyOpenSSL>=24.3.0, service_identity>=24.2.0 and cryptography>=41.0.5 are all required rather than optional extras. The consequence for a reader is concrete. Certificate verification is not something you opt into, and you cannot assemble a trimmed Scrapy from published metadata. If the actual task is a single site and a couple of pages, this is the wrong size of tool, and no amount of install flags will make it smaller.

PyPy and Python 3.14 change which packages actually land

Four dependency entries are keyed to the interpreter or the interpreter version, not to the operating system. PyDispatcher>=2.0.5 applies when platform_python_implementation is CPython, and PyPyDispatcher>=2.1.0 applies when it is PyPy. brotli>=1.2.0 applies when implementation_name is not pypy, and brotlicffi>=1.2.0.0 covers the PyPy case. backports.zstd>=1.3.0 applies when python_version is below 3.14. The README calls the project cross-platform, which is true, and says nothing about PyPy or about which of these combinations is exercised. The result for a reader is that the phrase cross-platform hides four distinct dependency graphs behind one command. A team pinning a single requirement across a fleet that mixes 3.10 through 3.15 will resolve different packages per environment, and the published metadata will not tell you which combinations anyone has run. The test entry points are tox.ini plus the tests/ and tests_typing/ directories, so the supported matrix is discoverable in the repository, not in the prose.

One install command, and no first spider to run it against

The entire installation instruction on this project is a single line, and it is the only command the README offers:

bash
pip install scrapy

Two badges point to where the built artifact lives: the PyPI project page, and a conda-forge channel. A conda-based machine therefore has a second route that the pip line does not cover, and that badge does not spell out the channel's own install syntax, so you are reading the conda documentation for the exact command. After the install line, the README points to the documentation at https://docs.scrapy.org/en/latest/, and for contributors to the contributing page under the same host. What is missing is the part people actually search for. There is no example spider, no command line for running a crawl, and no settings file in the repository's top-level listing, which holds directories such as extras/, docs/, scrapy/, sep/, tests/ and tests_typing/ alongside files like pyproject.toml, tox.ini and conftest.py. The consequence for a new user is direct: the first successful run cannot be assembled from this page. You leave for the documentation before you see anything work, and nothing here gives you a minimal spider to check your install against.

protego and defusedxml are required, so the safety pins cannot be removed

Two entries in the manifest say more about the project's defaults than any prose does. protego>=0.1.15 is a robots.txt parser, and it is a required dependency rather than an extra, so robots handling belongs to the base install instead of something you switch on per project. defusedxml>=0.7.1 is there because a crawler parses XML from servers you do not control, and it is required too, which means XML entity expansion is not left to your judgment at install time. service_identity>=24.2.0 and tldextract sit in the same category: unconditional, and both about how requests are addressed and identified. Here is the catch for anyone reasoning about runtime behaviour. No published setting changes how protego is consulted, and no default settings file appears among the top-level entries, so the defaults ship inside the package rather than as a root-level file you can read from the tree. You can inspect the installed package or read the documentation, but you cannot diff a configuration file to see which user agent string gets sent or how a disallowed path is handled.

The parser is one layer inside a crawler, not a stand-in for one

Read the dependency list in the order the work happens. queuelib>=1.6.1 holds what is still to be fetched. w3lib>=2.1.1, with Twisted>=21.7.0 and aiohttp>=3.13.3 above it, moves the bytes. parsel>=1.8.1, lxml>=4.6.4 and cssselect>=1.2.0 select from the response body. itemloaders>=1.0.1 and itemadapter>=0.1.0 shape what comes out, and zope.interface>=5.1.0 and platformdirs>=2.0.0 carry the extension contract and per-user state. An HTML parsing library covers the third line only, which is why comparisons between Scrapy and a parser tend to miss the point: the parser here is a component, wrapped in a scheduler and a retry story. Building the rest by hand means choosing a queue, a retry policy, a concurrency limit and an item pipeline yourself, then keeping their versions compatible by hand. The genuinely useful thing in this manifest is that every part has a stated floor. Those floors are the closest published statement of what a hand-built equivalent has to match, and no configuration file in the repository restates them. The README does not describe the item pipeline or the scheduler at all.

dynamic version means pyproject.toml cannot tell you what is installed

The build system is hatchling>=1.27.0 with build-backend hatchling.build, and the project table sets name = Scrapy alongside dynamic = version. The manifest therefore deliberately holds no version string, so there is nothing to grep in pyproject.toml when you need to know which release a machine is carrying. Releases are 2.19.0 on 2026-09-10, 2.18.0 on 2026-08-20 and 2.17.0 on 2026-07-07, an interval measured in weeks rather than quarters. Two markers make that cadence resolve differently across interpreters: typing-extensions applies below 3.13 and backports.zstd below 3.14, so a 3.10 environment ends up with a different package set than a 3.14 one on the very same Scrapy release. The consequence for anyone planning an upgrade is that one pinned requirement across a mixed interpreter fleet is not one resolved set, and reproduction depends on the interpreter you happen to run. The name of the build backend and the shape of the project table also tell you that packaging is modern, which is a small signal but the only one the file gives about release engineering.

The change record is NEWS, and the README never links it

Maintenance is current. The last push to the master branch landed on 2026-09-25, five days after the 2.19.0 tag, and the project carries the classifier Development Status :: 5 - Production/Stable along with Intended Audience :: Developers, Environment :: Console and Operating System :: OS Independent. The licence is BSD-3-Clause. The in-tree change record is the NEWS file at the repository root, and the README does not link it, so anyone trying to find out what changed between two versions has to go looking instead of following a pointer. The contributor surface is well marked by contrast: tox.ini, conftest.py, .pre-commit-config.yaml, .git-blame-ignore-revs, codecov.yml, .readthedocs.yml, CODE_OF_CONDUCT.md, SECURITY.md and AGENTS.md all sit at the top level next to CONTRIBUTING.md, which is the one page the README actually links. What the README cannot tell you is the cost of an upgrade: how many releases a given line covers, whether a major version needs spider changes, and what is coming next. That is the gap to close before scheduling one.

Editorial conclusion

Adopt Scrapy when the job is a multi-page crawl with a queue, retries and item pipelines, and stay with a plain HTTP client plus a parser when you need one page. Check first that your interpreter and Python version resolve the dependency markers you expect, because pyproject.toml makes the installed set differ per Python, and read NEWS before scheduling an upgrade since the README never links it. The deciding fact is plain: pip install scrapy hands you Twisted, aiohttp, lxml, parsel and queuelib as one indivisible set, and the project publishes no supported way to install a smaller one.

Frequently asked questions

How do I install Scrapy?

The documented command is pip install scrapy, and the project requires Python 3.10 or newer. A conda-forge badge points to a second distribution channel for machines that do not use pip.

What is Scrapy used for?

It is a web scraping framework for extracting structured data from websites, aimed at developers who need crawling, scheduling and item handling rather than HTML parsing alone.

How do I use Scrapy for web scraping?

Install it with pip install scrapy, then read the documentation at https://docs.scrapy.org/en/latest/. The README itself contains no example spider and no command for running a crawl.

Does Scrapy install differently on Windows?

No Windows specific step is described. The metadata lists Operating System :: OS Independent and calls the project cross-platform, so the same pip install scrapy line is what the project gives you everywhere.

Which Python versions does Scrapy support?

Python 3.10 or newer is required, and the classifiers name 3.10 through 3.15 on both CPython and PyPy. The resolved dependency set still differs by interpreter and by anything below Python 3.14.

Official sources

  1. Official documentation
  2. Official README
  3. Project repository
  4. Release notes
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/scrapy-scrapy.svg)](https://hysenlabs.com/projects/scrapy-scrapy)