scrapy-redis: one queue, many spiders
Redis-based components for Scrapy.
At a glance
- What is it?
- scrapy-redis provides Redis-based components for Scrapy, turning single-machine crawls into distributed ones through plug-and-play pieces, a Scheduler and duplication filter over a shared Redis queue, an Item Pipeline for distributed post-processing, and base spiders to inherit from. This fork adds JSON-encoded seed data with per-request metadata, idle-poll backoff tuned for fleets of workers, and careful handling of redis-py's RESP2 and RESP3 generations, with v0.9.1 released in July 2024 and the repository still receiving pushes in September 2026.
- Who is it for?
- Use scrapy-redis when a broad, multi-domain crawl needs horizontal scale and your coordination needs fit a shared Redis queue, the scheduler, filter, pipeline and base spiders drop into an existing Scrapy project with settings alone. Use the Frontera project when you need URL expiration or advanced prioritization, which the documentation itself names as beyond its scope, or stick with vanilla Scrapy when one machine suffices, since distribution adds a Redis to operate.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 15 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 2, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The plug-and-play answer to distributed crawling
scrapy-redis is a set of Redis-based components for Scrapy, and its architecture is expressed in the three component groups it ships. The Scheduler and Duplication Filter move Scrapy's request queue and seen-URL memory into Redis, so instead of each spider instance holding its own state, every instance pulls from and pushes to one shared queue, and the duplication filter becomes a shared memory of what the whole fleet has already visited. The Item Pipeline pushes scraped items into a Redis queue of their own, enabling distributed post-processing, as many consumer processes as needed working the items side by side with the crawlers. And Base Spiders provide the inheritance layer that ties a project to the components with minimal modification. The stated sweet spot is broad multi-domain crawls, the workload where one machine's bandwidth and politeness budget run out before the web does, and the documentation draws the boundary honestly, pointing readers with heavier needs, URL expiration, advanced prioritization, at the Frontera project rather than pretending this queue covers them. The design's quiet virtue is that nothing else about the spider changes, scheduling and filtering are swapped at the settings layer while parsing logic, middleware and item schemas stay exactly as the single-machine version left them, so the same project scales from one process to a fleet by editing configuration rather than rewriting code.
pip, GitHub, and a fork's JSON seeds
Two install paths exist, and which one you choose determines what you get, a distinction the README spells out with a warning. The standard path is pip:
pip install scrapy-redisThe GitHub path installs the fork's source directly:
git clone https://github.com/darkrho/scrapy-redis.git
cd scrapy-redis
python setup.py installThe fork's distinguishing feature, JSON supported data in Redis, only exists in the GitHub version, and the note is explicit that if the pip package is already installed it must be removed first:
pip uninstall scrapy-redisThe JSON feature addresses a real orchestration gap, seed records in Redis can carry a url, a nested meta object and optional parameters such as a cookie key, the example showing a job id and start date riding along inside meta, and the component extracts this data and issues a FormRequest carrying url, meta and additional formdata, so within the spider the values arrive naturally as request.url, request.meta and request.cookies. Distributing work across workers thus distributes context too, not just bare URLs, which matters for crawls whose behavior depends on per-job parameters, retry counts tied to a session or targeting data attached at enqueue time by an upstream service.
Idle backoff for fleets sharing one queue
A newer settings block addresses the economics of many workers polling one queue. Idle Redis queue polling can be reduced with opt-in, deadline-gated exponential backoff, four settings governing it, REDIS_IDLE_BACKOFF_ENABLED defaulting off, REDIS_IDLE_BACKOFF_MIN as the initial one-second delay, REDIS_IDLE_BACKOFF_MAX capping at thirty seconds, and REDIS_IDLE_BACKOFF_FACTOR doubling the delay after each empty poll. The mechanics are described precisely, when enabled the spider never blocks the reactor, while waiting for the next poll deadline it skips all Redis calls entirely, and the close deadline is still checked on every idle callback, so shutdown behavior is unaffected. The trade is stated rather than hidden, new work may take up to the backoff cap to be picked up, while idle Redis load drops when many workers share a queue. That framing, a latency-for-load dial tuned by two numbers, is the right abstraction for a fleet where dozens of idle spiders hammering BLPOP are pure waste, and the opt-in default keeps single-user crawls untouched.
One connection pool per crawler
The connection handling shows the same fleet-scale thinking. Components of one crawler share a single Redis connection pool when using the default client class and plain connection parameters, with pools scoped per Settings object, meaning per crawler, so a process running several crawlers keeps their connections separate while each crawler's scheduler, filter and pipeline multiplex one pool instead of opening sockets apiece. Configurations with custom client classes or client-only parameters such as ssl, unix_socket_path or single_connection_client keep the previous per-component behavior, an explicit compatibility carve-out rather than a silent change. REDIS_MAX_CONNECTIONS bounds the shared pool, raising ConnectionError when exhausted, with the documented remedy being a BlockingConnectionPool supplied through REDIS_PARAMS for waiting semantics instead of failing fast, and on redis-py 4.x and 5.x the default applies no override at all, deferring to the library's effectively unlimited pool. It is connection management expressed as deployment knobs, matching how the rest of the settings read.
RESP2, RESP3 and the redis-py generation table
The most modern section of the documentation wrestles with redis-py 8's switch of the default connection protocol to RESP3. The library adds a REDIS_PROTOCOL setting, default None to leave the installed redis-py default unchanged, with the recommended opt-in to RESP2 being a single line:
REDIS_PROTOCOL = 2The equivalent older form remains supported:
REDIS_PARAMS = {"protocol": 2}Both require redis-py 5 or newer, REDIS_PROTOCOL takes precedence when both are set, and an advanced option can omit redis-py's client-identification metadata to trim connection setup:
REDIS_PARAMS = {"driver_info": None}The notes are careful about what a protocol choice does and does not buy, it changes connection setup only and does not remove timeout risk from warm operations, and the historical retry_on_timeout=True default is marked deprecated with no effect on redis-py 6 and above, where the effective retry policy comes from the library's retry object, with advice to measure before lowering retries under churn. A testing observation with Redis 8.0.2, where CLIENT MAINT_NOTIFICATIONS was rejected during RESP3 setup, is scoped to the tested version with diagnostic guidance attached, the level of precision that suggests production debugging produced it.
A tox-powered test container and an example project
The development infrastructure is compact and self-contained. A docker-compose file brings up Redis 6.2 on Alpine beside a Python service that runs tox across three environments, security, flake8 and pytest, with the Redis host and port passed through the environment so integration tests talk to the real broker, and the Dockerfile is a python:3.11-slim image installing both runtime and test requirements before handing off to tox as its command. Quality tooling beyond tox includes bandit with its own configuration and a security badge, coverage configured through coveragerc and reported to Codecov, pre-commit hooks, isort, flake8 and pylint configurations, and a bumpversion setup for releases, the full lint-and-audit battery of a package that expects to be depended upon. Most useful for newcomers, an example-project directory exists in the repository, the place to see a working settings module wired to the scheduler, filter and pipeline before reading any documentation page, and the wiki carries the Usage guide that the README links as the primary entry point, with HISTORY for release notes and a Getting-Started page for contributors.
A 2010s classic, still moving in 2026
The project's position in the ecosystem is worth stating plainly, it is one of the long-standing answers to Scrapy distribution, maintained by R Max Espinoza per the setup metadata, with the repository's history reaching back through the darkrho lineage the install instructions preserve. The version line tells a story of consolidation and quiet care, v0.8.0, v0.9.0 and v0.9.1 all landed in the first week of July 2024, and the repository was last pushed on 2026-09-17, with the intervening period visible in exactly the features documented above, idle backoff, shared pools, RESP3 handling, the maintenance concerns of a library whose users run it against evolving Redis and redis-py generations rather than a library chasing new functionality. The setup classifiers still read Beta and list Pythons only through 3.10 even as requirements ask for 3.7-plus and modern tooling targets 3.11, small metadata drift of a working tool. For teams choosing today, the calculation is whether the shared-queue model fits the crawl, whether the fork's JSON seeds earn the GitHub install, and whether the documented Frontera handoff boundary is far enough away for the workload in hand.
Editorial conclusion
Use scrapy-redis when a broad, multi-domain crawl needs horizontal scale and your coordination needs fit a shared Redis queue, the scheduler, filter, pipeline and base spiders drop into an existing Scrapy project with settings alone. Use the Frontera project when you need URL expiration or advanced prioritization, which the documentation itself names as beyond its scope, or stick with vanilla Scrapy when one machine suffices, since distribution adds a Redis to operate. Verify first that your versions meet the floors, Scrapy 2.0, Redis 5.0 and redis-py 4.2, decide explicitly between RESP2 and RESP3 on modern redis-py, enable idle backoff only where pickup latency up to the cap is acceptable, and note the JSON seed feature requires the GitHub install rather than the pip package.
Frequently asked questions
What is scrapy-redis?
scrapy-redis provides Redis-based components for Scrapy: a Scheduler and Duplication Filter backed by a shared Redis queue, an Item Pipeline for distributed post-processing, and Base Spiders for inheritance. Multiple spider instances share one queue, which suits broad multi-domain crawls.
How do you install scrapy-redis?
Run pip install scrapy-redis for the standard package, or clone the repository from GitHub and run python setup.py install for the fork with JSON seed data support, uninstalling the pip version first if present. Requirements are Scrapy 2.0 or newer, Redis 5.0 or newer and redis-py 4.2 or newer.
Does scrapy-redis reduce Redis load for idle workers?
Yes, opt-in idle backoff reduces queue polling. REDIS_IDLE_BACKOFF_ENABLED turns it on, with MIN, MAX and FACTOR settings controlling the exponential delay from 1 to a default 30 second cap, trading slower pickup of new work for lower idle Redis load across many workers sharing a queue.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/rmax-scrapy-redis)