ProxyPool: a Redis-backed free proxy pool with a Getter, Tester and Flask API
An Efficient ProxyPool with Getter, Tester and Server
At a glance
- What is it?
- Python3WebSpider/ProxyPool scrapes public proxy lists, scores them in Redis and serves random working proxies over HTTP. It is a learning and infrastructure exercise, not a crawler workhorse.
- Who is it for?
- Adopt ProxyPool if you want to study or extend a three-process proxy pipeline, or if you already have proxy sources of your own to plug in. Do not adopt it as the proxy layer for a production crawl: the README states that availability is low and that it is not suitable for direct crawling.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 91 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 2, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What ProxyPool actually solves
Free proxy lists are noisy. A page of a thousand addresses may contain a handful that answer at all, and the working ones die within minutes. ProxyPool exists to automate the tedious part of that cycle: pull addresses from public sources on a schedule, store them in Redis, test them repeatedly, and expose only the survivors through a single HTTP endpoint. The README describes the four functions plainly: scheduled scraping of free proxy sites, Redis storage with availability ranking, periodic testing and filtering, and an API that returns a random passing proxy.
The audience is narrow and the README says so. Its 使用前注意 section warns that because the pool is built from public sources, availability is low, and that it is not suitable for direct use in crawler tasks. Readers who want to finish a crawl quickly are pointed at paid proxy providers instead. That leaves two honest use cases: learning how a proxy pool is architected, and running the pipeline against proxy sources you control or pay for. If you fall into the second group, the value is the plumbing, not the default source list.
Getter, Tester and Server: the three-process design
The project is three cooperating processes coordinated through Redis. Getter crawls proxy websites and writes candidate addresses into Redis. Tester periodically pulls those candidates, attempts connections, and keeps the ones that respond, so the stored set is ranked by usability rather than by recency. Server is a Flask application that answers HTTP requests for a random working proxy. Under Docker, supervisord starts all three; the README's sample startup log shows getter, server and tester each entering RUNNING state alongside Redis.
The separation matters because each stage has a different failure profile. Scraping is bursty and breaks whenever a source changes its HTML, which is why crawlers live in a directory you can extend. Testing is continuous and cheap per address but expensive across thousands of dead ones. Serving is low-latency and stateless. Running them as one process would couple those profiles; running them separately means a broken crawler does not stop the API from serving whatever Redis already holds. You can also start them individually with python3 run.py --processor getter, --processor tester or --processor server, which is useful while debugging one stage.
The storage layer is Redis, configured either through discrete host, port, password and database variables or through a single connection string in the form redis://[:password@]host[:port][/database]. The README notes that the PROXYPOOL_-prefixed variables override the shorter REDIS_ variants, so a container environment can be overridden locally without editing files.
Installing ProxyPool with Docker Compose
The README recommends Docker. Clone the repository and enter the directory first, then bring the stack up. Docker Compose builds the image, starts a redis:alpine container named redis4proxypool, and starts the pool on port 5555 with restart set to always.
git clone https://github.com/Python3WebSpider/ProxyPool.git
cd ProxyPool
docker-compose upThe compose file wires the two services together with a single environment variable, PROXYPOOL_REDIS_HOST: redis4proxypool, which is how the pool finds Redis by container name. The Dockerfile is a two-stage build on python:3.11-slim and installs gcc, g++, libxml2-dev and libxslt1-dev in the build stage to compile lxml. The final image sets APP_ENV=prod, exposes 5555 and declares a volume at /app/proxypool/crawlers/private, which is where you would mount your own crawler modules.
Once the containers are up, the README says to visit http://localhost:5555/random to get a proxy. Expect a plain-text address in host:port form, not JSON.
Running ProxyPool without Docker and pulling a proxy
The conventional route needs Python 3.6 or newer and a reachable Redis. Set the connection either as separate variables or as a connection string; the README shows both and says to pick one.
export PROXYPOOL_REDIS_HOST='localhost'
export PROXYPOOL_REDIS_PORT=6379
export PROXYPOOL_REDIS_PASSWORD=''
export PROXYPOOL_REDIS_DB=0Alternatively, a single string works: export PROXYPOOL_REDIS_CONNECTION_STRING='redis://localhost'. With the environment in place, install dependencies into a virtual environment and start all three processors at once.
pip3 install -r requirements.txt
python3 run.pyThe README's usage example fetches a proxy and immediately spends it on a request, which is the honest way to demonstrate the API because a proxy can die between the fetch and the crawl.
import requests
proxypool_url = 'http://127.0.0.1:5555/random'
target_url = 'http://httpbin.org/get'
proxy = requests.get(proxypool_url).text.strip()
proxies = {'http': 'http://' + proxy}
print(requests.get(target_url, proxies=proxies).text)You should see the proxy printed and an httpbin.org response whose origin field matches that address. Two query parameters extend the endpoint. count returns several proxies, one per line, and returns everything available when it exceeds the pool size. area filters by ISO country code through a bundled GeoLite2 database, so /random?area=CN returns Chinese addresses; proxies whose location cannot be resolved are excluded. Both combine with the key parameter.
Where ProxyPool is the wrong tool
The README's own warning is the first limitation, and it is not modesty. Public proxy sources yield low availability by nature, so a pool built only from them will frequently return addresses that fail on the next request. For a crawl that needs to finish today, the project directs you to paid providers. Treating the default configuration as a production proxy layer misreads the README.
The second limitation is freshness of the surrounding stack. The last push to the repository was on 2026-07-04, while the most recent tagged release is 20230301 from 2023-03-01. The dependency pins reflect that gap: Flask is capped below 3.0, Werkzeug below 3.0, attrs below 24.0 and lxml below 6.0, so installing this alongside a modern Flask application in the same environment will force downgrades. The Docker image sidesteps that by isolating dependencies, which is another argument for the container route.
The third is geographic resolution. The area filter depends on maxminddb_geolite2 pinned at exactly 2018.703, a 2018 vintage database. Country-level filtering will be wrong for addresses that were reassigned since then, and the README does not describe any update path for that data. If you need accurate geolocation, this is the wrong component.
Finally, the private crawler directory is a volume mount, not a plugin API. Adding a source means writing a crawler module that fits the project's internal conventions; the README points to the architecture article rather than documenting that interface itself.
How ProxyPool differs from a maintained proxy scraper like Monosans/proxy
The closest comparison in the related searches is Monosans/proxy, a repository people look for alongside this one. The difference is in what each produces. Monosans/proxy is a scraper-and-checker that publishes lists: its output is a file of verified addresses that consumers download and use directly, with no server process in the middle. ProxyPool produces a running service. Getter, Tester and Server stay resident, Redis holds the current state, and consumers ask an HTTP endpoint for a proxy at request time.
That distinction drives the operational trade-offs. A published list is trivially cacheable and needs no Redis or long-running process, but it is a snapshot, and you inherit the publisher's testing cadence. A live pool lets you control the test interval, plug in private sources, and filter by country at query time, at the cost of running three processes and a Redis instance. If your crawler can tolerate a list that refreshes hourly, the file-based approach is less machinery. If you need per-request selection with your own sources mixed in, ProxyPool's architecture is the reason to pick it. The project itself is not a proxy provider and does not sell addresses; the paid services it links to are a different category entirely.
Configuration, licence and the cost of keeping it running
Runtime behaviour is controlled by environment variables. ENABLE_TESTER, ENABLE_GETTER and ENABLE_SERVER each default to true and let you disable a stage, which is how you run a server against an externally maintained Redis or test a crawler in isolation. APP_ENV accepts dev, test or prod and defaults to dev. APP_DEBUG defaults to true, which is worth turning off if the Flask server is reachable from anywhere but localhost. APP_PROD_METHOD defaults to gevent and can be set to tornado or meinheld, though the README notes those require installing the corresponding module; tornado is already in requirements.txt, meinheld is not.
Upgrade cost is dominated by the dependency bounds rather than by the pool logic. Because the pins are ranges with low ceilings, a pip install -r requirements.txt into an existing environment can silently move Flask and Werkzeug backwards. The Dockerfile avoids this by building in an isolated image, and the README also suggests swapping in a faster PyPI mirror if the build is slow. There is no documented migration path between releases; the releases are dated snapshots (20230301, 20220306, 20220101) with no changelog in the repository.
The licence is MIT, which permits commercial use, modification and redistribution provided the copyright notice and permission notice are retained. That is a permissive arrangement, but it says nothing about the legality or terms of use of the proxy sources the crawlers contact, and the README does not address that question. Whether scraping a given proxy list is permitted is a separate matter you have to settle yourself.
Editorial conclusion
Adopt ProxyPool if you want to study or extend a three-process proxy pipeline, or if you already have proxy sources of your own to plug in. Do not adopt it as the proxy layer for a production crawl: the README states that availability is low and that it is not suitable for direct crawling. Before you commit, check whether the crawler for your target source exists under proxypool/crawlers, confirm your Redis connection string parses, and decide whether the GeoLite2 area filter matters for your traffic.
Frequently asked questions
What is a proxy pool, and what does ProxyPool do?
A proxy pool is a set of proxy addresses that a program can draw from, usually filtered to those currently working. ProxyPool builds one by scraping free proxy sites with a Getter, testing and ranking candidates in Redis with a Tester, and serving random passing proxies through a Flask API on port 5555.
Is ProxyPool suitable for a real crawling job?
The README states that because the pool is built from public proxy sources, availability is low and it is not suitable for direct use in crawler tasks. It recommends paid proxies or existing proxy resources if the goal is to finish a crawl quickly.
How do I install and run ProxyPool?
The README recommends Docker: clone the repository and run docker-compose up, which starts Redis and the pool with the Getter, Tester and Server processes. Without Docker you need Python 3.6 or newer and Redis, then pip3 install -r requirements.txt followed by python3 run.py.
How do I get more than one proxy at a time from ProxyPool?
Pass a count parameter to the /random endpoint, for example http://localhost:5555/random?count=5, which returns multiple proxies one per line. If count is omitted or set to 1 the behaviour is unchanged, and if it exceeds the available number the endpoint returns all available proxies.
Can ProxyPool return proxies from a specific country?
Yes. The /random and /all endpoints accept an area parameter holding an ISO country code, case-insensitive, so /random?area=CN returns Chinese proxies. Country data comes from a bundled GeoLite2 offline database, and proxies whose location cannot be resolved are excluded.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/python3webspider-proxypool)