jhao104/proxy_pool: a self-hosted free proxy pool for Python spiders
Python ProxyPool for web spider
At a glance
- What is it?
- proxy_pool collects free proxies from public sites, validates them on a schedule and serves them over a Flask API backed by Redis. It is a scraper convenience, not a privacy tool, and the README is explicit that free proxies are low quality.
- Who is it for?
- Adopt proxy_pool if you run Python crawlers and want a local Redis-backed pool of validated free proxies behind a small HTTP API, and you accept that free sources are unreliable. Do not adopt it if you need residential-grade anonymity or a hosted service; the README itself points to paid providers for that.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 106 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 27, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What proxy_pool actually solves for a spider
A crawler that sends every request from one IP gets rate-limited or blocked. The usual workaround is to rotate source addresses, but maintaining that rotation by hand means finding proxy lists, testing them, discarding dead entries and repeating the whole cycle as they expire. proxy_pool automates that loop. According to the README, it periodically collects free proxies published on public sites, validates them, stores the working ones, and exposes them through both an API and a CLI.
The audience is narrow and worth stating plainly: Python developers writing spiders who already have Redis available and want a local pool they control. It is not a privacy product. The proxies it manages are free public endpoints whose operators are unknown, so traffic through them is visible to whoever runs them. Anyone reaching for this as a VPN substitute has the wrong tool.
Schedule process, Redis store and Flask API
The project splits into two processes. python proxyPool.py schedule runs the collection and validation loop; python proxyPool.py server runs the HTTP API. Both read configuration from setting.py, which holds HOST, PORT and DB_CONN. The README's default DB_CONN example is a Redis URL of the form redis://:[email protected]:8888/0, so the pool's state lives in Redis rather than on disk, and the API process reads from the same store the scheduler writes to.
Proxy sources are pluggable. The README states that by default the scheduler scans fetcher/sources/ for every source with enabled=True, and that PROXY_FETCHER_EXCLUDE in setting.py takes a list of source names to skip. Adding a source means dropping a .py file into fetcher/sources/, subclassing BaseFetcher, declaring name, url and enabled, and implementing fetch() as a generator that yields host:port strings. No registry file needs editing. That auto-scan design is the strongest part of the architecture: extension is a file drop, and a broken source can be disabled by name without deleting code.
Installing proxy_pool and pulling a first proxy
The README gives a clone-then-pip path. Clone the repository and install the pinned dependencies from requirements.txt. Note the pins: Flask 2.1.1, requests 2.31.0, redis>=4.2.0, and APScheduler split by Python version (3.10.0 for 3.10 and above, 3.2.0 below).
git clone https://github.com/jhao104/proxy_pool.git
cd proxy_pool
pip install -r requirements.txtBefore starting anything, edit setting.py. The README shows HOST, PORT and DB_CONN as the keys to change, and DB_CONN must point at a reachable Redis instance. The default port in the README's API section is 5010, while the setting.py sample shows PORT = 5000, so check which value your checkout actually carries rather than assuming.
HOST = "0.0.0.0"
PORT = 5000
DB_CONN = 'redis://:[email protected]:8888/0'Start the scheduler in one terminal and the API in another. The scheduler needs time to fetch and validate before the pool holds anything usable.
python proxyPool.py schedule
python proxyPool.py serverWith the server up, the README documents GET /get for a random proxy, GET /pop to take one and remove it, GET /all for everything, GET /count for the size, and GET /delete?proxy=host:ip to drop a specific entry. The optional ?type=https parameter filters for proxies that support HTTPS. A first real use is a retry loop around a request, deleting the proxy when it fails, which the README demonstrates in Python.
import requests
def get_proxy():
return requests.get("http://127.0.0.1:5010/get/").json()
def delete_proxy(proxy):
requests.get("http://127.0.0.1:5010/delete/?proxy={}".format(proxy))The Docker path is shorter. The README shows pulling jhao104/proxy_pool and running it with DB_CONN passed as an environment variable and port 5010 published, and docker-compose.yml in the repository wires the app to a redis service on the same network.
Free proxies are the failure mode, not a footnote
The README does not pretend otherwise: it says the project ships several free proxy sources, that free proxies are limited in quality, and that running it as-is may yield unsatisfactory results. That is the honest core of this project. A pool built from public free lists will have low and volatile availability, and the table of supported sources rates update speed and availability with stars, which is a rough signal rather than a guarantee.
The second limitation is operational. The scheduler and API are separate processes sharing Redis, so if Redis is unreachable the API has nothing to serve. There is no documented fallback to an in-memory store. The README also does not document rollback, migration or backup for the Redis contents, so pool state is effectively disposable and rebuildable rather than something to protect. And because the sources are scraped from third-party sites, a layout change on any of those sites can silently break a fetcher until someone notices the count dropping.
How this differs from a paid proxy service
The obvious alternative is a commercial proxy provider, and the README itself recommends one, Bright Data, alongside the free sources. The difference in approach is not just price. A paid provider supplies addresses from its own network under a contract, with stated coverage and support; proxy_pool supplies addresses scraped from public lists whose operators are anonymous. With proxy_pool you own the validation loop, the Redis store and the failure handling; with a paid service you own only the integration and the bill.
A second alternative is writing the rotation yourself: a small script that fetches a list, tests each entry against a target URL and caches the survivors. That is essentially what proxy_pool already does, so rolling your own only makes sense if you need a different storage backend, a scheduler you already run, or validation semantics tuned to one specific target site. For most Python spider projects the existing fetcher base class and API are less work than a bespoke loop.
Maintenance, licence and upgrade cost
The repository is not archived, and the last push was on 2026-06-15. The most recent tagged release is 2.4.1 from 2023-02-23, with 2.4.0 in 2021 and 2.3.0 in 2021 before it, so release cadence is slow even though the default branch has seen later activity. Treat the code as stable rather than fast-moving, and expect to read commits rather than release notes when you upgrade.
The licence is MIT, which permits commercial use, modification and redistribution provided the copyright notice and permission notice are included. That is a permissive baseline, but it says nothing about the legality or terms of the proxy sources the fetchers scrape, and nothing about the operators of those proxies. Those are separate questions the MIT text does not answer, and this is not legal advice.
Upgrade cost is dominated by the pinned dependencies. Flask is held at 2.1.1 with werkzeug constrained to >=2.0,<2.2, and APScheduler is pinned differently for Python 3.10 and above versus below. Moving Python versions means re-checking that split. The Dockerfile builds on python:3.10-alpine and installs musl-dev, gcc, libxml2-dev and libxslt-dev at build time before removing the compilers, so image builds depend on those Alpine packages staying available.
Editorial conclusion
Adopt proxy_pool if you run Python crawlers and want a local Redis-backed pool of validated free proxies behind a small HTTP API, and you accept that free sources are unreliable. Do not adopt it if you need residential-grade anonymity or a hosted service; the README itself points to paid providers for that. Before committing, verify the Python version against the badges (3.8 to 3.11), confirm the Redis URL in setting.py, and run python proxyPool.py fetcher to see which sources are enabled in your checkout.
Frequently asked questions
What is proxy_pool in jhao104/proxy_pool?
It is a Python proxy pool for spiders that periodically collects free proxies published online, validates them, stores the working ones in Redis and serves them through an API and a CLI. The README describes it as a crawler proxy IP pool project.
How do I install jhao104/proxy_pool?
Clone the repository, run pip install -r requirements.txt, edit setting.py for HOST, PORT and DB_CONN, then start python proxyPool.py schedule and python proxyPool.py server in separate terminals. The README also documents a Docker image and a docker-compose setup.
Which API endpoints does proxy_pool expose?
The README lists GET / for the API introduction, /get for a random proxy, /pop to fetch and remove one, /all for every proxy, /count for the number held, and /delete with a proxy=host:ip parameter. The /get, /pop and /all endpoints accept an optional ?type=https filter.
Can I add my own proxy source to jhao104/proxy_pool?
Yes. The README says to create a .py file under fetcher/sources/, subclass BaseFetcher, declare name, url and enabled, and implement fetch() as a generator yielding host:port strings. The scheduler picks up new sources on its next run without configuration changes.
Are the free proxies in proxy_pool reliable?
The README states that free proxies are limited in quality and that running the project as-is may produce unsatisfactory results. It recommends paid providers for higher-quality IPs and rates the supported free sources with rough star marks for update speed and availability.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/jhao104-proxy-pool)