# wzdnzd/aggregator: a Python proxy pool built from crawled sources

> The repository crawls Telegram, GitHub and other public channels for proxy nodes, validates them, and republishes the survivors as Clash, V2Ray or SingBox subscriptions. It is a scraping pipeline, not a proxy client, and the README is explicit that it is meant for learning crawler technique.

**wzdnzd/aggregator** — One-stop Proxies Crawling and Aggregation Platform

- Repository: https://github.com/wzdnzd/aggregator
- Website: https://github.com/wzdnzd/aggregator
- Stars: 6,776 · Forks: 5,484
- Language: Python
- License: Apache-2.0
- Published: 2026-09-22 · Updated: 2026-09-22 · Language: en
- Canonical page: https://hysenlabs.com/projects/wzdnzd-aggregator

## The problem: public proxy announcements are scattered and short-lived

Free proxy nodes are announced in places that were never designed to be read by machines. The README lists Telegram, GitHub, Google, Yandex and Twitter among the crawl targets. Each of those has a different shape: a Telegram channel posts a base64 blob, a GitHub repository commits a YAML file, a search result page returns links that need a second fetch. Collecting by hand means checking each source, decoding whatever encoding it uses, and testing whether the node still answers.

The project is for someone who wants that loop automated and wants the output in a client format rather than a list of host:port pairs. It is not for someone who wants a proxy client. Aggregator produces subscription files; something else consumes them. The README's own framing is narrower still: the disclaimer says the project is for studying crawler technique and that commercial use is prohibited.

## Two entry points with different amounts of control

The repository ships two scripts under subscribe/. collect.py is the simplified path: it collects airport subscriptions, registers accounts, fetches the resulting subscription, validates the nodes, and uploads to a Gist. process.py is the full path: it loads a JSON config, crawls multiple sources, aggregates, runs quality checks, converts formats, and pushes to the configured storage backend.

The README's workflow diagram puts the difference plainly. collect.py runs a fixed sequence with no source configuration. process.py branches on config: sites you name, crawl sources you enable, groups that map to output targets, and a storage engine. If you only want a working subscription file, collect.py is fewer moving parts. If you want to decide which Telegram channels or which airport domains feed the pool, process.py is the only one of the two that accepts that input.

Validation is not a separate product. The README credits Mihomo as the proxy testing engine and Subconverter as the subscription conversion core, both vendored under the repository (clash/ and subconverter/). Format conversion is therefore delegated rather than reimplemented.

## Installing and running a first collection

The README gives no pip install line and no published package name; requirements.txt lists PyYAML, tqdm, geoip2, pycryptodomex and fofa-hack, so the intended path is a clone plus a dependency install. The Dockerfile is the more reproducible route and is the one worth reading first, because it pins the runtime and states its build arguments.

Build the image with the same command the Dockerfile documents in its header comment. The PIP_INDEX_URL build argument defaults to https://pypi.org/simple and can be pointed at a mirror:

```bash
docker buildx build --platform linux/amd64 -f Dockerfile -t wzdnzd/aggregator:tag --build-arg PIP_INDEX_URL="https://pypi.tuna.tsinghua.edu.cn/simple" .
```

The image sets three environment variables, GIST_PAT, GIST_LINK and CUSTOMIZE_LINK. GIST_LINK takes the form username/gist_id. Its default CMD runs collect.py with --all, --overwrite and --skip, so a bare docker run starts a full collection and skips work it has already done.

For the config-driven path, copy the example config and edit it. The README shows this minimal shape, with an airport site, a Telegram crawl source, a group, and a Gist storage item:

```json
{
    "sites": [
        {
            "name": "example-airport",
            "domain": "example.com",
            "push_to": ["free"]
        }
    ],
    "crawl": {
        "enable": true,
        "telegram": {
            "enable": true,
            "users": {
                "proxy_channel": {
                    "push_to": ["free"]
                }
            }
        }
    },
    "groups": {
        "free": {
            "targets": {"clash": "free-clash"}
        }
    },
    "storage": {
        "engine": "gist",
        "items": {
            "free-clash": {
                "username": "your-username",
                "gist_id": "your-gist-id",
                "filename": "clash.yaml"
            }
        }
    }
}
```

The push token is read from the environment rather than the config file:

```bash
export PUSH_TOKEN=your_github_token
```

Then run the processor against your edited config. The README's quick-collect variant takes the Gist path and token as flags and a list of target formats:

```bash
python subscribe/collect.py -g username/gist-id -k your-github-token -t clash v2ray singbox
```

What you should see after a run is a Gist item per group target, in the format named by the target key. If nothing appears, the README's troubleshooting table points at the crawl source configuration and the network connection first, and at token permissions second. It also offers a syntax check for the config itself:

```bash
python -m json.tool config.json
```

## Storage backends and the token they depend on

Output does not stay local. The README names GitHub Gist, PasteGG and Imperial as storage backends, selected through the storage.engine key. That design makes the tool useful for a cron job: the process can run on a machine that is not the one consuming the subscription, and the client reads the published URL.

It also concentrates risk in one credential. The README's FAQ lists an invalid token as a distinct failure and points at GitHub token permissions and expiry. The Dockerfile names the variable GIST_PAT, while the README's process.py example exports PUSH_TOKEN. Both appear in the repository, so a container run and a local run may need different variable names; check which one your path reads before assuming a failed push is a permissions problem.

PasteGG and Imperial are named but not documented in the README beyond the list, so treat Gist as the path with examples and the others as paths you will have to read the source for. That is a real gap for anyone whose threat model rules out publishing a subscription to a public Gist.

## Where it breaks, and where it is the wrong tool

The pipeline depends on other people's infrastructure. It scrapes public channels and airport sites, so when a channel changes its posting format or an airport changes its registration flow, the corresponding collector stops producing nodes. The README's TODO list acknowledges this class of problem indirectly: crawler pluginization, a plugin registration mechanism, and a plugin configuration standard are all unchecked. Until those land, extending a collector means editing the existing code, not dropping in a plugin, despite the README describing a plugin system as a core feature.

The same list also leaves core interfaces (ICrawler, IStorage, IConverter), base classes, and a factory pattern unchecked. So the extension story is a stated intention, not a documented contract. If your plan is to write a custom crawler against a stable interface, the repository does not give you one.

Two further limits are worth stating. First, node quality is whatever the sources publish; validation tells you a node answered, not that it will answer tomorrow, and the README's own sharing note warns users not to waste the shared subscriptions. Second, the disclaimer restricts the project to learning use and prohibits profitable use, which rules it out for anything commercial regardless of technical fit. If you need a proxy client, this is not one. If you need a guaranteed service level, no amount of crawling produces it.

## Alternatives and how they differ

The closest architectural comparison in the repository is Subconverter, which aggregator vendors under subconverter/ and credits as the conversion core. Subconverter takes subscription URLs you already have and rewrites them into a client format. Aggregator sits one step earlier: it goes looking for the nodes, tests them, and only then hands the survivors to Subconverter. If you already have subscription URLs you trust, Subconverter alone is the smaller dependency.

Mihomo appears in the acknowledgements as the proxy testing engine and is vendored under clash/. It is a proxy runtime that also exposes a testing mode. Aggregator uses it as a component; it is not a replacement for the aggregation layer, because it does not crawl or deduplicate sources.

Against hand-maintained subscription lists, the difference is the validation step. A curated list is only as fresh as its maintainer's last edit. Aggregator's design assumes sources are noisy and filters them on every run, which is more work per run and less work per month.

## Maintenance, licence and what to verify before you commit

The repository is not archived and the last push was on 2026-09-20. That is recent enough to call it maintained, but the TODO list is the more useful signal: it describes a planned refactor of the core interfaces, the plugin system, the configuration system, and the exception and logging layers. Work of that scope means file layout and internal APIs may move. Pin a commit or an image tag rather than tracking main if you build on top of it.

The licence is Apache-2.0, which permits commercial use and modification with the usual notice and patent terms. The README's disclaimer is separate from the licence and prohibits profitable use of the project. Those two statements point in different directions, and resolving that tension is a question for a lawyer, not for this article. The disclaimer also requires users to follow local law and site terms of service, which matters because the crawlers fetch third-party pages and channels.

Before running it, verify that python -m json.tool accepts your config, that the token variable your chosen entry point reads is set (PUSH_TOKEN for the README's process.py example, GIST_PAT in the Dockerfile), and that the storage engine you selected is one the README actually documents. Gist is; PasteGG and Imperial are named only.

## Conclusion

Adopt wzdnzd/aggregator if you want a self-hosted pipeline that turns scattered public proxy announcements into a validated subscription file, and you are willing to run it on a schedule against your own Gist. Do not adopt it if you need a proxy client, a stable commercial service, or a tool with a documented plugin API, because the README lists interfaces such as ICrawler and IStorage as unchecked roadmap items. Verify two things before committing: that your config passes python -m json.tool, and that PUSH_TOKEN has the Gist scope the storage backend needs.

## FAQ

### What is wzdnzd/aggregator and who is it for?

It is a Python tool that crawls public sources such as Telegram, GitHub, Google, Yandex and Twitter for proxy nodes, validates them, and converts them into client formats including Clash, V2Ray and SingBox. The README positions it as a tool for studying crawler technique, not as a proxy client.

### How do I install and run wzdnzd/aggregator?

The README gives no pip package, so the documented paths are a clone with the dependencies in requirements.txt, or the Dockerfile, which builds on python:3.12.3-slim and defaults to running subscribe/collect.py with --all, --overwrite and --skip. For the config-driven path, copy subscribe/examples/config.default.json, set PUSH_TOKEN, and run python subscribe/process.py -s my-config.json.

### Where does wzdnzd/aggregator store the proxies it collects?

The README names GitHub Gist, PasteGG and Imperial as storage backends, selected with the storage.engine key in the config. Only Gist has a worked config example in the README, and its items take username, gist_id and filename.

### What does wzdnzd/aggregator use to test and convert proxies?

The acknowledgements credit Subconverter as the subscription conversion core and Mihomo as the proxy testing engine, and both are vendored in the repository under subconverter/ and clash/. The Dockerfile copies clash-linux-amd and Country.mmdb into the image.

## Sources

- [Issues](https://github.com/wzdnzd/aggregator/issues)
- [License: Apache-2.0](https://github.com/wzdnzd/aggregator/blob/main/LICENSE)
- [Project website](https://github.com/wzdnzd/aggregator)
- [README](https://github.com/wzdnzd/aggregator/blob/main/README.md)
- [wzdnzd/aggregator on GitHub](https://github.com/wzdnzd/aggregator)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/wzdnzd-aggregator
