# wechat-spider: a mitmproxy-based crawler for WeChat official account articles, read counts and comments

> striver-ing/wechat-spider intercepts WeChat client traffic through a local mitmproxy to collect official account article lists, read and like counts, and comment text into MySQL. It is a deployment-heavy tool for people who accept a proxy certificate on a phone or desktop and want the data in their own database.

**striver-ing/wechat-spider** — 开源微信爬虫：爬取公众号所有 文章、阅读量、点赞量和评论内容。易部署。持续维护！！！

- Repository: https://github.com/striver-ing/wechat-spider
- Stars: 2,973 · Forks: 639
- Language: Python
- License: not declared
- Published: 2026-09-24 · Updated: 2026-09-24 · Language: en
- Canonical page: https://hysenlabs.com/projects/striver-ing-wechat-spider

## What wechat-spider collects and who it is aimed at

The project targets a specific gap: WeChat official accounts do not expose a public API for article lists, read counts, like counts or comment text. wechat-spider's README lists what it captures: daily new articles from an account, account information, article lists, article details, read and like and comment counts, comment content, and conversion of temporary links to permanent links. The output lands in MySQL, which the README presents as a benefit because it makes the data easy to process downstream.

The intended user is someone who already operates a WeChat client and is willing to route its traffic through a proxy. The README's feature list mentions Android, iPhone, Mac and Windows clients, and the repository's docker-compose.yml maps port 8080 and 8081, which matches the README's instruction to point a device's proxy at the host IP and port 8080. This is not a library you import. It is a service you run next to a database and a message queue, then feed account identifiers into.

## How the interception works: mitmproxy, Redis and MySQL

The mechanism is a man-in-the-middle proxy. requirements.txt pins mitmproxy==7.0.3, and the README's setup section tells the reader to install a mitmproxy certificate, then configure the phone or desktop to use the host running wechat-spider as an HTTPS proxy on port 8080. Once the certificate is trusted, the crawler can read WeChat client requests and responses as they pass through.

Two supporting services do the bookkeeping. MySQL stores both the captured data and a task table, and Redis acts as a task cache to reduce how often MySQL is hit. The README describes the queue design as task-state marking, which it says prevents missing any account or article. The docker-compose.yml shows the same shape in code: a wechat-spider service that waits for mariadb:3306 before running python3 ./run.py, plus a redis:3.2 service and a bitnami/mariadb:10.5-debian-10 service. The config directory and the mitmproxy directory are mounted as volumes, so the certificate and config.yaml survive container restarts.

One design detail matters for anyone extending it. The README says the crawler starts by reading wechat_account_task and that you only need to fill in the __biz value, giving MzIxNzg1ODQ0MQ== as an example. Everything else follows from that identifier.

## Installing wechat-spider with Docker and running a first collection

The repository ships a Dockerfile and a docker-compose.yml, so the shortest path is Compose. The Dockerfile is based on python:3.7 and installs dependencies from requirements.txt using an Aliyun pip mirror. The compose file builds an image tagged wechat-spider:latest and starts MariaDB with root credentials root/root, a utf8mb4 character set, and a healthcheck.

From the repository root, the documented entry point is:

```bash
docker compose up -d
```

The wechat-spider service waits for MariaDB on 3306, then runs python3 ./run.py. Ports 8080 and 8081 are published to the host. The mitmproxy directory is mounted from ./mitmproxy and the config directory from ./config, so you edit config.yaml on the host rather than inside the container.

Before the first run, create the database named wechat. The README says that when auto_create_tables in config.yaml is true, the MySQL tables are created automatically, and it recommends setting it to true on the first start and false afterwards to speed up startup. The README's troubleshooting entry for a MySQL connection error shows the shape of the config:

```yaml
mysqldb:
  ip: localhost
  port: 3306
  db: wechat
  user: root
  passwd: "123456"
  auto_create_tables: true
```

The quoting is not cosmetic. The README states that an integer password such as 123456 causes an object supporting the buffer api required exception, and that wrapping it in double quotes fixes it.

Next, install the mitmproxy certificate on the device whose traffic you will capture. The README points to mitm.it for the download and gives per-platform trust steps: on iPhone, enable the mitmproxy entry under Settings, General, About, Certificate Trust Settings; on Android, confirm the certificate under Settings, Security, Trusted credentials; on Windows, place it in Trusted Root Certification Authorities; on macOS, set it to Always Trust in Keychain Access. Then set the device's HTTPS proxy to the host IP and port 8080. If the device is a phone, the README notes it must be on the same router as the machine running wechat-spider.

Finally, insert a task. Add a row to wechat_account_task with only the __biz value, for example MzIxNzg1ODQ0MQ==. When you then open that account's history in the WeChat client and the client shows the expected prompt, the README says the setup is working and the data will appear in the database shortly.

## The account-ban risk and the timeliness ceiling

The README opens with a warning that the crawler is a man-in-the-middle approach, that its timeliness is not high, and that it may get an account banned, and it tells the reader to use it at their own discretion. That single sentence rules out several use cases. If you need continuous real-time monitoring of many accounts, the project's own documentation points elsewhere.

The operational failure modes are also documented rather than hidden. The README's FAQ covers a certificate or security warning that appears even after correct proxy configuration, which it attributes to an expired certificate, and links to external instructions for reinstalling it. It covers an Exception: DISCARD without MULTI, and a case where the software starts normally but captures no packets, with two checks: whether the proxy is set and whether the port is occupied. A no-task message means wechat_account_task has no __biz row, and the README suggests inserting several for testing.

There is also a structural constraint. The whole pipeline depends on a WeChat client generating traffic through the proxy, so the crawl runs at the pace of a human opening account history views, not on a schedule you control. The task-state marking in Redis and MySQL prevents duplicates and omissions, but it does not create traffic.

## How wechat-spider differs from an article exporter

A tool in the article-exporter family, such as the Wechat article-exporter pattern that appears in related searches, typically works from article URLs or a subscription list and produces files. The difference here is the data source. wechat-spider never asks for a URL list; it observes the WeChat client's own requests, which is why it can return read counts, like counts and comment text that a URL-based exporter cannot see from the page alone.

The trade-off is the reverse of an exporter's. An exporter can run headless on a server and produce a document per article. wechat-spider needs a trusted certificate on a device, a proxy pointed at port 8080, and a person or automation driving the client. It gives you a relational schema in MySQL instead of files, which is better for querying counts over time and worse for handing a single article to someone who does not have database access. If your goal is archiving article text, the exporter approach is the simpler fit. If your goal is measuring engagement on accounts you follow, the interception approach is the only one of the two that can produce those numbers from client traffic.

## Maintenance, licence and upgrade cost

The last push to the repository was on 2026-08-28, which is recent, and the repository is not archived. The README's own description claims continuous maintenance, and the commit history is the only evidence available for that claim.

The dependency set is the real upgrade cost. requirements.txt pins mitmproxy==7.0.3, w3lib==1.22.0, PyMySQL==0.10.0, redis==2.10.6, DBUtils==1.3, PyYAML==5.4, parsel==1.6.0, six==1.15.0 and better_exceptions==0.2.2. The Dockerfile builds on python:3.7, which is past its end of life, and the compose file pins redis:3.2 and bitnami/mariadb:10.5-debian-10. Moving any of these forward means retesting the interception path, because mitmproxy's certificate and addon interfaces are the part most likely to change between major versions.

The repository does not state a licence. The README has no licence section and the top-level listing shows no LICENSE file, so the terms under which you may use, modify or redistribute the code are not established. Treat that as an open question to resolve with the author before any commercial or redistributed use; this is not legal advice, and only the copyright holder can grant terms.

## Conclusion

Adopt wechat-spider if you can run MySQL and Redis, install the mitmproxy certificate on a WeChat client, and accept that collection depends on a live client session rather than a stable public API. Do not adopt it if you need unattended long-term monitoring of many accounts, since the README itself warns the man-in-the-middle approach has limited timeliness and may get an account banned. Before deploying, verify three things: that the mitmproxy certificate installs and is trusted on your target device, that config.yaml connects to MySQL and Redis with the password quoted if it is numeric, and that a row in wechat_account_task with a correct __biz produces captured rows in the wechat database.

## FAQ

### What is wechat-spider and what does it do?

It is an open source Python crawler for WeChat official accounts that captures daily new articles, account information, article lists and details, read and like and comment counts, and comment content into MySQL. It works by intercepting WeChat client traffic through a mitmproxy certificate installed on the device.

### How do I install wechat-spider?

The repository provides a Dockerfile and docker-compose.yml, so you can build and start it with Compose; the service waits for MariaDB on 3306 and then runs python3 ./run.py. You also need to install the mitmproxy certificate on your device, set its HTTPS proxy to the host IP and port 8080, and create a database named wechat.

### Why does wechat-spider show object supporting the buffer api required when connecting to MySQL?

According to the README's troubleshooting section, this happens when the MySQL password is an integer such as 123456. Wrapping the value in double quotes in config.yaml, for example passwd: "123456", resolves it.

### Why does wechat-spider report no tasks?

The README says this means no __biz has been inserted into the wechat_account_task table. It suggests inserting several __biz values for testing.

### Why does wechat-spider start normally but capture no packets?

The README lists two checks: whether the proxy is actually configured on the device, and whether the port is occupied.

## Sources

- [Issues](https://github.com/striver-ing/wechat-spider/issues)
- [README](https://github.com/striver-ing/wechat-spider/blob/master/README.md)
- [striver-ing/wechat-spider on GitHub](https://github.com/striver-ing/wechat-spider)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/striver-ing-wechat-spider
