# DouYin_Spider: a Python toolkit for Douyin data collection, live-room events and direct messages

> DouYin_Spider wraps Douyin's web APIs behind a Python client: profile and comment scraping, a WebSocket listener for live-room danmaku and gifts, and direct-message send and receive. It suits engineers who already have a logged-in Douyin account and want programmatic access, not a managed data service.

**cv-cat/DouYin_Spider** — 抖音逆向，抖音爬虫，抖音全部api、私信、直播间监听

- Repository: https://github.com/cv-cat/DouYin_Spider
- Stars: 3,241 · Forks: 866
- Language: Python
- License: not declared
- Published: 2026-09-24 · Updated: 2026-09-24 · Language: en
- Canonical page: https://hysenlabs.com/projects/cv-cat-douyin-spider

## What DouYin_Spider is for, and who should not use it

DouYin_Spider targets a specific gap: Douyin has no public API for the data and events this project exposes. The README frames the motivation as giving AI agents a way to reach the platform, saying that the first wall is usually not model capability but the absence of platform communication capability. Concretely, the project covers profile and post detail, comment threads including nested replies, search across video, user and live, follower and following lists, notifications, favorites and the recommendation feed. It also listens to live rooms for danmaku, gifts with the recipient, entries, follows, likes and room heat, and it sends and receives direct messages over WebSocket.

The audience is narrow. You need a Douyin account you are willing to log into from a browser, extract a cookie from, and keep alive. You need to accept that the project scrapes undocumented endpoints rather than calling a contract. And you need to be comfortable with the legal framing the README itself gives: the project is stated to be for learning and technical research only, with a warning against publishing harmful or illegal content and a note that the author disclaims responsibility for consequences. If your use case requires a documented SLA or a data-processing agreement, this is not the tool.

## How the pieces fit: cookies, a TLS-fingerprinting HTTP client and two WebSocket paths

The repository splits into three runnable entry points plus shared utilities. main.py is the scraping entry point. dy_live/server.py holds the live-room WebSocket listener. dy_apis/douyin_recv_msg.py handles real-time direct-message reception, and dy_apis/douyin_api.py is described in the README as containing all the API wrappers, including live-room likes, message sending and direct-message send and receive.

Authentication is cookie-based. The README states that configuring DY_COOKIES alone is now recommended, because the main-site Auth value is reused by the post endpoints, the live REST and WebSocket paths, and the creator center. DY_LIVE_COOKIES still exists as an explicit override for the older standalone live browser, but it is no longer required. A second credential set, DY_TICKET, DY_TS_SIGN, DY_CLIENT_CERT and DY_PRIVATE_KEY, is written automatically after a QR-code login via login_grab_ticket and is used for direct-message signing.

The HTTP layer is worth noting. requirements.txt pins curl_cffi and explains that all Douyin interface calls go through utils/http_client.py, which impersonates Chrome's TLS and HTTP/2 fingerprint. The comment states that requests remains only in utils/data_util.py for downloading CDN media, a path that does not touch risk control. That is a deliberate design choice: fingerprint-level impersonation rather than header spoofing. It also means the project is coupled to whatever curl_cffi supports.

## Installing DouYin_Spider and running a first scrape

The README lists Python 3.7+ and Node.js 18+ as the runtime requirements, and the Dockerfile builds from python:3.10-slim and exposes port 5000. Dependencies split into a Python set and a Node set.

```bash
pip install -r requirements.txt
npm install
```

Before running anything, create the .env file at the project root from .env.example. The README instructs you to open the browser developer console on a logged-in Douyin session, go to the network tab, filter for fetch, open any request and copy its cookie value. The README states explicitly that only a cookie taken after logging in is valid. The .env.example file lists the keys as follows.

```bash
DY_COOKIES=''
DY_LIVE_COOKIES=''
DY_TICKET=''
DY_TS_SIGN=''
DY_CLIENT_CERT=''
DY_PRIVATE_KEY=''
```

The four ticket and key fields are not filled by hand: the README says login_grab_ticket writes them after a QR-code login, and they are used for direct-message signing. With DY_COOKIES populated, the scraping entry point runs directly.

```bash
python main.py
```

Output is written to a structured directory and formatted as JSON, Excel or media, according to the README's feature list. The other two entry points are separate processes: python dy_live/server.py for live-room danmaku, gifts and likes, and python dy_apis/douyin_recv_msg.py for the direct-message WebSocket. The README notes that main.py is meant to be edited for your own calls, and that dy_apis/douyin_api.py is where the full API surface lives.

## The protobuf pin is the install step most likely to break

requirements.txt carries an unusually explicit warning. static/Response_pb2.py is generated by a newer protoc and needs the protobuf>=5.27 runtime_version. The comment states that blackboxprotobuf declares a dependency on protobuf==3.10, which would downgrade protobuf and produce ImportError: cannot import name 'runtime_version'. The file pins protobuf>=5.27 and notes that the two can coexist at runtime, but asks you to confirm protobuf was not downgraded after installing.

This is a real constraint, not a formality. If you install blackboxprotobuf into the same environment without the pin in place, the resolution can pull protobuf down and the protobuf-based response parsing fails at import time. The requirements.txt comment asks you to confirm the version after installing, and the same file is the place to check what is pinned.

A second operational limitation is stated in the changelog rather than the setup section: an entry dated 23/10/28 says that when a captcha appears, you must click it manually. There is no documented automated captcha path. Any long-running collection job therefore has a human in the loop, and the README does not document what happens to an in-flight scrape when the challenge appears.

## Live-room and direct-message work is a different operating model

Scraping a profile is a request-response job. Listening to a live room is not. dy_live/server.py holds a WebSocket open and emits events as they arrive: danmaku, gifts including the recipient, entries, follows, likes and room heat. The README also lists sending danmaku and liking a room from the same surface. Direct messages run on their own WebSocket in dy_apis/douyin_recv_msg.py and handle text, stickers, voice, images and shared videos, with active sending and conversation list creation and query alongside.

The README claims automatic retry and reconnect behavior for this architecture, but it does not document the reconnect policy: no backoff parameters, no maximum attempt count, no statement about whether events missed during a disconnect are recovered. That matters if you are building anything that must not lose gifts or messages. Treat the listener as best-effort until you have measured it against your own rooms.

The signing requirement is the other asymmetry. Direct messages depend on DY_TICKET, DY_TS_SIGN, DY_CLIENT_CERT and DY_PRIVATE_KEY, which are produced by the QR-code login flow. The README does not document how long those credentials remain valid or what the failure mode looks like when they expire. Live-room listening, by contrast, is described as reusing the main Auth cookie. So a deployment that only reads live rooms has a simpler credential story than one that sends or receives direct messages.

## Alternatives and the honest comparison

The closest alternative in practice is writing your own client against the same endpoints. That is not a trivial substitution: the project's value is concentrated in two places you would otherwise have to reconstruct, the curl_cffi fingerprint impersonation in utils/http_client.py and the protobuf schema in static/Response_pb2.py, plus the WebSocket framing for live rooms and direct messages. If you only need one endpoint, a small script is less to maintain. If you need the breadth the README lists, rebuilding it is the larger project.

A second alternative is a general-purpose browser automation stack driving a logged-in session. That approach sidesteps fingerprint work because the real browser does it, and it handles captchas interactively by design. The difference in approach is fundamental: browser automation pays in latency, memory and fragility against UI changes, while DouYin_Spider pays in coupling to undocumented endpoints and to a pinned protobuf runtime. Neither is strictly better; they fail in different places. The project's own changelog shows the maintenance cost of the endpoint route, with entries for fixing live-room monitoring and gift information across 2023 and 2026.

When this is the wrong tool: if you need historical bulk data with reproducible schemas, or if your compliance posture requires a documented data source, no scraping library of this kind fits.

## Maintenance, licence and upgrade cost

The repository is not archived and the last push was on 2026-09-19, five days before this writing. Release v2.0.0 is dated 2026-09-13; the previous release, v1.1.0, is dated 2023-10-21. That gap is the useful signal. The changelog jumps from 25/06/07, when previously closed-source code for scraping and live-room listening was opened, to 26/04/09, which added gift recipient information, live-room likes, live-room danmaku sending, and direct-message receive and send. So the direct-message and interaction features are recent, and the older scraping surface has had years of bug-fix entries.

Upgrade cost is dominated by the protobuf pin and by cookie churn. The requirements.txt comment asks you to confirm protobuf is not downgraded after install, which means every dependency change is a potential runtime break. The README states that the project adapts to Douyin's latest API, and the changelog entries about fixing live-room monitoring show that this adaptation is continuous rather than one-off. Budget for periodic updates rather than a set-and-forget install.

On licensing: the repository has no LICENSE file, and the README carries a disclaimer that the project is for learning and technical research only. Without a licence, the default position is that no rights are granted beyond what the platform itself allows, so redistribution or commercial use is undefined. That is a question for your own counsel, not something this review can settle.

## Conclusion

Adopt DouYin_Spider if you already operate a logged-in Douyin account, can refresh cookies and tickets yourself, and want one Python client that covers profiles, comments, live-room events and direct messages. Do not adopt it if you need a documented, versioned API contract, a managed service, or a guarantee about account safety: the README states it is for learning and technical research only, warns that it must not be used to publish harmful or illegal content, and the repository has no LICENSE file, so redistribution rights are undefined. Before building on it, verify which of the three entry points you actually need (main.py, dy_live/server.py, dy_apis/douyin_recv_msg.py), confirm that protobuf is still at 5.27 or above after pip install, and test the cookie and ticket flow against your own account.

## FAQ

### What are the runtime requirements for DouYin_Spider?

The README lists Python 3.7+ and Node.js 18+, and the Dockerfile builds from python:3.10-slim with port 5000 exposed. Dependencies install through pip install -r requirements.txt and npm install.

### Does DouYin_Spider need a logged-in Douyin account?

Yes. The README instructs you to copy the cookie from a browser session that is already logged into Douyin and states that a cookie taken without logging in is not valid. DY_COOKIES is the recommended single setting, with DY_LIVE_COOKIES kept only as an override for the older standalone live browser.

### How do I run the live-room danmaku and gift listener?

The README gives python dy_live/server.py as the entry point for live-room listening, covering danmaku, gifts including the recipient, entries, follows, likes and room heat. It is a separate process from the scraper in main.py.

### Why does installing DouYin_Spider downgrade protobuf?

requirements.txt explains that blackboxprotobuf declares a dependency on protobuf==3.10, which would downgrade the runtime and cause ImportError: cannot import name 'runtime_version' because static/Response_pb2.py needs protobuf>=5.27. The file pins protobuf>=5.27 and asks you to confirm it was not downgraded after installing.

### What happens when Douyin shows a captcha during a scrape?

The changelog entry dated 23/10/28 states that you must click the captcha manually. The README does not document an automated captcha path or what happens to a running job while the challenge is open.

## Sources

- [cv-cat/DouYin_Spider on GitHub](https://github.com/cv-cat/DouYin_Spider)
- [Issues](https://github.com/cv-cat/DouYin_Spider/issues)
- [README](https://github.com/cv-cat/DouYin_Spider/blob/master/README.md)
- [Releases](https://github.com/cv-cat/DouYin_Spider/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/cv-cat-douyin-spider
