# MediaCrawler: the crawler drives the Chrome session you already logged in to

> Instead of reverse-engineering signatures, MediaCrawler attaches to your own browser over a debugging port and asks the page for them. That is the technical trick, and it is also why a run needs your real account, a live confirmation click, and a careful reading of the disclaimer at the top of the file.

**NanmiCoder/MediaCrawler** — 小红书笔记 | 评论爬虫、抖音视频 | 评论爬虫、快手视频 | 评论爬虫、B 站视频 ｜ 评论爬虫、微博帖子 ｜ 评论爬虫、百度贴吧帖子 ｜ 百度贴吧评论回复爬虫  | 知乎问答文章｜评论爬虫

- Repository: https://github.com/NanmiCoder/MediaCrawler
- Website: https://nanmicoder.github.io/MediaCrawler/
- Stars: 65,902 · Forks: 12,695
- Language: Python
- License: NOASSERTION
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/nanmicoder-mediacrawler

## It attaches to your own Chrome over a debugging port

The technical claim at the top of the README is that no JavaScript reverse engineering is needed. Instead of decoding a platform's signature algorithm, the crawler runs inside a browser context that already holds your login state and asks the page for its signature parameters with JavaScript expressions. The default mode is CDP, which connects to a Chrome instance you already have open, so the session, cookies and extensions are the real ones.

Setting that up is deliberately hands-on. You need Chrome 144 or newer, you enable remote debugging through `chrome://inspect/#remote-debugging` by ticking "Allow remote debugging for this browser instance", and the browser then reports `Server running at: 127.0.0.1:9222`. When a crawl starts, Chrome raises a confirmation dialog and the program waits for you to accept it, with sixty seconds to do so. Set `ENABLE_CDP_MODE = False` in `config/base_config.py` to fall back to standard Playwright mode, which is also the only mode that needs `uv run playwright install`.

The consequence is that the credential here is your account rather than a token, and the first minute of every run is interactive. That is precisely what lowers the risk of platform-side detection, and it is also why this is not something to point at a schedule without reading the section below.

## The disclaimer is the first substantive thing in the file

Before any project description there is a disclaimer, and it is unusually specific. Use the repository for learning. All content is for study and reference. Commercial use is prohibited. Nobody may use it for illegal purposes or to infringe the rights of others. The crawler techniques are for learning and research, and must not be used for large-scale crawling of other platforms or other illegal behaviour. The project accepts no liability for legal responsibility arising from your use, and using the repository means accepting those terms. It links to a separate repository of illegal crawler cases as its illustration.

Two details sharpen this. A `LICENSE` file sits at the top of the tree, yet GitHub's own classification of the project comes back as NOASSERTION rather than a recognised identifier, so automated compliance tooling will not tell you what the grant actually is. And a prominent block above the disclaimer is a paid sponsor, a commercial scraping service that advertises stealth browsing, CAPTCHA handling and residential proxies.

The consequence is that a README-level restriction on commercial use and a LICENSE file whose terms you must read are two different things, and neither is machine-readable. Anyone considering this for work, as opposed to study, has to resolve that before writing a line of configuration.

## The open version cannot resume a crawl, and the paid one can

The README devotes a long section to MediaCrawlerPro, a separate subscription product, and the comparison is the most useful summary of the open version's limits. Listed as headline Pro features are resumable crawling, multi-account plus IP proxy pool support, removal of the Playwright dependency, full Linux environment support, a refactored codebase with the JavaScript signature logic decoupled, a desktop video downloader, homepage feed recommendations, and AI agent skill support.

The first two matter most for anyone with a real crawl in mind. In the open version, an interrupted run starts over, and there is no multi-account rotation, so all traffic comes from the single profile the CDP browser is holding. The project's own feature matrix lists login state caching and an IP proxy pool as available, which is true, but the proxy pool does not give you the account rotation that would make large parallel collection practical.

The consequence is a direct trade. You get the seven-platform matrix, comment export and word cloud generation for free, and you pay, or accept the risk, for the two capabilities that make long craw survivable. Read the Pro section as a specification of what the open version deliberately leaves out.

## A uniform seven-by-seven matrix hides uneven media support

The feature table is the strongest thing in the README: seven platforms, each with a checkmark for keyword search, crawling a specified post ID, secondary comments, a specified creator's homepage, login state caching, an IP proxy pool, and generating a comment word cloud image. The platforms are Xiaohongshu, Douyin, Kuaishou, Bilibili, Weibo, Tieba and Zhihu. Media download is off by default and enabled with `--get_media true` or `ENABLE_GET_MEDIA = True`, writing to `data/xhs/media/{post ID}/` with a cover, a video and numbered images per post.

The matrix is misleading in one place, and the media section is honest about it. Media download covers five platforms. Tieba and Zhihu have no media fields in their data structures and are not supported, which means a uniform-looking table hides a real gap for anyone whose target is Zhihu. Bilibili also has its own path: install ffmpeg and it takes the DASH route for the highest quality, merging separate audio and video tracks, and without ffmpeg it falls back to an mp4 direct link named `video-durl.mp4`. Quality is set by `BILI_QN` in `config/bilibili_config.py`, defaulting to 80 for 1080p, and steps down automatically when not logged in or not permitted.

Two more behaviours are worth knowing before a long run. Media download failures are logged and do not interrupt the crawl, existing files are skipped, and downloads are serial rather than concurrent specifically to avoid pressuring platform CDNs.

## Two dependency manifests that disagree, and a Chinese package mirror by default

There are two dependency files. `pyproject.toml` names the project mediacrawler, version 0.1.0, authored by 程序员阿江-Relakkes, and requires Python 3.11 or newer, with a long list running from fastapi and playwright through pandas, sqlalchemy, motor, redis, jieba, wordcloud, matplotlib, opencv-python and pyexecjs. `requirements.txt` lists a shorter, slightly different set. The two do not match: `asyncpg`, `websockets` and `pre-commit` appear only in the pyproject list, and opencv-python is pinned there but bare in requirements.

The recommended install path is uv, and the README says so while marking the plain venv route as not recommended:

```shell
cd MediaCrawler
uv sync
```

One line in that file is easy to miss and affects everyone outside China. The pyproject configures a default package index pointing at `https://pypi.tuna.tsinghua.edu.cn/simple`, so a plain `uv sync` resolves from the Tsinghua mirror unless you override it. Node.js 16 or newer is also required, and the venv notes say it is needed specifically when crawling Douyin and Zhihu, which points at the JavaScript evaluation the signature logic depends on.

The consequence is that `uv sync` and `pip install -r requirements.txt` give you different environments from the same repository, the mirror choice is baked in rather than left to you, and the README warns that the requirements are built against Python 3.11.

## The WebUI is a second process on a second port

There is a browser interface, and it is not a single command. In development you run the FastAPI backend and the Vite dev server separately:

```shell
uv run uvicorn api.main:app --port 8080 --reload
```

The frontend then runs in a second terminal from the `webui` directory with `npm install` and `npm run dev`, starting on port 5173 and proxying `/api` to 8080. For production you build the frontend with `npm run build`, which outputs into `api/webui/`, after which the API server alone serves the interface on 8080. The Node side of the repository is only this: the root package.json contains three VitePress documentation scripts and nothing else.

The first page load calls `/api/env/check` to verify the environment, and the README notes there is a button to skip the check if it fails. That is worth pausing on, because it means the interface will let you proceed with an environment it has just failed to validate.

The consequence is that the WebUI is a development-mode arrangement with a proxy rule, not a packaged application, and a skipped environment check trades a useful early failure for a confusing later one.

## recv_sms.py and the .env.example show what a run really costs

Two files at the top of the tree describe the operational footprint better than the feature matrix does. `recv_sms.py` sits in the repository root, and login is described in the commands as opening the corresponding app and scanning a QR code, with a `--lt qrcode` flag on the run command, which points at a family of login types rather than a single mechanism.

Then there is `.env.example`, which is configured for four databases at once, MySQL on 3306, Redis on 6379, MongoDB on 27017 and PostgreSQL on 5432, plus three commercial proxy vendors: Wandou HTTP, Kuaidaili and JiSu HTTP, each with its own key, signature or credential fields. The example values are placeholders, and they are the weakest possible placeholders, with `123456` as a password for MySQL, Redis and PostgreSQL and `root` or `postgres` as the user.

The consequence is that a realistic deployment is a browser profile with a logged-in account, at least one database, and possibly a paid proxy subscription, and the example file will happily start you with a trivial root password if you copy it without editing. The remaining directories, `store/`, `database/`, `cache/`, `proxy/`, `model/`, `media_platform/`, `cmd_arg/` and `const` under `constant/`, are the same story told in directory names.

## Conclusion

MediaCrawler is a well-built study of how far you can get by reusing a real browser session instead of reimplementing a platform's signing scheme, and the seven-platform matrix with login caching, proxy pools and comment export is more complete than most projects in this space. Three things decide whether it is appropriate for you. It runs as your logged-in account, so the risk profile is your account's. The open version cannot resume a crawl, and resumable crawling, multi-account rotation and the removal of the Playwright dependency are all listed as paid Pro features. And the project's own disclaimer restricts use to learning and forbids commercial use and large-scale crawling. Before running it, verify three things: the disclaimer terms against your situation, whether a crawl of your size survives an interruption, and which platform data structures you actually need, since media download covers five of the seven platforms.

## FAQ

### Is MediaCrawler legal to use?

The project's own disclaimer says to use the repository for learning, prohibits commercial use, and states that the crawler techniques must not be used for large-scale crawling of other platforms or for other illegal behaviour. It also disclaims liability for legal consequences. A LICENSE file exists at the root of the repository, though GitHub classifies the project as NOASSERTION rather than a recognised licence identifier, so read the file yourself before any use that is not study.

### How do I install and run MediaCrawler?

Install uv, then run `cd MediaCrawler` and `uv sync` in the project directory. Node.js 16 or newer is required. A typical run is `uv run main.py --platform xhs --lt qrcode --type search` for a keyword search or `--type detail` for specified post IDs, and `uv run main.py --help` lists the other platforms. Comment crawling is off by default and is enabled with `ENABLE_GET_COMMENTS` in `config/base_config.py`.

### Does MediaCrawler need to solve CAPTCHAs or reverse engineer signatures?

The README states the opposite approach. It attaches to a Chrome instance you already have open over CDP, reuses the existing login state and cookies, and obtains signature parameters by evaluating JavaScript inside that already-authenticated page context, so no reverse engineering of the encryption algorithm is required. In the default CDP mode you also do not need to install a browser driver; `uv run playwright install` is only for standard Playwright mode.

## Sources

- [Issues](https://github.com/NanmiCoder/MediaCrawler/issues)
- [NanmiCoder/MediaCrawler on GitHub](https://github.com/NanmiCoder/MediaCrawler)
- [Project website](https://nanmicoder.github.io/MediaCrawler/)
- [README](https://github.com/NanmiCoder/MediaCrawler/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/nanmicoder-mediacrawler
