Model or dataset
cv-cat/Spider_XHS avatar
cv-cat/Spider_XHS

Spider_XHS: a Python client library that reverses Xiaohongshu request signing

小红书爬虫数据采集,小红书逆向,私信,直播,小红书全域运营解决方案

7,789 stars1,348 forksPythonLicense varies

At a glance

What is it?
A collection of API clients for the Xiaohongshu PC web app, the creator platform, the Pugongying KOL marketplace and the Qianfan distribution platform, with local signature computation so a script can read and write notes without a browser.
Who is it for?
Spider_XHS is a signing and API layer rather than a crawler, and that framing explains the codebase. The hard problem is not walking pages, it is producing the `a1`, `b1`, `x-s`, `x-t`, `sign` and `q-signature` values that the platform expects, so the repository spends its weight on auth objects, a small Node runtime for signature math and per-platform API classes.
Can I use it commercially?
Not without permission. GitHub finds no licence file in the repository, and without a licence all rights are reserved by default: you may read the code but not reuse it. Check the README, or ask the authors, before using it.
Is it still maintained?
Yes. The repository last received commits 29 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 21, 2026, and from our analysis. They are not legal advice.

Editorial analysis

Why the project exists: no official content API

The README's framing is blunt about the reason. Xiaohongshu does not offer a complete content operations interface, so anyone wanting to build an automated workflow around collecting notes, rewriting them with a model and publishing the result has to do platform-level plumbing first. The repository positions itself as that plumbing: you supply the model, it handles reading and writing platform data.

The diagram the README uses makes the intended pipeline clear, from collecting competitor notes through the library into an AI agent that rewrites or analyses, and then out to an automatic publish, with account management looping back. The project description is broader again, covering data collection, reverse engineering, private messaging, live streaming and a full-domain operations approach.

The technical claim at the centre is signature reconstruction. The README says the project reverses the signing algorithms used by the Xiaohongshu PC web app and the creator platform, and it names the parameters involved: `a1`, `web_id`, `b1`, `websectiga`, `sec_poison_id`, `gid`, `x-s`, `x-t`, `x-s-common`, `x-b3-traceid`, `x-xray-traceid`, `x-rap-param`, `search_id`, `request_id`, `sign` and `q-signature`. Once those are generated locally, the core HTTP interfaces are wrapped with the signing already handled, which is why the library can run without a browser at all.

One note on scope before going further: the README carries an explicit warning that the project is for study and exchange only and that commercial use is prohibited. That is the author's stated boundary, and it is worth respecting both on legal grounds and because platform terms change independently of this repository.

Four surfaces covered: PC web, creator, Pugongying and Qianfan

The feature table in the README is organised by platform, and that structure mirrors the code layout. The tree has `apis/`, `spider/` and `xhs_utils/`, and the import paths give the same shape: `apis.xhs_pc_apis`, `apis.xhs_creator_apis` and `apis.xhs_pugongying_apis`.

For the PC web app the covered operations are QR code login and phone verification login, reading a profile's channels and recommended notes, fetching user profiles and your own account details, listing published, liked and saved notes, retrieving full note content including watermarked-free images and video, searching notes and users, reading comments, and handling unread messages, comment mentions, likes, saves and new follows. The live and private messaging surface includes connecting to a live room and listening to its events, with danmaku-style comments, likes, join and gift events, plus sending and receiving private messages over WebSocket with an HTTP fallback.

The creator platform side covers the same two login methods, session-level automatic retry behind what the README calls a 406 probability gate, uploading image gallery posts, uploading video posts with a transcode polling loop, listing published work, and computing the Creator RAP publish signature locally with no server call.

The two commercial surfaces are read and outreach rather than publishing. Pugongying, the KOL marketplace, returns blogger lists and detailed data, fan profile breakdowns and historical trends, and can initiate a collaboration invitation. Qianfan, the distribution platform, returns reseller lists with details on collaboration categories, shops and products. That split is the practical shape of the library: heavy on the consumer app, complete on your own publishing, and read-mostly on the business platforms.

Getting it running without a browser, and the three login modes

Requirements are Python 3.10 or newer and Node.js 20 or newer, which is a strange pairing for a Python project until you look at `package.json`, where the single dependency is `crypto-js` and the description reads XHS API clients and local PC signing runtime. Signature computation needs a JavaScript engine, and `PyExecJS` in `requirements.txt` provides one, with Node supplying it.

Installation is two commands, one per ecosystem:

bash
pip install -r requirements.txt
npm install

The rest of `requirements.txt` is a reasonable picture of what the client actually does: `requests`, `curl_cffi` pinned at 0.15.0 for browser-fingerprinted TLS, `loguru` for logging, `python-dotenv`, `retry`, `openpyxl` for the saved Excel output, `aiohttp` for the async live and messaging paths, `opencv-python` and `numpy`, and `qrcode` for rendering login codes locally.

Login is configured in `spider/spider.py` with a single setting, and the project explicitly does not drive a browser:

python
login_type = 'cookie'  # cookie / qrcode / phone

`qrcode` fetches a QR code through local requests for scanning with the Xiaohongshu app, `phone` calls the SMS verification login interface directly, and `cookie` takes a full Cookie string copied after logging in elsewhere. Only the cookie mode needs a `.env` file, copied from `.env.example`, which contains a single `COOKIES` entry.

The PC side routes everything through `XHSPcAuth`, which manages login state, `b1`, DS, MNS environment material and session counters. Its three factory methods all return a bootstrapped PC login state:

python
from apis.xhs_pc_apis import XHS_Apis
from xhs_utils.xhs_pc import XHSPcAuth

That single object is where the signing work lives, which is why the rest of the API classes read as thin wrappers over HTTP calls.

The agent integration patterns the README demonstrates

The README walks through three usage patterns, and they are the fastest way to understand the API surface without reading the source. The first is the full loop: build a PC auth and a creator auth from cookies, collect a competitor note, hand its content to any model, then post the rewritten result to the creator platform.

python
from apis.xhs_pc_apis import XHS_Apis
from apis.xhs_creator_apis import XHS_Creator_Apis
from xhs_utils.xhs_pc import XHSPcAuth
from xhs_utils.xhs_creator import XHSCreatorAuth

The first pattern bootstraps both auth objects, calls `get_note_info` with a note URL and gets back a success flag, a message and the note payload. The rewrite step is left as a placeholder for whatever model you prefer, with GPT, Claude, Qwen and local models all named as options, and the publish step calls `post_note` with a title, description, media type and image list.

The second pattern is monitoring rather than publishing: search notes for a keyword with a requested count and hand the result set to a model for trend analysis. The third is KOL selection, using `PuGongYingAPI` to fetch a batch of bloggers for a category and score them against a brand profile.

The README also documents a skills-based integration path alongside the Python API. A separate XhsSkills repository holds Agent Skills packaged from this project, and it names Clawbot, Claude Code and Codex as tools that can consume skills directly. That is a meaningful addition because it means the platform client can be driven by an agent toolchain rather than only by a script you write yourself.

A larger product built on top, sold separately

Before describing the repository itself, the README describes a finished product: XHS_ALL_IN_ONE. It is a multi-account matrix with QR login, SMS verification and cookie import, encrypted cookie storage and a health check every two hours with expiry notification. It adds AI image retouching that takes a reference image and an instruction, generates a revised image and swaps it into the draft in place, with the original and the result side by side. And it has a publish centre that previews the draft, picks the creator account, sets visibility and publishing mode for immediate or scheduled posting, and publishes to the creator platform after validation passes.

That product is presented as the built-on result, and the distinction is worth keeping clear when judging this repository. What is open here is the client library and the login layer; the account matrix, the image pipeline and the web interface live elsewhere.

The repository scale is the other useful signal. Around 7,800 stars and 1,350 forks, Python, with 123 open issues. That issue count is high relative to the star count and is the clearest external hint at how much churn the platform causes.

Release history is short and terse. v5.0.0 on 2026-09-07 has notes that amount to private messaging and live streaming, a tag named `xhs_new_api` carries v4.0.0 from 2026-04-15, and v3.0.0 from 2026-03-19 is tagged `xs_xt_26_3` with the body code version. The `xs_xt_26_3` naming pattern is itself informative, since it reads like a pointer to a signature token schema that has to track a platform change.

Docker path, data output and where the documentation stops

A `Dockerfile` sits at the root and is short enough to read in full. It starts from `python:3.10-slim`, installs curl, gnupg, build-essential and git, then pulls Node 20 from the NodeSource setup script, prints all three versions, copies in `requirements.txt` and installs with pip using no cache, copies the source, exposes port 5000 and sets `PYTHONUNBUFFERED` and `NODE_ENV=production`. The entry point is the spider module run directly:

text
EXPOSE 5000
ENV PYTHONUNBUFFERED=1
ENV NODE_ENV=production
CMD ["python", "-m", "spider.spider"]

The port and production environment variable suggest an HTTP service component alongside the library, which fits with the live and private messaging paths needing long-lived connections, though the README does not document an API surface for port 5000 or describe the endpoints.

The documented outputs are screenshots and an Excel file, the latter explained by `openpyxl` in the requirements. That tells you scraped note data is meant to land in a spreadsheet rather than a database, which is a reasonable choice for one-off collection and a poor one for continuous monitoring.

So the documentation is thorough on installation, login configuration and the feature list, and thin on exactly the things a maintainer would want written down: what the HTTP service on 5000 exposes, how rate limits are handled beyond the 406 retry gate, how the KOL and Qianfan endpoints authenticate, and what the signing layer breaks on when the platform rotates a token. The README also documents a small `demo.py` in the tree without describing it, which is another gap worth knowing before you start reading code.

Editorial conclusion

Spider_XHS is a signing and API layer rather than a crawler, and that framing explains the codebase. The hard problem is not walking pages, it is producing the `a1`, `b1`, `x-s`, `x-t`, `sign` and `q-signature` values that the platform expects, so the repository spends its weight on auth objects, a small Node runtime for signature math and per-platform API classes. The feature table is broad, covering notes, search, comments, messages, live events, private messages, publishing, KOL data and distribution data, and the honest gap is stability: 123 open issues and a v5.0.0 release whose notes are two words is what upstream signature churn looks like from the outside. Read `spider/spider.py` for the login mode, start with one cookie, and keep the README's own warning about scope in mind.

Frequently asked questions

What is Spider_XHS and what platforms does it cover?

It is a Python client library for Xiaohongshu that computes request signatures locally so scripts can read and write platform data without a browser. It covers four surfaces: the PC web app for notes, search, comments, live events and private messages, the creator platform for uploading and publishing, the Pugongying KOL marketplace, and the Qianfan distribution platform.

What are the login options and how do I set one up?

You set `login_type` in `spider/spider.py` to one of `cookie`, `qrcode` or `phone`. QR mode fetches a code locally for scanning with the app, phone mode calls the SMS verification interface, and cookie mode takes a full Cookie string. Only cookie mode needs a `.env` file copied from `.env.example` with a `COOKIES` entry.

Why does a Python project also need Node.js installed?

Because the PC signature values are computed through a small local JavaScript runtime. `package.json` has one dependency, `crypto-js`, and describes the project as XHS API clients and local PC signing runtime, while `PyExecJS` in the requirements provides the bridge to Node. The Dockerfile installs Node 20 for exactly this reason.

Can the library publish notes automatically?

Yes, through the creator platform client, which supports uploading image galleries and video with transcode polling, and computes the Creator RAP publish signature locally. The README's first agent example collects a note, hands it to a model and calls `post_note` with a title, description, media type and images. The author's stated scope is study and exchange, with commercial use prohibited.

Official sources

  1. cv-cat/Spider_XHS on GitHub
  2. Issues
  3. README
  4. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/cv-cat-spider-xhs.svg)](https://hysenlabs.com/projects/cv-cat-spider-xhs)