Library / SDK
dataabc/weibo-search avatar
dataabc/weibo-search

weibo-search: getting a whole keyword, one hour at a time

获取微博搜索结果信息,搜索即可以是微博关键词搜索,也可以是微博话题搜索

2,328 stars430 forksPythonLicense varies

At a glance

What is it?
weibo-search is a Scrapy crawler that works around Weibo's fifty-page result cap by subdividing a date range into hours and then into minutes, and the README claims that recovers all or nearly all matching posts, up to a hundred million for a popular keyword. The algorithm is carefully documented and the operational hygiene is thin: no licence file, no statement of terms of service, and a full session cookie pasted into settings.py in plaintext.
Who is it for?
Use weibo-search if you need a reproducible archive of what a keyword produced on Weibo over a bounded date range and you accept the platform's terms, because the completeness argument is the only reason to choose this over a simpler scraper: the recursive subdivision is a real answer to the result cap rather than a workaround for a missing API.
Can I use it commercially?
Not without permission. GitHub finds no licence file in the repository, and without a licence all rights are reserved by default: you may read the code but not reuse it. Check the README, or ask the authors, before using it.
Is it still maintained?
Yes. The repository last received commits 118 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The subdivision algorithm is the real engineering

The feature that makes this crawler worth reading is FURTHER_THRESHOLD, and the explanation is seven sentences of the clearest documentation in the repository.

FURTHER_THRESHOLD is the threshold at which the program searches further. The problem it solves is that Weibo's search caps at fifty pages of results. When a search condition returns a lot of results there should be about fifty pages of posts, and more than fifty is not displayed. So the program treats a page count equal to the threshold as evidence that the results were truncated, and responds by subdividing the search.

The subdivision is recursive and time-based. If the current search is by day, the program splits that one search into twenty-four hourly searches. If the hourly search again comes back at the threshold, it subdivides again, and again after that, so a day eventually becomes hours, then minutes. The README says this continues in the same way, 以此类推.

That is a clean and correct answer to a pagination cap, and it is worth being clear about why it is not simply rate limiting or persistence. A crawler that just asks for pages 51 to 500 gets nothing, because the cap is on the result set rather than on the request. The only way past it is to make the result set smaller, and the only axis that reliably shrinks it without dropping matches is time. So a day-long window with a hundred thousand hits becomes twenty-four hour-long windows, and each of those is small enough to fit under the cap.

The subtlety, and this is what the threshold parameter is for, is that hitting the cap is not proof of truncation. The README says some keywords, even popular ones, only ever show forty-odd pages. So a threshold of fifty would treat a genuinely complete forty-two-page result as truncated and subdivide unnecessarily, and a threshold that is too low subdivides everything, including results of a few pages, which slows the whole crawl for no benefit.

The recommendation follows from the two failure modes: the threshold should be a number smaller than fifty, and the README suggests setting it between forty and forty-six. The example configuration is FURTHER_THRESHOLD = 46.

This is the part of the project that is properly engineered, and it is the part a reader should take away. The rest of the crawler is configuration surface, but this one parameter is a genuine algorithmic contribution to a real problem, it is documented with both of its failure modes, and the recommended range is justified rather than guessed.

Ten million posts claimed, and a date field with day precision

The scale claim in this README is extraordinary and the output schema quietly contradicts it.

The claim is in the first section. Search results are enormous: for a very popular keyword, within a one-day time range you can obtain more than ten million search results. Note that one day here means the time filter range, and how long it takes to actually download those ten million depends on the speed. Ten million is only what one day of range can yield, and if you want more you increase the range: ten days gives at most ten million times ten, which is a hundred million results, and you can go wider still. The README then makes the general claim: for most keywords, related search results in a day should be below ten million, so this program can obtain all or nearly all of the search results for a specified keyword.

Now the output specification. The publish-time field is described as the time the post was published, 精确到天, that is, precise to the day.

So the crawler subdivides internally down to minutes, in order to work around the pagination cap, and then writes out a date that has been thrown away to the day. Every one of those hundred million rows carries a day but not a time.

That is a real and consequential gap for the use the tool is built for. If you are archiving what a keyword produced, chronological ordering within a day is not something you can reconstruct, and for a platform where recency drives visibility, the hour a post appeared in is often the interesting part. You also cannot align the archive against anything that has finer resolution.

The gap is easy to fix and is not fixed, which suggests it is an oversight rather than a design decision, and it is worth noting because a hundred million rows with day-only timestamps is a lot of storage for a lot of lost information.

Two other output details are worth reading closely. The original-post id field is specific to reposts and holds the id of the post that was reposted, and the README says that post is also stored with the same fields except that this field is empty for it. So the crawler follows the repost chain and stores both the repost and its original. For a keyword where most posts are reposts, that roughly doubles the row count, and the de-duplication pipeline that ITEM_PIPELINES lists first is presumably what stops the duplicates.

The separator conventions are inconsistent in a way that will bite anyone parsing the CSV. Multiple image URLs are separated by an English comma, and multiple video URLs by an English semicolon. Topics and at-mentioned users also use commas. So a single CSV column can use two different delimiters depending on whether it holds images or video, and a naive split on comma will not parse the video field and a naive split on semicolon will not parse the images.

The other field that deserves attention is the publishing client, which records the device or app the post was made from, with the examples given being an iPhone client and a HUAWEI Mate 20 Pro. Combined with the user_authentication field, which classifies each account as a blue-V, yellow-V, red-V, gold-V or ordinary user, that is a fairly detailed profile of who posted and from what, for every row, which is exactly the kind of data whose handling deserves a deliberate decision.

Your session cookie goes into settings.py as plaintext

The setup instructions are fourteen numbered steps, and step four is the one with a lasting consequence.

It says the cookie in DEFAULT_REQUEST_HEADERS is the value you need to fill in, and to replace "your cookie" with the real cookie once you have obtained it. All configuration for the program is done in setting.py, located at weibo\settings.py.

So the mechanism is a full logged-in session cookie for a Weibo account, stored as a literal string in a Python configuration file in the working tree. The README then devotes a section to how to get that cookie: open weibo.com in Chrome, click the login button, complete the private-message verification or the SMS code verification, get into the new Weibo, press F12, and in the developer tools go to Network, find the weibo.cn entry, look at Request Headers and copy the value after Cookie. A second, compatibility section gives the same procedure for the older Weibo interface through passport.weibo.cn.

The instructions are precise and a reader can follow them. What is missing is any guidance on what to do with the value afterwards.

There is no .gitignore mention of settings.py in what is shown, and the repository's .gitignore exists but its contents are not documented. So a user who initialises a git repository, copies the settings file in, and commits it will commit a live session credential. That credential is not a password: it is a bearer token that authenticates as the account until it expires, and on a social platform it is enough to read and post as that user. It is the kind of value that belongs in an environment variable or a local file that is explicitly ignored, and the project's configuration model, which is a Python file with literal assignments, makes that harder than it needs to be.

The cookie instructions also compound the exposure in a small way. They direct the user to log in interactively, complete a verification challenge, and then extract the resulting session from devtools. That is a normal way to obtain a session cookie and there is nothing improper about it for one's own account, but it does mean the tool's setup requires a person to be actively logged in and then to hand the credential to a script.

The right mitigation is not difficult: keep the cookie in an environment variable, read it in settings.py, and add settings.py to .gitignore. That is a small change to a project that already has good documentation for the hard part, and its absence is the gap.

No licence file, and no statement of terms of service

Two absences in this repository, one of which will stop an adoption outright and one of which is a judgement call.

The first is the licence. The repository's licence field reads as unknown, and the top-level listing is .gitignore, README.md, requirements.txt, scrapy.cfg and weibo/, with no LICENSE file among them. So the code is under default copyright. That means you may read it, and you may not copy it, modify it, redistribute it or incorporate any of it into a project of your own without permission from the author. For a public repository that many people will find and want to use, that is the most consequential omission in it, and it is a one-file fix.

The second is terms of service, and here the judgement is more subtle.

The function of this program is bulk collection from a social platform. It authenticates with a real account, walks through search result pages with a ten-second delay, downloads post text, images and videos, and is explicitly designed to work around the platform's result cap so that it can retrieve the complete set rather than the truncated set. There is no mention anywhere in the README of Weibo's terms of service, of the legal position on automated collection, of the rights of the people whose posts are being collected, or of personal data handling.

None of that makes the tool wrong to write. Web scraping is a legitimate activity, the code is public, and the person running it is responsible for their own use. But a README that helps you configure a collector and says nothing about whether you may is leaving the most consequential decision to the reader without telling them there is a decision to make.

The rate limiting deserves its own note. DOWNLOAD_DELAY is the wait between finishing one page and starting the next, defaulting to ten seconds, and the README shows how to change it to fifteen. Ten seconds is a substantial politeness interval, far slower than a naive scraper would be, and whoever wrote this clearly was not trying to hammer the platform. But a delay is a rate limit, not a permission, and the difference matters: one is a technical courtesy and the other is a legal position, and only one of them is a decision the user can make for themselves.

The personal data dimension is also unaddressed. The output includes the account verification tier for every post, so blue-V and gold-V verified accounts are distinguishable from ordinary users, and it includes the exact publishing device, so the same account can be tracked across an iPhone and a specific Huawei model. Aggregated over a hundred million rows, that is a dataset about identifiable people, and the README does not say what a user of it should consider.

None of this is a reason to skip the project. It is a reason to read it with the operational and legal questions in view, and to notice that the repository is a good tutorial in one specific technique and thin in every other respect.

Four requirements, no database driver, and pytest in the runtime list

The requirements file is four lines long, and comparing it to the feature list is instructive.

The four entries are Scrapy>=2.11,<3, requests>=2.31,<3, Pillow>=8.1.1, and pytest>=8,<9.

Two of those are upper-bounded and one is not. Scrapy and requests are both capped below version 3, which is the responsible thing to do for a project that has been around a while and does not want to break on a major release of its two most important dependencies. Pillow has a floor of 8.1.1 and no ceiling, and pytest has a floor of 8 and a ceiling below 9. So two entries are pinned at both ends, one at the bottom, and one at both ends. That inconsistency is small and worth noting because it is the kind of thing that means the file was assembled incrementally rather than generated.

The bigger finding is what is missing. The README describes four output targets: a CSV file as the default, and MySQL, MongoDB and Sqlite as options, plus optional image and video downloads. There is no database driver in requirements.txt. Not mysql-connector, not pymongo, not psycopg, nothing. The README does address this in passing, saying that if you want to write to a database you need to fill in the relevant configuration in setting.py, with MONGO_URI for MongoDB and the MYSQL-prefixed settings for MySQL. It does not say you also have to install the driver yourself.

So the SQLite path is the only database target that works from a clean install, and the README says as much, noting that Sqlite needs no external installation and is more convenient than MySQL and MongoDB. That is an honest and correct observation, and it is presented as a convenience rather than as the only option that will work. A user picking MySQL will discover at run time that a package is missing.

The last entry is the one that is simply miscategorised. pytest is a test framework, and it belongs in a development requirements file, not in the file the README tells you to install with pip install -r requirements.txt as step three. Its presence there means everyone who wants to run the crawler also installs a test runner, which is harmless and slightly untidy.

The dependency list is otherwise a good sign for the project's age. Scrapy at a 2.11 floor with a sub-3 ceiling is a current, maintained Scrapy, not a fossilised 1.x pin, and requests at a 2.31 floor is similarly current. So the tool is not depending on abandoned versions, which for a scraper that talks to a website's search interface is the more important thing to check, because the framework underneath is what determines whether the code still runs at all.

The rest of the repository is minimal in a way that suits the tool. scrapy.cfg at the root is Scrapy's project configuration, and weibo/ is the single Python package holding the spider, the settings, the pipelines and, per the README's mention of a region.py file, the list of supported province and municipality names.

Configuration is a Python file, and the docs assume Windows

Every knob in this crawler is a Python assignment in settings.py, and the fourteen-step setup walks through them in order. The list is worth reading as a description of the API, because it is the entire surface.

KEYWORD_LIST takes a single keyword string, a list of keyword strings searched separately, a list containing one space-separated string to require all of them, a hashtag-delimited string to search a topic, or the path to a text file with one keyword per line. So the AND-search form is a single string with a space in it rather than a nested structure, which is compact and slightly ambiguous if a keyword legitimately contains a space.

START_DATE and END_DATE are yyyy-mm-dd strings and the range is inclusive of both boundaries. The README notes that the publish time is only recorded to the day, which is consistent with a day-granular input.

WEIBO_TYPE filters by post type: zero for all, one for original posts only, two for hot posts, three for posts from people you follow, four for verified users, five for media accounts, six for opinion posts. CONTAIN_TYPE filters by required content: zero for no filter, one for posts with images, two with video, three with music, four with a short link. REGION filters by province or direct-administered municipality, and the instructions are unusually precise about the format: the value should not include the characters for province or city, so use 北京 rather than 北京市 and 安徽 rather than 安徽省, multiple regions can be given, only province and municipality level names are supported, city names below a province and district names below a municipality are not, and 全部 means no filter. The supported names are in region.py.

DOWNLOAD_DELAY is the seconds between pages, ten by default. ITEM_PIPELINES is an ordered list where the numbers are execution priorities: first de-duplication, then the CSV writer, then MySQL, then MongoDB, then image download, then video download, and the README's instruction is to comment out the ones you do not want in order to save resources.

The one genuinely fiddly operational instruction is the last one. To save progress so the crawl can resume, you add a JOBDIR argument to the crawl command rather than running a bare scrapy crawl search. To stop safely you press Ctrl+C once, and the README bolds that word twice in the same sentence:

bash
scrapy crawl search -s JOBDIR=crawls/search

That is the whole run command, and the JOBDIR argument is the whole of the resumability mechanism. After one press the program keeps running for a while, saving fetched data and progress, and you are asked to be patient. If the resume produces nothing, the likely cause is progress that was not saved correctly, and the fix is to delete the progress file inside the crawls folder and run the command again.

That is a real gotcha. A crawler that needs a single interrupt to shut down cleanly is asking the user to distinguish one Ctrl+C from two, and the documentation compensates with emphasis rather than with a signal handler. It is the kind of rough edge that a user hits once and remembers.

One last documentation detail. The path to the settings file is given with Windows separators, weibo\settings.py, in a project that is otherwise written for a cross-platform Python tool. It is a small inconsistency in a README that is otherwise thorough, and it is the kind of thing that suggests the documentation was written on one machine for a long time.

Editorial conclusion

Use weibo-search if you need a reproducible archive of what a keyword produced on Weibo over a bounded date range and you accept the platform's terms, because the completeness argument is the only reason to choose this over a simpler scraper: the recursive subdivision is a real answer to the result cap rather than a workaround for a missing API. Do not use it on a repository or team where the absence of a licence file is a blocker, since there is no LICENSE in the listing and the default position on undeclared copyright is that you may read the code and not redistribute or incorporate it. Do not put a production Weibo account's cookie in it, because the session cookie is a full logged-in credential written into a plaintext settings file with no guidance on keeping it out of version control. Verify four things. Read the ToS position yourself, because the README does not mention terms of service at all and the only rate limiting is a ten-second delay between pages. Decide whether day-precision dates are enough, because the output field is precise only to the day even though the crawler subdivides to minutes internally. Check what you will do with the personal data, since the output includes the posting device string and the account verification tier for every post. And set FURTHER_THRESHOLD to something in the recommended forty-to-forty-six range, because a value of fifty is the default mental model from the cap and it is exactly the value that silently loses posts. The deciding fact is that the collection technique here is genuinely good and the surrounding project governance is not, so the risk in adopting it is legal and operational rather than technical.

Frequently asked questions

How do I install and run weibo-search?

Clone the repository, install Scrapy with pip install scrapy, then install the remaining dependencies with pip install -r requirements.txt. Configure the cookie in DEFAULT_REQUEST_HEADERS in the weibo/settings.py file, set your keywords in KEYWORD_LIST and a date range with START_DATE and END_DATE, then run scrapy crawl search. Adding -s JOBDIR=crawls/search saves progress so the crawl can resume.

What does FURTHER_THRESHOLD do?

It is the page count at which the crawler assumes Weibo truncated the results. Weibo's search stops at fifty pages, so a result set sitting at the threshold is treated as incomplete and the search is subdivided by time: a day becomes twenty-four hours, and an hour subdivides again, recursively. The README recommends a value between 40 and 46, because 50 would stop early on keywords that genuinely only have forty-odd pages.

What licence is weibo-search released under?

None that is declared. The repository's licence field reads as unknown and there is no LICENSE file among the top-level entries, which are .gitignore, README.md, requirements.txt, scrapy.cfg and weibo/. Under the default position on undeclared copyright you may read the code but not copy, modify or redistribute it without permission from the author.

What data does weibo-search collect?

Post id, bid, post text, article url, image urls, video urls, publish location, publish time precise only to the day, like, repost and comment counts, the publishing client such as an iPhone or a specific Huawei model, hashtags, at-mentioned users, the original post id for reposts, and the account verification tier, which is blue-V, yellow-V, red-V, gold-V or ordinary user. Images and videos can optionally be downloaded alongside.

Does weibo-search need any database drivers installed?

It lists none. requirements.txt contains only Scrapy, requests, Pillow and pytest, so SQLite is the only database target that works from a clean install, which the README notes as more convenient. For MySQL and MongoDB you have to install the driver yourself in addition to filling in the configuration in settings.py, where MONGO_URI holds the MongoDB settings and the MYSQL-prefixed entries hold the MySQL ones.

What keyword forms does weibo-search support?

KEYWORD_LIST accepts a single keyword string, a list of strings searched separately, a single string with spaces to require all of those words, a hashtag-delimited string such as #keyword# to search a topic, or the path to a text file with one keyword per line. Search can be further filtered by post type with WEIBO_TYPE, by required content with CONTAIN_TYPE, and by province or direct-administered municipality with REGION, where the name must be given without the province or city suffix.

Official sources

  1. dataabc/weibo-search on GitHub
  2. Issues
  3. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/dataabc-weibo-search.svg)](https://hysenlabs.com/projects/dataabc-weibo-search)