Open-source project
cxyfreedom/website-hot-hub avatar
cxyfreedom/website-hot-hub

website-hot-hub: hourly archives of Chinese platform trending lists

36Kr bilibili GitHub 2023-10-25 Releases .

414 stars60 forksPythonMIT

At a glance

What is it?
A Python scraper that records the 36Kr, bilibili, GitHub, Douyin, Juejin, WeRead and Kuaishou hot lists once an hour and commits them as dated markdown. It is a data-collection tool, not a dashboard.
Who is it for?
Adopt it if you need a self-hosted, plain-text record of Chinese platform trending lists and you are willing to keep the per-site scrapers working yourself. Do not adopt it if you want a live dashboard, an API, or a guaranteed daily history, because the README documents no scheduler and the archives only exist where someone ran the script.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What website-hot-hub records, and who needs a record like this

website-hot-hub is a scraper, not a product with a user interface. Its README states the purpose plainly: it records the hot lists of 36Kr, bilibili, GitHub, Douyin, Juejin, WeRead and Kuaishou from 2023-10-25 onward, fetching once per hour and archiving by day. The interesting word there is archive. Most trending-list tools answer the question of what is popular right now, and that answer expires within hours. This project answers a different question: what was on the list at a particular hour on a particular day, kept in a form you can grep.

The audience is narrow and fairly technical. Anyone doing retrospective analysis of Chinese tech and entertainment coverage needs the list as it stood, not as it looks after a week of editing and reordering. A researcher tracking how a story moved between 36Kr and Juejin, or how a repository entered the GitHub list, needs dated snapshots. A journalist reconstructing what was visible on a given morning needs the same thing. If you only want to read today's list, this repository is more machinery than the task requires.

The scope is defined by the modules present at the top level. There is a website_36kr.py, website_bilibili.py, website_douyin.py, website_github.py, website_juejin.py, website_kuaishou.py, website_sspai.py and website_weread.py. That is the whole coverage. There is no generic plugin interface documented in the README, so adding a ninth platform means writing a module that matches whatever convention the existing ones follow.

How the hourly fetch and the dated archive fit together

The repository layout tells most of the story. main.py sits at the top level alongside utils.py and the per-site modules. There is a raw/ directory and an archives/ directory, plus a template/ directory and a README.md that the project evidently rewrites as new data arrives. The README embeds generated blocks between HTML comment markers, for example an opening marker for 36Kr and a closing one, with a last-updated timestamp inside. That is the mechanism: the scraper writes both a machine-readable record and a human-readable section of the README.

The data flow is one direction. A per-site module fetches a page, parses it with requests and BeautifulSoup, and returns a ranked list of titles with links. utils.py holds the shared work: writing files, formatting dates, updating the README markers. The archives directory is organised by day, which is what makes the project useful after the fact. The raw directory appears to hold intermediate output, though the README does not document its format.

Two dependencies are pinned in requirements.txt: requests==2.32.4 and beautifulsoup4==4.12.3. That is a deliberately small footprint, and it also tells you what the parsing strategy is. There is no headless browser in the dependency list, so every module works against server-rendered HTML. Sites that render their lists client-side cannot be handled by this approach without adding a browser dependency, and the README does not describe any such fallback.

Installing website-hot-hub and running the first fetch

The README does not include an installation section, so the steps below follow from the repository files rather than from documented instructions. The project is a plain Python script with two third-party dependencies, so setup is short.

Clone the repository and install the pinned dependencies. The requirements.txt pins requests and beautifulsoup4 to specific versions, which matters because the parsers depend on the shape of the HTML those versions return.

bash
git clone https://github.com/cxyfreedom/website-hot-hub.git
cd website-hot-hub
pip install -r requirements.txt

After this the interpreter should be able to import requests and bs4 without error. If you use a virtual environment, the project does not document one, but nothing in the layout prevents it.

The entry point is main.py. The README does not document any command-line flags, so the safest assumption is that running it with no arguments performs a fetch of every configured platform and writes the results into raw/, archives/ and the README blocks.

bash
python main.py

What you should see afterwards is a new dated entry under archives/ and updated sections in README.md between the platform markers, each carrying a last-updated timestamp in the form the README shows, for example a date and time followed by +0800. The timestamp is in China Standard Time, which is worth noting if your server runs on UTC: the day boundary in the archives follows Beijing time, not your local clock.

If you only want one platform, the per-site modules are separate files, so calling into website_36kr.py or website_github.py directly is possible in principle. The README does not document a supported way to do that, so treat it as reading the source rather than following an interface.

Where website-hot-hub breaks, and when it is the wrong choice

The main limitation is structural: every module is a scraper tied to one site's markup. When 36Kr or Douyin changes its page structure, the corresponding module stops returning useful data, and there is no test suite or fixture in the repository layout to catch that before it reaches the archive. A silent failure is worse than a crash here, because an empty or partial list written into a dated file looks like a quiet news day rather than a broken parser.

The README documents no scheduler. The claim that data is fetched once per hour describes the intended cadence, not a guarantee that ships with the code. If you want hourly coverage, you supply the cron job or systemd timer yourself, and you accept that a machine outage means a permanent hole in the archive. There is no backfill mechanism described, and the archives directory is the only record.

The project is also not a dashboard. There is no server, no query interface and no API in the repository layout. The output is markdown and files on disk. Anyone expecting to point a browser at it and browse trends by date will be disappointed; they would need to build that layer on top of the archives.

Finally, consider the maintenance signal. The last push was on 2026-02-07, and the most recent release is v0.0.2, tagged as Archive-2025. The previous release, v0.0.1, carries the tag Archive-2023-2024. The release naming suggests the maintainer treats the work as periodic archival passes rather than continuous development, and the gap since the last push is consistent with that. Nothing here is archived, but this is not a project you should expect to receive a fix for a broken parser on short notice.

How it differs from a general scraping framework

The obvious alternative is a general-purpose scraping or scheduling framework such as Scrapy, or a hosted change-detection service. The difference is in what is fixed and what is flexible. A framework gives you a pipeline abstraction and expects you to define spiders, item schemas and export formats. website-hot-hub gives you the opposite: the platforms are hard-coded, the output format is decided, and the archive layout is already chosen for you.

That trade is real in both directions. With a framework you would write the per-site parsing yourself anyway, and you would still be responsible for the archive layout and the README generation, which is exactly the work this repository has already done for eight platforms. What you gain from a framework is retry logic, concurrency and middleware for handling blocks and rate limits. The dependency list here, just requests and beautifulsoup4, suggests none of that is present, so a fetch that fails is a fetch that fails.

A second alternative is simply not archiving at all and reading the platforms' own historical pages. Some of these sites keep a browsable past, but not uniformly, and not at hourly resolution. The specific thing website-hot-hub provides is a uniform, dated, plain-text snapshot across platforms that have nothing in common otherwise. If you only need one platform, use that platform's own tooling and skip this.

Licence, upgrade cost and the maintenance burden you are taking on

The repository is MIT licensed, with a LICENSE file at the top level. In practical terms that permits commercial and private use, modification and redistribution provided the copyright notice and permission notice are retained. This is a summary of the licence text, not legal advice; if you plan to redistribute the archives or embed the code in a product, read the LICENSE file and take your own counsel.

One point worth flagging for anyone republishing the output: the MIT licence covers the code, not the content the code collects. The archives contain headlines and links from 36Kr, bilibili, Douyin, Juejin, WeRead, Kuaishou and GitHub. The README says nothing about the terms under which that content may be redistributed, so the licence on the repository should not be read as permission to republish the headlines themselves.

Upgrade cost is low in the ordinary sense, because the dependency list is two packages and the code is a handful of Python files. The real cost is ongoing: each website_*.py module is a maintenance liability that fails when the target site changes. The release history reinforces this. v0.0.1 covered 2023-2024, v0.0.2 is labelled Archive-2025, and the last push was on 2026-02-07. A team adopting this should budget for reading each module and repairing it when a platform changes, rather than expecting upstream to do it.

Editorial conclusion

Adopt it if you need a self-hosted, plain-text record of Chinese platform trending lists and you are willing to keep the per-site scrapers working yourself. Do not adopt it if you want a live dashboard, an API, or a guaranteed daily history, because the README documents no scheduler and the archives only exist where someone ran the script. Before relying on it, check the archives directory for gaps on the dates you care about, and confirm that each website_*.py module you need still parses the current page markup.

Frequently asked questions

Which platforms does website-hot-hub collect trending lists from?

The README names 36Kr, bilibili, GitHub, Douyin, Juejin, WeRead and Kuaishou, and the repository also contains a website_sspai.py module. Each platform has its own scraper file at the top level.

How often does website-hot-hub fetch data, and where is it stored?

The README states that data is fetched once per hour and archived by day, starting from 2023-10-25. The output lands in the archives directory and in generated blocks inside README.md.

Does website-hot-hub provide a web interface or an API?

No. The repository layout contains no server component or API layer; the output is markdown and files on disk, so any browsing interface would have to be built on top of the archives.

What are the dependencies for running website-hot-hub?

requirements.txt pins requests==2.32.4 and beautifulsoup4==4.12.3, and the entry point is main.py. There is no headless browser dependency, so the scrapers work against server-rendered HTML.

What is the current state of website-hot-hub releases?

The last push was on 2026-02-07 and the most recent release is v0.0.2, tagged Archive-2025; the earlier v0.0.1 is tagged Archive-2023-2024. The repository is not archived, and the release naming points to periodic archival passes.

Official sources

  1. Official README
  2. Project repository
  3. Release notes
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/cxyfreedom-website-hot-hub.svg)](https://hysenlabs.com/projects/cxyfreedom-website-hot-hub)