Open-source project
hiddendevj/Crawler_Illegal_Cases_In_China avatar
hiddendevj/Crawler_Illegal_Cases_In_China

Crawler_Illegal_Cases_In_China: A Case Archive for Chinese Web Scraping Law

Collection of China illegal cases about web crawler 本项目用来整理所有中国大陆爬虫开发者涉诉与违规相关的新闻、资料与法律法规。致力于帮助在中国大陆工作的爬虫行业从业者了解我国相关法律,避免触碰数据合规红线。

4,737 stars320 forksHTMLLicense varies

At a glance

What is it?
This repository collects real Chinese prosecution cases, news reports, and statutory text about web crawler developers who faced criminal or civil penalties. It is aimed at developers working in mainland China who need concrete examples of where data collection crosses legal lines.
Who is it for?
Crawler_Illegal_Cases_In_China is worth reading before any data collection project that touches Chinese websites. It shows which statutes prosecutors have actually invoked, what thresholds triggered charges, and that liability extends upward to managers.
Can I use it commercially?
Not without permission. GitHub finds no licence file in the repository, and without a licence all rights are reserved by default: you may read the code but not reuse it. Check the README, or ask the authors, before using it.
Is it still maintained?
Activity is slowing. The repository last received commits 6 months ago.
What is it written in?
Mainly HTML, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What This Repository Is and Who Needs It

Crawler_Illegal_Cases_In_China is a manually curated reference for engineers and data professionals whose work involves collecting information from websites under Chinese jurisdiction. It does not offer legal advice. Instead, it provides what statutory text alone cannot: a map of how prosecutors and courts have actually classified scraping-related conduct.

The repository owner describes it as a resource to help crawler developers and data industry professionals understand Chinese law and avoid crossing compliance lines. The collection covers news reports, court judgment summaries, lawyer commentary, and excerpts from the relevant criminal statutes. Every entry is sourced: each case either links to the original news article or points to a subdirectory in the repository with further documentation.

For anyone writing scrapers that will run against Chinese websites, the value is specificity. Knowing that selling scraped personal data for 11,584 yuan resulted in a criminal conviction is more concrete guidance than reading that personal information infringement carries penalties for "serious" violations.

Five Risk Categories and How the Cases Are Organized

The README structures the 16 documented cases into five zones, each corresponding to a category of conduct that has led to prosecution.

Zone 1 covers developers who provided scraping infrastructure to organizations already engaged in illegal activity. Two cases involve CAPTCHA-breaking service operators, and one involves a group that manipulated search engine results for profit.

Zone 2 covers personal privacy data collection and resale. Six cases fall here. They include a resume data company prosecuted for bulk collection of personal records, a developer convicted for selling user data scraped from a professional networking platform, and an operator who stored more than 20 million citizen account credentials.

Zone 3 covers scraping commercially sensitive data owned by another company and profiting from it. Five cases include a vehicle tracking application that allegedly copied bus schedule data, a scraper selling access to a court document database, and a platform that provided livestream data from a major e-commerce site.

Zone 4 contains one case: a company whose scraper sent 183 requests per second to a government residence permit website until the system crashed. Both the CTO and the programmer were convicted.

Zone 5 contains a single case involving a former executive at a major internet company, with no explicit category name assigned. The README labels it as an open category.

The repository also includes two further sections: one with excerpts from the applicable statutes and links to their full texts, and one linking to analytical articles written by named practicing lawyers.

The Four Criminal Statutes That Drive Most Charges

Article 285 of the Criminal Law of the People's Republic of China covers illegal acquisition of computer system data outside of national defense and state affairs systems. It applies when a person gains unauthorized access to a computer information system and obtains data stored, processed, or transmitted there. The statute provides for a sentence of up to three years for serious cases and three to seven years for especially serious cases, in both instances with a fine.

Article 286 covers disruption of computer information systems. It applies when conduct causes a system to malfunction or deletes, modifies, or adds data or programs. For serious harm the sentence is up to five years; for especially severe harm, five years or more. Article 286 also explicitly addresses unit liability, stating that when an organization commits the offense, the managers directly responsible are punished under the same terms as individual offenders. The README notes this in connection with the Zone 4 case, where the CTO received a harsher sentence than the programmer.

The infringement of citizen personal information crime, codified in Article 253(a) of the amended Criminal Law, applies when personal data is collected, used, or sold without authorization. A supplementary judicial interpretation sets the thresholds for "serious" conduct: 50 or more records of sensitive categories such as travel routes or financial data; 500 or more records of transaction, health, or communication data; and 5,000 or more records of other personal information.

The Anti-Unfair Competition Law, Article 9, addresses commercial secrets. It applies when scraped data qualifies as a trade secret of the source organization and the scraper uses or discloses it without authorization. No criminal sentence threshold is given in the README excerpt; the focus is on the act of misappropriation.

What the Documented Cases Show About Prosecution Thresholds

The cases reveal patterns that go beyond the statutory thresholds. Resale is the most common aggravating factor, appearing across Zones 2 and 3. The Maimai personal information case resulted in conviction for profit of 11,584 yuan, which is a low absolute amount. The novel-scraping case in Zone 3, where scraped content was provided free to readers, resulted in prosecution for profits of roughly 10 million yuan.

Scale is a separate pathway to prosecution. The Zone 2 case involving the social insurance application that exposed user data resulted in the application being removed from distribution. The Zone 2 case involving 20 million stored credentials resulted in a conviction regardless of whether those credentials were sold.

Operational impact is the third pathway. The Zone 4 case demonstrates that a developer who neither sells data nor targets personal information can still face criminal liability if the scraping disrupts the target system. At 183 requests per second, the supercomputer supporting the residence permit website became unavailable, and both the software developer and the CTO were held accountable.

The repository also documents a managerial liability issue that is easy to overlook. Article 286 holds direct supervisors responsible under the same statutory framework as the individual who wrote the code. In the Zone 4 case, the CTO was sentenced more severely than the programmer. Any developer working inside an organization should be aware that their manager can be charged for work performed under their supervision.

Where This Repository Has Real Gaps

Several cases in the collection lack final court judgments. The README explicitly flags CASE13, involving a livestream data platform whose CEO was arrested. The entry states that a judgment will be added once it becomes public. Any case without a final judgment should be read as an arrest record, not a conviction record, and the outcome should not be assumed.

The repository does not address the Personal Information Protection Law enacted in November 2021 or the Data Security Law enacted in September 2021. Those statutes introduced consent requirements, cross-border data transfer rules, and tiered obligations based on data sensitivity and volume. A developer relying only on this repository would miss the regulatory layer that both laws added on top of the Criminal Law framework documented here.

The list has no search functionality. Finding all cases involving a particular statute, a particular industry sector, or a particular outcome requires reading the README from top to bottom or opening each case subdirectory. The flat structure works for an initial orientation but does not serve iterative research well.

The repository has no stated license, which introduces ambiguity about whether its compiled summaries can be reproduced in a compliance document, a training dataset, or a published analysis.

A Professional Legal Database as an Alternative

Pkulaw (北大法宝), operated by Peking University, is a subscription-based legal research platform providing full-text court judgments, statutory text, legislative histories, and practitioner commentary in Chinese. It covers the same statutes documented in Crawler_Illegal_Cases_In_China and a far larger set of judgments, updated continuously as new decisions are published.

The practical difference is access and depth. Pkulaw requires institutional or paid individual access, and it is designed for legal professionals who need to search across thousands of judgments by charge, outcome, or court level. This repository is public, free, and carries editorial context written for software engineers rather than lawyers. For an engineer doing a first pass to understand what laws apply and what sentences have been imposed, the public repository is more immediately readable. For a compliance officer building a data handling policy or a lawyer preparing a formal legal opinion, a judgment database provides complete documentation that no community-compiled list can match.

Maintenance Status and the License Question

The last push to this repository was on 2026-03-12. No GitHub releases have been tagged. The repository metadata does not list a license, though the README links to external content under various terms. An engineer who wants to reproduce the case summaries in internal documentation should investigate the licensing status of each external source independently.

The README acknowledges its own incompleteness in two places: the CASE13 entry notes that the judgment is not yet public, and the description of Zone 5 leaves the category name blank. The three lawyer commentary articles linked in the final section are hosted on external platforms; whether those links remain valid over time is not guaranteed. The repository is best read as a starting point for understanding the enforcement landscape that existed through early 2026, not as an authoritative or exhaustive legal reference.

Editorial conclusion

Crawler_Illegal_Cases_In_China is worth reading before any data collection project that touches Chinese websites. It shows which statutes prosecutors have actually invoked, what thresholds triggered charges, and that liability extends upward to managers. The project does not cover newer frameworks like PIPL or DSL, and some entries still lack final court judgments. Treat it as a map of past enforcement, then verify current obligations against the source statutes before drawing compliance conclusions.

Frequently asked questions

What types of crawler activity does Crawler_Illegal_Cases_In_China document as criminal?

The repository documents five categories: providing scraping services to illegal organizations such as CAPTCHA-breaking or SEO manipulation operations; collecting and reselling personal data; profiting from commercially sensitive data owned by another company; causing server outages through excessive request rates; and one case involving an internet executive that the README does not explicitly categorize.

Does Crawler_Illegal_Cases_In_China include the text of Chinese criminal statutes?

Yes. The README includes excerpts from Article 285 and Article 286 of the Criminal Law, the personal information infringement provisions from the amended Criminal Law, the Anti-Unfair Competition Law on trade secrets, the Cybersecurity Law, and the Civil Code, along with links to the full statutory texts.

Are all 16 cases in the repository accompanied by final court judgments?

No. At least one case, CASE13 involving a livestream data platform, is noted in the README as pending a public judgment. Some entries link to news articles rather than court documents. The collection should be read as an enforcement record, not a complete verdict archive.

Official sources

  1. hiddendevj/Crawler_Illegal_Cases_In_China on GitHub
  2. Issues
  3. Project website
  4. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/hiddendevj-crawler-illegal-cases-in-china.svg)](https://hysenlabs.com/projects/hiddendevj-crawler-illegal-cases-in-china)