Open-source project
chinese-poetry/chinese-poetry avatar
chinese-poetry/chinese-poetry

chinese-poetry: a 55,000-poem JSON corpus whose only release tag is from 2017

GitHub describes it as The most comprehensive database of Chinese poetry 🧶最全中华古诗词数据库, 唐宋两朝近一万四千古诗人, 接近5.5万首唐诗加26万宋诗. 两宋时期1564位词人,21050首词。. The repository metadata lists JavaScript as its primary language. The metadata lists the MIT license. This article stays within the project description and details documented in the GitHub repository README.

53,498 stars10,740 forksJavaScriptMIT

At a glance

What is it?
chinese-poetry is a large MIT licensed dump of Chinese classical texts distributed as JSON files on a branch, with 55,000 Tang poems, 260,000 Song poems, 21,000 Song ci, and exactly one release, v0.1, from May 2017. It is a corpus to vendor, not a library to depend on.
Who is it for?
Use chinese-poetry when you need the raw Chinese text of Tang verse, Song verse, or Song ci and are willing to own the parser, the storage, and the versioning yourself. Do not treat it as an authority you can cite without checking: the data was gathered from the internet without a recorded process, no per-record source is stored, and a disputed line is settled by community vote rather than by evidence held in the data.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 105 days ago.
What is it written in?
Mainly JavaScript, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 27, 2026, and from our analysis. They are not legal advice.

Editorial analysis

One release tagged v0.1 in 2017 while master was pushed in June 2026

The release history is a single line. The repository has one release, v0.1, dated 2017-05-26, while the default branch master was last pushed 2026-06-17. Nothing in between produces a version number to pin, and the project's own purpose is stated as generating formatted JSON data so that developers can build poetry applications more quickly, which means the artifact people actually depend on is a directory tree on a branch rather than a package. The consequence is a dependency with no pins. There is no semantic versioning contract, no release tarball, and no changelog to diff, so when a corpus file is corrected, added, or renamed there is no way for a downstream project to find out except by watching master. Projects listed as building on this data include an Android app holding an offline copy of the full Tang poems and a MySQL backed web search, and both pinned themselves against whatever tree they happened to clone.

The crawl was never recorded, so a disputed verse cannot be settled from the data

Provenance is the softest part of the dataset, and the project is direct about why. The text states the data was gathered from the internet, and it explains that the collection process was not recorded because the corpus is large, the target sites impose limits, and the crawl was repeatedly interrupted for more than a week. The single published write-up covers the crawl behind the Complete Song Ci added in 2017. What that means for a user is concrete. There is no capture timestamp on a record, no stored snapshot of the page it came from, and no field pointing back at a source, so when two readings of a line disagree the database cannot adjudicate from its own contents. The project routes that outward instead: corrections sent as pull requests are asked to state their source, and records the maintainer treats as disputed are decided by community vote. The verification effort therefore lands on whoever files the correction, not on the data pipeline.

Four dataset directories sit in the tree without appearing in the dataset list

The dataset list and the directory tree do not line up. What is listed is the Tang and Song poems, the Complete Song Ci, the Five Dynasties Flower Between collection and the Southern Tang rulers' ci, then the Analects, the Book of Songs, 幽梦影, the Four Books and Five Classics, primer texts, the Nalan Xingde poetry collection, and the Imperial Complete Tang Poems. The top level of the tree holds more than that. Alongside the listed ones there are directories for 元曲, 曹操诗集, 楚辞, and 水墨唐诗, and none of the four is named in the dataset list. A consumer planning an ingestion from the documentation alone would skip Chu ci and the Yuan qu without knowing they exist, and no schema or file layout is given for the unlisted directories, so their structure has to be discovered by opening the files. The working directories loader/, rank/, and strains/ are also unlisted.

requirements.txt pins a single test runner and nothing else

Distribution is by clone, and the toolchain shows it. requirements.txt contains exactly one entry, pytest==5.3.2, and the only test file in the tree is test_poetry.py, so the sole pinned Python dependency is a test runner rather than anything needed to read the corpus. The primary language is recorded as JavaScript, the working directories are loader/, rank/, and strains/, and a log.log file is committed at the top level. There is no install command in the documentation, no npm or pip package name, and no HTTP endpoint; the browser front door is the project homepage at awesome-poetry.top, and the source of truth is the JSON sitting on master. A team that wants these poems inside an application is expected to write its own reader, which is precisely what each of the listed downstream projects did on its own terms.

The frequency tables are derived, undated, and not regenerable from the documentation

Not everything in the repository is raw text. Alongside the corpora the documentation carries frequency analysis in collapsible sections: a chart of the most popular Song ci tune names, which is open by default, then high frequency word tables and author work rankings for Song ci, Tang poems, and Song poems, each collapsed until clicked open. That is a real convenience for anyone doing frequency work on classical Chinese. It is also an undated derived product. The documentation does not state which corpus snapshot, which tokenisation, or which refresh date produced any of those tables, so a count copied from the Song ci word table cannot be reproduced from the repository alone. The texts are in the tree and the counts are beside it, and because loader/ and rank/ are not explained, there is no documented command that rebuilds the tables, which means the two can drift apart unnoticed.

Every downstream project had to build its own ingestion on top of the same JSON

The showcase list is the best available evidence of what the data is good for, and every entry does its own ingestion. There is a browser poetry site at chinese-poetry.github.io carrying three hundred Tang poems and three hundred Song ci, an offline Android app holding the full Tang collection, a character level recurrent network written in PyTorch, a repository of Tang poem generation with PyTorch that can produce acrostics and take a mood and a prefix, a desktop poetry app, a mini program version, a clean searchable site, a MySQL integration with web side search and retrieval, and a PaddleNLP notebook generating poetry with ERNIE-GEN. That is eight independent parsing layers sitting on one set of files. The upside is that the layout has survived contact with very different consumers. The downside is that none of them is a supported interface, and a change to a file on master reaches all of them at once with no deprecation path, because the corpus has never been versioned.

Corrections run through one mailbox, and the sponsors list is empty

Governance is one inbox. Additions and corrections arrive as pull requests or issues, disputed records are settled by community vote, a correction is asked to state its source, and the detailed rules live in a wiki page rather than inside the repository. Funding routes are listed as a Chinese patronage platform, a Patreon subscription, one time Alipay and WeChat payments where you leave an email address, and a direct address, [email protected]. The sponsors section currently reads none. The project describes itself as short handed and asks for more maintainers. For a corpus of 55,000 Tang poems, 260,000 Song poems, and 21,000 Song ci, the implication for a downstream project is that correction throughput depends on one person reading issues, with no published turnaround, no reviewer roster, and no second maintainer named. Continuous integration is Travis running that one pytest file, so a green badge is evidence about test_poetry.py and not about every file in the collections.

Editorial conclusion

Use chinese-poetry when you need the raw Chinese text of Tang verse, Song verse, or Song ci and are willing to own the parser, the storage, and the versioning yourself. Do not treat it as an authority you can cite without checking: the data was gathered from the internet without a recorded process, no per-record source is stored, and a disputed line is settled by community vote rather than by evidence held in the data. Before building on it, open the JSON for the collection you actually need and confirm the field names for yourself, check whether that collection is one of the four directories missing from the dataset list, and pin a commit rather than tracking master, since the only release is v0.1 from 2017-05-26 while the branch itself was last pushed 2026-06-17.

Frequently asked questions

what is chinese poetry

In this repository, chinese poetry means a database of Chinese classical literature rather than a literary category: 55,000 Tang poems, 260,000 Song poems, 21,000 Song ci, nearly 14,000 Tang and Song poets, and about 1,500 Song ci poets. The data is distributed as JSON under the MIT license, with the project homepage at awesome-poetry.top.

chinese poetry types

The repository separates Tang verse, Song verse, and Song ci, and adds the Five Dynasties Flower Between collection and the Southern Tang rulers' ci as separate links. It also carries the Analects, the Book of Songs, 幽梦影, the Four Books and Five Classics, primer texts, the Nalan Xingde poetry collection, and the Imperial Complete Tang Poems, plus directories for 元曲, 曹操诗集, 楚辞, and 水墨唐诗 that the dataset list does not mention.

What is the oldest collection of chinese poetry?

The repository does not make that claim anywhere. It hosts the Book of Songs, the Analects, the Four Books and Five Classics, primer texts, 幽梦影, the Nalan Xingde poetry collection, the Imperial Complete Tang Poems, a Chu ci directory, and the Tang, Song, and Five Dynasties verse collections, and it leaves the ordering of those to the reader.

does chinese poetry rhyme

The documentation does not describe a rhyme, metre, or pinyin field, because the data is distributed as JSON text without an annotated schema. What it does publish is frequency analysis: high frequency word tables and author work rankings for Tang poems, Song poems, and Song ci, plus a chart of the most popular ci tune names.

what is yijing in chinese poetry

The I Ching is not named in the dataset list. The closest entry is the Four Books and Five Classics directory, and the listed collections otherwise cover Tang verse, Song verse, Song ci, the Five Dynasties collections, the Book of Songs, the Analects, primer texts, and 幽梦影, so a search for a Yijing text would have to go through the 四书五经 directory.

Official sources

  1. Official documentation
  2. Official README
  3. Project repository
  4. Release notes
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/chinese-poetry-chinese-poetry.svg)](https://hysenlabs.com/projects/chinese-poetry-chinese-poetry)