PyCantonese: Cantonese Linguistics and NLP in Python
Cantonese Linguistics and NLP
At a glance
- What is it?
- PyCantonese is an MIT-licensed Python library for Cantonese corpus access, Jyutping romanization, word segmentation and part-of-speech tagging, and since v4.0.0 it delegates the heavy lifting to the Rustling Rust library. It is aimed at linguists and NLP developers working with Jyutping-tagged Cantonese text, not at general Chinese-language processing.
- Who is it for?
- Adopt PyCantonese if your data is Cantonese written in Chinese characters and Jyutping, and you need corpus search, romanization conversion, segmentation or POS tagging inside a Python pipeline. Do not adopt it if you need Mandarin, general-purpose Chinese segmentation, or a pure-Python dependency tree with no compiled extension.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 134 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 7, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What PyCantonese is for, and who it is not for
Cantonese is written with Chinese characters, but the romanization most linguistic work depends on is Jyutping, and the two do not map onto each other one-to-one. A pipeline that needs to search a spoken Cantonese corpus, convert between characters and Jyutping, split a sentence into words, or assign part-of-speech tags has to handle that relationship explicitly. PyCantonese packages those operations behind a Python API. The README lists the implemented features plainly: accessing and searching corpus data, parsing and conversion tools for Jyutping romanization, parsing Cantonese text, stop words, word segmentation, and part-of-speech tagging.
The intended audience is narrower than the name suggests. The classifiers in pyproject.toml declare an audience of developers, education, information technology and research, and the natural languages are Traditional Chinese and Cantonese. If you are building a general Chinese text pipeline, or you need Mandarin, this is the wrong library: nothing in the README claims coverage outside Cantonese. The project states that its design prioritizes ease of use and linguistic knowledge, and that it has been used by academic and commercial organizations, including major US tech companies. That is an assertion in the README, not a benchmark, and it is the only usage claim the repository makes.
How the Python layer and the Rustling dependency fit together
Since v4.0.0 (March 2026), PyCantonese depends on Rustling, described in the README as a library for efficient CHAT data handling, word segmentation, and part-of-speech tagging. The repository layout shows what that means in practice. There is a Cargo.toml at the top level alongside pyproject.toml, and the build backend declared in pyproject.toml is maturin, the tool that builds and ships Rust extensions for Python. The Rust crate is named pycantonese with a cdylib library target called _pycantonese, so the compiled module that Python imports carries an underscore prefix while the distributed package is pycantonese.
The dependency graph is explicit. Cargo.toml pulls in pyo3 0.28 with the abi3-py310 feature, rustling 0.8.0 with default features off and the pyo3 feature on, and regex. The default Cargo feature set is parallel, which enables rustling/parallel. On the Python side, pyproject.toml lists a single runtime dependency: rustling >= 0.8.0. So the architecture is a thin Python package that wraps a Rust core, with the linguistic work happening in Rustling and the corpus data shipped as package data (the README points to src/pycantonese/data for dataset documentation). The practical consequence is that installing PyCantonese is not a pure-Python affair; pip has to resolve a binary wheel or build one.
Installing PyCantonese and running a first Cantonese text through it
The README gives two install routes. The pip route is the one most users will take. Run it in a virtual environment so the compiled extension does not collide with system packages.
pip install --upgrade pycantoneseAfter this, `import pycantonese` should resolve. If no prebuilt wheel matches your platform, maturin will attempt a Rust build, which requires a working Rust toolchain and will take noticeably longer than a wheel download. Conda users have a second route, which avoids the local build entirely if the conda-forge package covers your platform.
conda install -c conda-forge pycantoneseThe README also states that PyCantonese works in JavaScript, linking to the quickstart page for that path. For Python work, the Quickstart at docs.pycantonese.org/stable/quickstart.html is the next stop; the README does not inline worked examples, so the exact function names for corpus search, Jyutping conversion, segmentation and tagging come from that page rather than from the README. The minimum interpreter is worth checking before you start: pyproject.toml sets requires-python to >= 3.10, and the classifiers list 3.10 through 3.14. The abi3-py310 feature in Cargo.toml is what makes one wheel usable across those versions, so a 3.9 environment will not install the package at all.
The compiled extension is the main operational constraint
The most concrete limitation visible in the repository is that PyCantonese is no longer pure Python. Because the build backend is maturin and the crate produces a cdylib, any environment that cannot consume a prebuilt wheel needs a Rust toolchain at install time. That matters for locked-down CI images, for Alpine or other musl-based containers where wheel availability differs from glibc, and for architectures where rustling 0.8.0 may not publish a wheel. The README does not document a fallback, and it does not describe how to build from source; the build configuration is only visible in pyproject.toml and Cargo.toml.
The second constraint is scope. The bundled data comes from five sources with five different licences, listed in the README: the Hong Kong Cantonese Corpus under CC BY, CantoMap under GPL-3.0, rime-cantonese under CC BY 4.0, Common Voice Cantonese under MPL 2.0, and the Cantonese-Traditional Chinese Parallel Corpus under CC0 1.0. The package itself is MIT, but the data it carries is not uniformly MIT. Anyone redistributing PyCantonese as part of a larger product, or extracting the corpora into their own dataset, is dealing with a mixed-licence bundle. The README points to the data directory documentation for details; it does not summarize the obligations per dataset. A third limitation is simply linguistic: if your text is Mandarin, or code-switched in ways the Cantonese tagger was not built for, the segmentation and tagging features are not the right instrument, and the README makes no claim that they generalize.
How PyCantonese differs from a general Chinese NLP toolkit
The obvious alternative for Chinese text processing is a general-purpose library such as jieba, which segments Simplified and Traditional Chinese using statistical models trained largely on written Mandarin. The difference in approach is not just language coverage, it is what the output looks like. A general Chinese segmenter returns tokens; PyCantonese is built around Jyutping, so romanization conversion and Cantonese-specific part-of-speech tagging sit alongside segmentation. If your downstream task needs Jyutping at all, a Mandarin-oriented segmenter gives you nothing to work with and you would have to bolt on a separate romanization dictionary.
The trade-off runs the other way too. A general Chinese toolkit has a much larger user base, broader documentation, and no compiled extension to worry about. PyCantonese is a focused library for one variety, and its feature list is short by design: corpus access, Jyutping parsing and conversion, text parsing, stop words, segmentation, tagging. There is no claim of entity recognition, translation, or model training. Choosing PyCantonese means accepting a narrower surface in exchange for Cantonese-specific linguistic data, including the corpora bundled with the package.
Maintenance, versioning and what upgrading costs
The last push to the repository was on 2026-05-26, and the most recent release is v5.0.0 on the same date, following v4.3.0 on 2026-05-07 and v4.2.0 on 2026-03-27. The repository is not archived. The release cadence in that window is tight, and the major version bump from 4.x to 5.0.0, together with the note that v4.0.0 introduced the Rustling dependency, tells you where the churn has been: the Rust core and its Python bindings.
For anyone upgrading, the version pins are the thing to read. pyproject.toml requires rustling >= 0.8.0, and Cargo.toml pins pyo3 to 0.28 and rustling to 0.8.0 with default features disabled. A project that pins pycantonese to a 4.x release while also pinning rustling will need to move both together. The MIT licence on the package is permissive and imposes no copyleft on your own code, but the MIT grant covers the PyCantonese code, not the bundled datasets, which carry their own terms as listed above. That distinction is worth resolving before shipping a product that embeds the corpora. This is a description of what the repository states, not legal advice.
Editorial conclusion
Adopt PyCantonese if your data is Cantonese written in Chinese characters and Jyutping, and you need corpus search, romanization conversion, segmentation or POS tagging inside a Python pipeline. Do not adopt it if you need Mandarin, general-purpose Chinese segmentation, or a pure-Python dependency tree with no compiled extension. Before committing, verify three things: that your Python is 3.10 or newer, that a rustling wheel exists for your platform so the maturin build does not fall back to compiling Rust, and that the bundled corpora licences (CC BY, GPL-3.0, CC BY 4.0, MPL 2.0, CC0 1.0) are compatible with how you plan to redistribute the data.
Frequently asked questions
Is Cantonese the same as Chinese?
No. PyCantonese treats Cantonese as its own language: pyproject.toml lists the natural languages as Traditional Chinese and Cantonese, and the library's romanization work is built around Jyutping rather than Mandarin pinyin.
What is hello in Cantonese?
The repository does not provide a phrasebook or translation feature, so it does not answer this. Its features are corpus access and search, Jyutping parsing and conversion, text parsing, stop words, word segmentation and part-of-speech tagging.
What country speaks Cantonese?
The repository does not cover where Cantonese is spoken. What it does show is the data PyCantonese ships, including the Hong Kong Cantonese Corpus, CantoMap, rime-cantonese, Common Voice Cantonese and the Cantonese-Traditional Chinese Parallel Corpus.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/jacksonllee-pycantonese)