cnchar: pinyin, stroke order and Chinese text utilities as a JavaScript library
A full-featured Chinese character utility library (pinyin, strokes, radicals, idioms, speech, visualization and more).
At a glance
- What is it?
- A character utility library covering pinyin, stroke order, traditional conversion, idioms and drawing, split across a core package and separate plugins.
- Who is it for?
- cnchar is a good fit when you are sorting, searching or teaching Chinese text in the browser, and specifically when stroke order is part of the requirement, since most alternatives stop at pinyin. The core package is pure JavaScript and loads without a framework, and each capability beyond pinyin and strokes is a separate plugin so you only pay for what you use.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 69 days ago.
- What is it written in?
- Mainly TypeScript, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 9, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What the library actually covers
The repository describes itself as a comprehensive Chinese character utility library covering pinyin, strokes, radicals, idioms, voice and visualisation, and the documentation table of contents shows how wide that claim is. There are fourteen numbered API sections, and the list is worth reading as a capability inventory rather than a summary.
The foundations are pinyin with tone, including polyphone handling for characters that read differently by word, and stroke order as an ordered sequence. On top of those sit drawing, which animates a character stroke by stroke on canvas, conversion between simplified, traditional and what the documentation calls Martian text, and three reverse lookups: stroke sequence back to a character, pinyin back to characters, and stroke count back to characters.
The language layer is large enough to be its own category. Idioms, xiehouyu, radicals, word formation and word explanations each have a section. Voice has three separate APIs: voice, speak and recognise, meaning the library addresses pronunciation data, speech synthesis and speech recognition.
The final section is the one that tells you how extensible it is. There are setters for custom data: `setSpell`, `setSpellDefault`, `setStrokeCount`, `setPolyPhrase`, `setOrder`, `setRadical` and `addXhy`. Those let you correct or extend the bundled dictionary rather than forking it.
Core plus plugins, and why that split matters
The build scripts in `package.json` show the packaging model. There is a `build:main` target for the core bundle, and a separate `build:cnchar-all` target that takes a plugin name through an environment variable:
npm run build:cnchar-all --env.pluginname=allThat script name is the clearest evidence for how the library divides itself. Pinyin and stroke data is core, and everything else arrives as a plugin, which is why the README has separate pages for traditional conversion, idioms, voice and drawing. Each plugin brings its own data file, and the reason to care is bundle size: a page that only needs pinyin should not ship an idiom dictionary and a speech recognition code path.
The topics confirm the shape of the feature set: `chinese-characters`, `draw`, `pinyin`, `speak`, `spell-stroke` and `voice-recognition`.
The web tooling is older but complete for what it does. `vuepress/` and the `dev:docs` and `build:docs` scripts drive the documentation site, `webpack-config/` holds the bundler setup, `helper/` contains the project's own build, release and publish scripts including a CDN purge step, and `jest.config.js` with `test` and `jest` scripts covers testing. The `public/` directory serves the live demos, and several of the README's examples are runnable pages rather than screenshots.
Installing it and loading it in each environment
The README documents two installation routes, npm and a CDN include, and then three loading situations: a webpack browser environment, a Node environment, and a plain browser environment. That separation matters more than it sounds, because the drawing API needs a canvas and will not work in Node.
From a build perspective the entry point is `cnchar`, and the package's own metadata sets `main` to `index.html`, which is unusual for an npm package and points at the unpkg style entry that lets browser tooling resolve the prebuilt bundle. The version in `package.json` is `3.2.6`, matching the most recent release tag.
The plugin pattern shows up in how the bundles are produced, where the core and a named plugin are built separately and then combined:
npm run build:main
npm run build:cnchar-all --env.pluginname=allThe first target produces the core bundle and the second passes a plugin name through an environment variable, which is the whole packaging story in two lines.
If you are not using a bundler, the CDN route exists for exactly that case, and the docs site at `theajack.github.io/cnchar` is both the reference and the place where the interactive examples run, which is the fastest way to see whether the API shape fits your code.
One housekeeping note for a project of this shape: the repository tree contains both `README.md` and `README.en.md`, so English documentation does exist even though the primary README is written in Chinese.
Version history and what the repository does not settle
Three releases are on record: `v3.2.4` on 2023-04-09, `v3.2.5` on 2023-12-16, and `v3.2.6` on 2024-03-23. The last push was on 2026-08-02, which means roughly eighteen months of commits sit on `master` without a release behind them.
That gap is the main thing to weigh. The `package.json` version still reads `3.2.6`, so anyone installing from npm gets the March 2024 build regardless of how recent the source looks. For a character data library that is a specific kind of staleness: the code may not have changed much while the dictionaries were extended, and the npm artifact would not carry that.
The repository does not answer when the next release is planned, and there is no changelog file in the tree, though the README links to one at `helper/version.md`, which is the release notes location.
A second thing the repository leaves open is data provenance. There is no documentation here stating where the pinyin, stroke order and idiom data came from or how a disputed reading gets resolved. The custom data setters exist precisely because the bundled data sometimes needs correcting, and knowing how often to reach for them is a judgement call the project does not help with.
Where cnchar fits against the alternatives
The nearest comparison is a server-side approach: sending Chinese text to a service that returns pinyin and strokes. That is often the wrong direction for this problem, because transliteration is per-character and context-sensitive in Chinese, so doing it correctly means segmenting text into words first. A library that runs where your text already is avoids a network round trip on every keystroke, which is the difference between a typing helper that feels instant and one that does not.
The other comparison is a minimal pinyin library. cnchar is considerably larger in scope, and if all you need is pinyin for a search box, a smaller dependency is the better answer. What the smaller option will not give you is stroke order as an ordered sequence, which is what you need for drawing, for handwriting input, for teaching, and for any check that cares about stroke count rather than just reading.
Its own README lists the kind of things people build with it: a character typing game, a name picker, an idiom chain game, contact list sorting, an input method, xiehouyu lookup and simplified to traditional conversion. Those examples are the honest scope. This is a text and UI utility library for Chinese characters, not a natural language processing stack, and it does not do tokenisation, sentence analysis or translation.
Licensing is MIT, which is the permissive end and imposes nothing on a commercial application. The maintainer lists PayPal, Ko-fi and WeChat donation options, and the project has a Gitee mirror alongside GitHub, which matters for anyone outside China who needs a reliable fetch.
Editorial conclusion
cnchar is a good fit when you are sorting, searching or teaching Chinese text in the browser, and specifically when stroke order is part of the requirement, since most alternatives stop at pinyin. The core package is pure JavaScript and loads without a framework, and each capability beyond pinyin and strokes is a separate plugin so you only pay for what you use. Check three things before committing: which plugins your use case needs, because idioms, traditional conversion and voice are not in the core, whether the browser bundle size matters to you since the README carries a minified size badge, and whether the last release in March 2024 is recent enough given that commits have continued since. The upstream rule set also matters: rare characters and polyphonic readings are exactly where a character dictionary has to be complete.
Frequently asked questions
How do I get pinyin and strokes for a Chinese character with cnchar?
The core package covers pinyin with tone and stroke order as an ordered sequence, which is the foundation the other plugins build on. Load the core for those two capabilities and add a plugin when you need stroke animation, since drawing ships separately.
Which parts of cnchar are separate plugins?
Pinyin and stroke order are core. Drawing, simplified and traditional conversion, idioms, xiehouyu, radicals, word data and the voice and speech APIs are documented as their own sections and are loaded as plugins through cnchar.use(). The build script build:cnchar-all takes a plugin name through env.pluginname.
How current is the published cnchar package?
The latest release tag is v3.2.6 from 2024-03-23, which is also the version in package.json, and the last push was on 2026-08-02. Installing from npm therefore gives you the March 2024 build rather than the current state of the source tree.
Can cnchar correct or extend its character data?
Yes, there is a dedicated API section for custom data with setters including setSpell, setStrokeCount, setPolyPhrase, setOrder, setRadical and addXhy. That is the supported route for fixing a wrong reading or adding a rare character rather than patching the bundled dictionary.
Does cnchar work in Node as well as the browser?
The documentation covers webpack browser environments, Node environments and plain browser environments separately, so the core text operations are usable outside a browser. The drawing capability is a canvas feature and is browser oriented, which is why the environments are documented apart.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/theajack-cnchar)