Library / SDK
zh-lx/pinyin-pro avatar
zh-lx/pinyin-pro

pinyin-pro: A JavaScript Library for Hanzi to Pinyin Conversion, Matching and Segmentation

中文转拼音、拼音音调、拼音声母、拼音韵母、多音字拼音、姓氏拼音、拼音匹配、中文分词

4,740 stars401 forksTypeScriptMIT

At a glance

What is it?
pinyin-pro converts Chinese characters to pinyin with tone marks, numbers or no tones, and adds pinyin matching, segmentation and HTML ruby output. It is an MIT-licensed TypeScript package for Node and the browser, with a large dictionary that shapes both its accuracy and its bundle size.
Who is it for?
pinyin-pro fits JavaScript and TypeScript projects that need pinyin conversion, pinyin-aware search matching or ruby annotation in the browser or in Node, and the README's own comparison puts it ahead of pinyin and @napi-rs/pinyin on accuracy. It is the wrong choice when bundle size dominates: the README measures the pinyin API at 306.41 KB minified and 134.52 KB gzipped in ESM, so a page that only needs a small lookup table should look elsewhere.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 20 days ago.
What is it written in?
Mainly TypeScript, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 5, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What pinyin-pro solves, and who ends up using it

Converting Hanzi to pinyin is not a lookup table problem in practice. A character can carry several readings, and the correct one depends on the word it sits in. The README's example is 睡着了, which pinyin-pro returns as "shuì zháo le"; a naive per-character mapping would not pick zháo there. The library exists to make that contextual choice, and to expose the pieces around it: initials, finals, tone marks, tone numbers, surnames, matching and segmentation.

The audience is JavaScript and TypeScript developers. The README shows an npm install and a script tag pointing at unpkg, so the same package serves a Node service and a browser page. That matters for products where a search box accepts pinyin input, where a reading aid annotates Chinese text, or where a backend needs to index Chinese content by its pronunciation. The library is written in TypeScript and published under MIT, which is about as permissive as a dependency gets.

How the conversion, matching and segmentation APIs fit together

The package exposes several named exports rather than one function with modes. pinyin() is the core converter and takes options such as type: "array" and toneType: "none" or "num". match() takes a Chinese string and a pinyin query and returns the index positions that matched, and the README shows three query styles against 中文拼音: the initials "zwp", the full spelling "zhongwenpin", and the mixed "zhongwp", all returning [0, 1, 2]. That mixed mode is the interesting one, because real users type a blend of full syllables and abbreviations and rarely announce which they meant.

segment() splits text into words and returns both the original slice and its pinyin, with an OutputFormat enum to switch the result shape. html() returns a ruby-annotated HTML string with classes like py-result-item, py-chinese-item and py-pinyin-item, which is a rendering concern rather than a data one. convert() moves between tone notations, for example from "pin1 yin1" to 'pīn yīn' and back with format: "symbolToNum", plus a toneNone option and erhua handling in the README's dou4 zhi1r example.

The architecture visible from the repository is a pnpm monorepo. The root package.json is private and names three pieces: pinyin-pro, @pinyin-pro/data and the documentation site. The dictionary lives in its own package, the library consumes it, and the docs are built separately. For anyone evaluating the dependency, that split is the reason the size numbers in the README are large: the data is the product.

Installing pinyin-pro and running a first conversion

The README gives two installation paths. For a project with a package manager, install from npm:

bash
npm install pinyin-pro

For a plain HTML page, the README loads the UMD build from unpkg:

html
<script src="https://unpkg.com/pinyin-pro"></script>

Once installed, a first real use is converting a string and reading the tone marks back. The README's own example imports the named export and passes a string:

js
import { pinyin } from "pinyin-pro";

pinyin("汉语拼音"); // 'hàn yǔ pīn yīn'

The return value is a space-separated string with tone marks. If you need to process syllables individually, the README shows the array form, and it also shows how to drop tones or switch to numbers:

js
pinyin("汉语拼音", { type: "array" }); // ["hàn", "yǔ", "pīn", "yīn"]
pinyin("汉语拼音", { toneType: "none" }); // "han yu pin yin"
pinyin("汉语拼音", { toneType: "num" }); // "han4 yu3 pin1 yin1"

The three options above are the ones most projects start with: array output for per-syllable work, toneType: "none" when you are building a search key, and toneType: "num" when the consumer is a database or a plain-text index. The README points to the online documentation at pinyin-pro.cn for the full option list, so treat the snippet as a starting point rather than the whole surface.

Where pinyin-pro is the wrong tool: bundle size and dictionary weight

The README publishes its own size table, and it is the most useful thing on the page for anyone deciding whether to adopt the library. The pinyin API is listed at 306.41 KB minified and 134.52 KB gzipped in ESM. The UMD build, which cannot tree-shake per API, is listed at 316.72 KB minified and 138.03 KB gzipped for the whole bundle. match is the lightest entry shown at 185.54 KB minified, though the README's table is truncated in the excerpt available here, so the gzipped figure for match cannot be quoted.

Those numbers are the cost of the dictionary, and they are not small for a browser bundle. A page that only needs to look up a few hundred characters, or that ships to users on slow connections, may be better served by a smaller table it controls. The README also states that UMD does not support tree shaking, so a script-tag integration pays the full weight regardless of which functions you call. If your build pipeline consumes the ESM output and the bundler can drop unused exports, the per-API figures are the ones that apply; if it cannot, assume the larger number.

There is a second boundary worth naming. The README describes a surname mode as a distinct feature, which implies the default reading is not always the surname reading. Names are exactly where pinyin conversion gets contentious, and a library that offers a mode for them is telling you the general case does not cover them. Anyone converting a contact list or a roster should test that mode rather than assume the default is correct. The README does not document a rollback or versioning policy for dictionary changes, so a dictionary correction between releases is something you would discover by upgrading.

pinyin-pro against pinyin and @napi-rs/pinyin

The README benchmarks pinyin-pro against two alternatives, pinyin and @napi-rs/pinyin, and publishes both the scripts and the results. The accuracy table lists pinyin at 94.097% for its Node version and 91.170% for its web version, @napi-rs/pinyin at 94.097%, and pinyin-pro at 99.846%. Those are the project's own measurements, run from scripts the README links to, so they are reproducible rather than third-party.

The speed table tells a more mixed story, and it is worth reading rather than skimming. Dictionary initialisation is listed at 8.412ms for pinyin-pro, 14.261ms for pinyin and 160.769ms for @napi-rs/pinyin. On 10k characters, @napi-rs/pinyin is fastest at 4.298ms against pinyin-pro's 7.216ms. On 100k characters the gap narrows: 29.32ms versus 45.471ms. At one million characters pinyin-pro is listed at 328.338ms and @napi-rs/pinyin at 297.41ms, and at ten million characters they are 3375.192ms and 3907.278ms respectively. The pinyin package runs out of memory at one million characters, according to the table.

The architectural difference explains the pattern. @napi-rs/pinyin is a native Node addon, and the README's compatibility row marks it as not supporting the web environment at all. That is the real trade: a native binding can be faster on bulk conversion in Node, but it cannot run in a browser, so it is not a substitute if your conversion happens client-side. pinyin is pure JavaScript and supports both environments, but the README's figures put it behind on accuracy and it fails outright on large inputs. pinyin-pro's pitch is the combination: browser and Node support, high accuracy, and throughput that stays within the same order of magnitude as the native addon at scale.

Maintenance, licence and the cost of upgrading

The repository is not archived, and the last push was on 2026-09-15. Recent releases are 3.29.4 on 2026-09-11, 3.29.3 on 2026-08-19 and 3.29.2 on 2026-08-15, so the project is being released on a short cadence. For a dependency whose value is a data table, that cadence cuts both ways: dictionary corrections arrive quickly, and so do version bumps you have to absorb.

The root package.json pins the toolchain to [email protected] and Node >=18, and it lists @napi-rs/pinyin and pinyin as devDependencies, which is how the comparison benchmarks run. That means the competitor figures in the README are generated from the repository itself. If you want to check the accuracy claim against your own corpus, the scripts are named in the root package.json as accuracy and speed, and the README links the corresponding files under packages/pinyin-pro/scripts/benchmark.

The licence is MIT, which places few conditions on commercial use. This is not legal advice; read the LICENSE file in the repository and your own organisation's policy before relying on that. The practical licence question for this library is not the terms but the data: the dictionary is a separate package, @pinyin-pro/data, and if you redistribute a built artefact you are redistributing that data along with the code.

Editorial conclusion

pinyin-pro fits JavaScript and TypeScript projects that need pinyin conversion, pinyin-aware search matching or ruby annotation in the browser or in Node, and the README's own comparison puts it ahead of pinyin and @napi-rs/pinyin on accuracy. It is the wrong choice when bundle size dominates: the README measures the pinyin API at 306.41 KB minified and 134.52 KB gzipped in ESM, so a page that only needs a small lookup table should look elsewhere. Before adopting it, check the size table for the exact API you plan to import, confirm whether your bundler tree-shakes ESM output, and decide whether you need the surname mode, since the README lists it as a separate feature rather than the default behaviour.

Frequently asked questions

How do I install pinyin-pro?

The README gives npm install pinyin-pro for package-manager projects, or a script tag loading https://unpkg.com/pinyin-pro for a plain browser page. The README notes that the UMD build does not support per-API tree shaking, so the script-tag route loads the full bundle.

Can pinyin-pro match text against pinyin input?

Yes. The match API returns the index positions that matched, and the README shows it working with initials ("zwp"), full spelling ("zhongwenpin") and a mix of the two ("zhongwp") against the string 中文拼音, each returning [0, 1, 2].

How large is pinyin-pro in a bundle?

The README's size table lists the pinyin API at 306.41 KB minified and 134.52 KB gzipped for ESM, and the UMD build at 316.72 KB minified and 138.03 KB gzipped. The match API is listed at 185.54 KB minified in ESM.

Does pinyin-pro work in the browser as well as Node?

The README's compatibility table marks pinyin-pro as supporting both the web and Node environments. The same table marks @napi-rs/pinyin as not supporting the web environment, which is the main architectural difference between the two.

What is pinyin-pro's licence?

The repository and package metadata state MIT. The dictionary is distributed as a separate package, @pinyin-pro/data, so a built artefact carries that data as well as the code.

Official sources

  1. License: MIT
  2. Project website
  3. README
  4. Releases
  5. zh-lx/pinyin-pro on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/zh-lx-pinyin-pro.svg)](https://hysenlabs.com/projects/zh-lx-pinyin-pro)