# pinyin-pro: one toneType option, six APIs, and a dictionary you pay for in kilobytes

> pinyin-pro is a TypeScript library that turns Chinese characters into pinyin, with tone marks, numeric tones or no tones, and adds matching, segmentation, format conversion and ruby annotated HTML. The interesting engineering is in what it costs you: the dictionary dominates the bundle, and the maintainer's own benchmark scripts compare it against the two libraries it lists as devDependencies.

**zh-lx/pinyin-pro** — 中文转拼音、拼音音调、拼音声母、拼音韵母、多音字拼音、姓氏拼音、拼音匹配、中文分词

- Repository: https://github.com/zh-lx/pinyin-pro
- Website: https://pinyin-pro.cn
- Stars: 4,740 · Forks: 401
- Language: TypeScript
- License: MIT
- Published: 2026-09-23 · Updated: 2026-09-23 · Language: en
- Canonical page: https://hysenlabs.com/projects/zh-lx-pinyin-pro

## toneType picks the shape of the output, nothing else changes

The main entry point takes a string and returns pinyin with tone marks, which is the default:

```js
import { pinyin } from "pinyin-pro";
pinyin("汉语拼音"); // 'hàn yǔ pīn yīn'
pinyin("汉语拼音", { type: "array" }); // ["hàn", "yǔ", "pīn", "yīn"]
pinyin("汉语拼音", { toneType: "none" }); // "han yu pin yin"
pinyin("汉语拼音", { toneType: "num" }); // "han4 yu3 pin1 yin1"
pinyin("睡着了"); // "shuì zháo le"
```

Three options do the work. `type` decides whether you get a joined string or an array of per-character readings, `toneType` decides whether the tone appears as a mark, as a trailing digit, or not at all, and the two combine freely, so `{ toneType: "none", type: "array" }` gives you `["han", "yu", "pin", "yin"]`.

The last line is the one that separates this from a lookup table. 睡着 is polyphonic, and the library resolves it to shuì zháo le without you passing a dictionary. The feature list also names initials, finals, first letters, tones and all information as separate outputs, plus a surname mode, custom pinyin, every pinyin a character has, and conversion of pinyin typed back into Chinese.

## match() accepts zwp, zhongwenpin, or zhongwp

Matching is the feature that turns the library into a search primitive, and it is three lines:

```js
import { match } from "pinyin-pro";
match("中文拼音", "zwp"); // [0, 1, 2]
match("中文拼音", "zhongwenpin"); // [0, 1, 2]
match("中文拼音", "zhongwp"); // [0, 1, 2]
```

The return value is an array of matched character indices, so `[0, 1, 2]` means the query covered the first three characters of 中文拼音. That shape is what you want for highlighting results or for narrowing a candidate list, and it is why the function is named match rather than includes.

What the three examples demonstrate is that the same call accepts three query styles: initials only, full pinyin, and a mixture of the two. For a Chinese input method or a search box over a Chinese catalogue, that is the difference between three separate matchers and one, and the cost is the `match` module, which the size table puts at 185.54 KB before gzip against 306.41 KB for the main pinyin API.

## convert() moves between numeric, symbol and toneless forms, erhua included

Format conversion is the cheap API and the odd one out. `convert("pin1 yin1")` gives `pīn yīn`, the reverse needs `{ format: "symbolToNum" }`, and `{ format: "toneNone" }` strips the marks. The interesting cases are the erhua forms, where the r suffix is a real part of the syllable rather than decoration:

```js
import { convert } from "pinyin-pro";
convert("dou4 zhi1r") // dòu zhīr
convert("dòu zhīr", { format: "symbolToNum" }) // dou4 zhi1r
convert("dòu zhīr", { format: "toneNone" }); // 'dou zhir'
```

Round tripping an erhua syllable keeps the r in all three forms, which is the kind of detail that separates a converter that understands the notation from one that only maps characters.

And this API carries almost nothing. The generated size table lists convert at 1.78 KB, 0.98 KB gzipped, against roughly 306 KB for pinyin. If all you need is rewriting tones, you do not pay for a dictionary, which makes convert the one function here that is safe to add to any bundle without asking.

## segment() returns words with their pinyin already joined

Word segmentation is bundled rather than delegated, and the return value keeps each word and its reading together:

```js
import { segment, OutputFormat } from "pinyin-pro";
segment("我喜欢学习汉语");
// [
//   { origin: "我", result: "wǒ" },
//   { origin: "喜欢", result: "xǐhuān" },
//   { origin: "学习", result: "xuéxí" },
//   { origin: "汉语", result: "hànyǔ" }
// ]
segment("我喜欢学习汉语", { format: OutputFormat.PinyinString });
// "wǒ xǐhuān xuéxí hànyǔ"
```

The default gives you objects with an `origin` and a `result`, which is what you want when a downstream step needs the word boundaries. The `OutputFormat.PinyinString` option collapses the same call into a single string of per-word pinyin.

Note what the readings look like inside a word. 喜欢 comes back as xǐhuān, one string with no separator, and 学习 as xuéxí, so the joining is per word rather than per character. That differs from the main pinyin API, which separates every character with a space, and it is the detail to check if you plan to line the output up under the source text. The feature list also credits Chinese word segmentation as a first class capability, along with surname pinyin for names.

## html() emits ruby markup with class names you can style

The fifth API exists for reading aids and it returns HTML rather than data:

```js
import { html } from "pinyin-pro";
html("汉语拼音");
```

The output is ruby annotation, one `<ruby>` element per character, with the character in a `py-chinese-item` span, the reading in an `rt` with class `py-pinyin-item`, and `<rp>` fallbacks around the reading so a browser without ruby support shows parentheses. Each character is wrapped in an outer `py-result-item` span.

Those four class names are the contract. They are fixed strings in the output rather than options, so styling the annotation is a CSS question, not a configuration question, and a caller who needs different markup has to post-process the string. If you are building something where the pinyin should be selectable, sortable or hidden behind a toggle, that is the seam to look at first, because the library hands you markup and hands you nothing else about it.

## The accuracy and speed numbers come from scripts in the repository

The README carries a comparison against two other packages, `pinyin` and `@napi-rs/pinyin`, and it points at the scripts that produce it: an accuracy file at packages/pinyin-pro/scripts/benchmark/accuracy.ts and a speed file at packages/pinyin-pro/scripts/benchmark/speed.ts. Both competitors are devDependencies of the root package.json, which is how a benchmark can call the libraries it is measuring.

| Comparison | pinyin | @napi-rs/pinyin | pinyin-pro |
| --- | --- | --- | --- |
| Accuracy, Node build | 94.097% | 94.097% | 99.846% |
| Accuracy, Web build | 91.170% | not supported | 99.846% |
| First dictionary init | 14.261ms | 160.769ms | 8.412ms |
| 10k characters | 74.442ms | 4.298ms | 7.216ms |
| 100k characters | 6287.332ms | 29.32ms | 45.471ms |
| 1m characters | out of memory, conversion failed | 297.41ms | 328.338ms |
| 10m characters | out of memory, conversion failed | 3907.278ms | 3375.192ms |
| Web environment | supported | not supported | supported |
| Node environment | supported | supported | supported |

Read it as a maintainer's benchmark, not as a neutral one, and note where it does not lead. `@napi-rs/pinyin` is faster on the 10k and 100k conversions, 4.298ms against 7.216ms and 29.32ms against 45.471ms, while pinyin-pro is faster on first dictionary initialisation at 8.412ms. pinyin-pro is also the only one of the three that completes at one and ten million characters, and the only one besides pinyin that runs in a browser.

## Tree shaking works per API, and convert is the exception

There is a second generated table, produced by `pnpm size`, that answers the question the first one raises. Each API is bundled separately under ESM with tree shaking enabled, while the UMD build cannot be shaken and ships one file for everything at 316.72 KB, 138.03 KB gzipped.

| API | ESM size |
| --- | --- |
| pinyin | 306.41 KB (gzip 134.52 KB) |
| segment | 305.16 KB (gzip 133.75 KB) |
| match | 185.54 KB (gzip 80.87 KB) |
| html | 307.25 KB (gzip 134.83 KB) |
| polyphonic | 180.75 KB (gzip 78.83 KB) |
| convert | 1.78 KB (gzip 0.98 KB) |
| Total | 559.55 KB (gzip 157.79 KB) |

The shape of those numbers is the design. Anything that needs the character dictionary costs about 134 KB gzipped, whether you want pinyin, segmentation, HTML or polyphonic lookup, because the dictionary is the payload. Only convert escapes it, at 1.78 KB, since it rewrites notation without consulting a dictionary at all.

So the decision is not which API to call but whether you need the dictionary. A page that only reformats tones takes 1.78 KB. A page that annotates Chinese text takes roughly a third of that budget gzipped before your own code is counted.

## A pnpm workspace of three packages, on Node 18 or newer

The repository is a monorepo, and the published library is one package inside it. `packages/pinyin-pro` is the core package, `packages/data` is the `@pinyin-pro/data` extension dictionary with its data processing scripts, and `packages/docs` holds the Chinese and English VitePress documentation. The root package.json is named pinyin-pro-repository, marked private, and its version field reads 1.0.0, which is unrelated to the library versions: the recent releases are 3.29.4 on 2026-09-11, 3.29.3 on 2026-08-19 and 3.29.2 on 2026-08-15.

The tooling is pinned in the file. `packageManager` is pnpm@10.8.0 and `engines` asks for Node 18 or newer, so an older runtime is out by declaration rather than by surprise. The scripts delegate into the workspace with filters, for example `pnpm --filter pinyin-pro test` and `pnpm --filter @pinyin-pro/data build`, and there are dedicated ones for the benchmarks and the size report: `compare`, `speed`, `accuracy` and `size`. Documentation builds are separate again, with `build:docs`, `deploy:docs` running `bash ./deploy.sh`, and a dev server per language.

Installing the library is the short version of all this. `npm install pinyin-pro` for a project, or a script tag from unpkg for a page with no build step.

## Conclusion

pinyin-pro fits a JavaScript or TypeScript product that needs annotated Chinese text, search that accepts zhongwp as well as full pinyin, or ruby markup for a reading aid, and that runs in both Node and the browser. It does not fit a bundle budget that cannot absorb roughly 134 KB gzipped of dictionary for the main API, where the convert API at 1.78 KB is the cheap exception. Before you commit, run the accuracy and speed scripts yourself, since the published comparison is the maintainer's own, and check that your Node version clears the stated floor of 18.

## FAQ

### How do I install pinyin-pro?

From npm with `npm install pinyin-pro`, or in a page with no build step by adding `<script src="https://unpkg.com/pinyin-pro"></script>`. The package is MIT licensed, and the repository asks for Node 18 or newer.

### How do I convert Chinese text to pinyin without tone marks?

Pass `{ toneType: "none" }`, which turns `pinyin("汉语拼音")` from 'hàn yǔ pīn yīn' into "han yu pin yin". Combine it with `{ type: "array" }` to get `["han", "yu", "pin", "yin"]`, or use `{ toneType: "num" }` for trailing digits instead.

### How does pinyin-pro handle polyphonic characters?

It resolves them from context without a dictionary, so `pinyin("睡着了")` returns "shuì zháo le". The feature list also covers surnames, custom pinyin overrides, and returning every pinyin a character has, and the `polyphonic` API is listed separately in the size table.

### How big is the pinyin-pro bundle?

The generated size table lists the ESM `pinyin` API at 306.41 KB, 134.52 KB gzipped, with `segment` and `html` at a similar weight and `match` at 185.54 KB. The exception is `convert` at 1.78 KB, 0.98 KB gzipped, since it rewrites notation without a dictionary, and the UMD build cannot be tree shaken and is 316.72 KB for all APIs.

## Sources

- [License: MIT](https://github.com/zh-lx/pinyin-pro/blob/main/LICENSE)
- [Project website](https://pinyin-pro.cn)
- [README](https://github.com/zh-lx/pinyin-pro/blob/main/README.md)
- [Releases](https://github.com/zh-lx/pinyin-pro/releases)
- [zh-lx/pinyin-pro on GitHub](https://github.com/zh-lx/pinyin-pro)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/zh-lx-pinyin-pro
