Library / SDK
BYVoid/OpenCC avatar
BYVoid/OpenCC

OpenCC: Deterministic Traditional and Simplified Chinese Conversion

Library for conversion between Traditional and Simplified Chinese

10,020 stars1,073 forksC++Apache-2.0

At a glance

What is it?
OpenCC is a dictionary-driven C++ library and CLI for converting between Traditional Chinese, Simplified Chinese, regional wording and Japanese Shinjitai. It is for pipelines that need repeatable, offline output rather than a model that guesses.
Who is it for?
Adopt OpenCC when you need repeatable, offline conversion between Traditional and Simplified Chinese, regional wording for Mainland China, Taiwan and Hong Kong, or the limited Japanese Shinjitai and Kyujitai mapping, and when you are willing to pin a dictionary version and re-run your own test corpus on upgrade.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 3 days ago.
What is it written in?
Mainly C++, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem OpenCC solves, and who actually needs it

Chinese text is not one target. A product page written in Simplified Chinese for Mainland readers uses 鼠标, while the same page for Taiwan uses 滑鼠. Character variants split further: 裏 and 裡 are both legitimate Traditional forms, but which one a Taiwanese reader expects differs from what a Hong Kong reader expects. Handling this with ad hoc string replacement tables breaks the moment a term has more than one plausible mapping.

OpenCC addresses that with reviewed dictionaries plus a conversion engine. The README describes the project as providing "dictionaries, a reusable library, conversion tools, and dictionary generation tools" for Traditional Chinese, Simplified Chinese, Japanese Shinjitai and regional wording across Mainland China, Taiwan and Hong Kong. The audience is therefore not end users typing into a web box, although an online converter exists at opencc.js.org. The real audience is engineers embedding conversion into a build step, a CMS import, a search indexer or a subtitle pipeline.

The stated design principle matters here: OpenCC follows "能分則不合", separate whenever distinguishable, and the README says it strictly distinguishes Simplified-Traditional mappings from character variant mappings. That is a deliberate choice to keep one-to-many entries auditable rather than collapsing them for convenience.

How the conversion engine works: dictionaries, configs, and no model

The README is explicit that conversion is "deterministic, dictionary-based conversion without large language models (LLMs)". There is no inference step and no network call, so the same input plus the same config plus the same dictionary version produces the same output. For a content pipeline that property is worth more than fluency.

The second structural decision is decoupling. The README states that dictionaries are decoupled from the library, "allowing custom modification and extension". A conversion is therefore parameterised by a config file rather than hard-coded into the binary. The README's own examples use config names such as s2t.json for Simplified to Traditional, and the online converter is linked with a config parameter in the URL. Configs live under data/config in the repository, dictionaries under data/dictionary, and data/scripts holds the generation scripts that build dictionaries from source data.

That layout explains the upgrade story. Because dictionaries ship as data, a dictionary correction is a data change, not an API change. The 1.4.2 release notes make the boundary concrete: the C++ ABI is the same as 1.4.1, SOVERSION 1.4, so downstream C++ programs upgrading from 1.4.0 or 1.4.1 do not need to relink. The same notes also mention a fix for loading legacy .ocd dictionaries on big-endian platforms, which previously produced an empty dictionary. That is the kind of failure a deterministic engine can still have: not a wrong guess, but a silently empty lookup table.

Installing OpenCC and running a first conversion

The README lists package manager routes for Debian, Ubuntu, Fedora, Arch Linux, macOS via Homebrew, WinGet, Bazel, Node.js and Python. On macOS the documented command is short:

bash
brew install opencc

After that, the native CLI is on your PATH. A first conversion uses a config name and input and output files. The npm CLI example in the README is the clearest template:

bash
opencc -c s2t.json -i input.txt -o output.txt

The -c flag selects the config, -i the input file, -o the output file. s2t.json is Simplified to Traditional. If your input is Traditional and your target is Simplified, you would pick the corresponding config from data/config rather than s2t.json.

For Python, the README gives `pip install opencc`, which installs both the Python API and a Python CLI. For Node.js, `npm install -g opencc` installs the Node.js CLI. The npm package requires Node.js >=20.17 and installs a native addon through prebuilt packages named @opencc/opencc-<platform>-<arch>, covering macOS x64 and arm64, Linux x64 and arm64, and Windows x64. The README warns that on platforms without a prebuilt binary, npm install builds from source with Bazel, compiling the addon and regenerating dictionaries, which needs a C++ toolchain and network access. If you are deploying to an unusual architecture, check that warning before you put opencc in a Dockerfile.

The npm CLI is deliberately narrower than the native one. The README states that plugins, --inspect, --segmentation and --ambiguities require the native OpenCC CLI. If your workflow depends on inspecting which dictionary entry fired, install the native binary, not the npm one.

Where OpenCC is the wrong tool

The most important limitation is stated plainly in the README: dictionaries are based on Mandarin vocabulary, and translation between languages, "such as between Mandarin and Cantonese, Southern Min, or Japanese", is not supported. If your requirement is Cantonese written form, this is not a configuration problem you can solve by editing dictionaries; the project excludes it by scope.

Japanese coverage is likewise bounded. The README says conversion between Japanese Shinjitai and Kyujitai is supported "to a limited extent". Limited is the operative word. Do not treat OpenCC as a Japanese normalisation library.

The second class of limitation is the npm surface. Because the npm CLI omits segmentation, ambiguity inspection and plugins, a team that prototypes with `npm install -g opencc` and later needs ambiguity data will have to switch distribution channels. That is a migration, not a flag.

The third is packaging. The README documents that source builds regenerate dictionaries and pull a hermetic Python toolchain through Bazel. That is a heavier build than a typical C++ library, and it is the reason prebuilt binaries exist for Windows x64 and for Debian and Ubuntu amd64 and arm64. The Debian and Ubuntu zips bundle the opencc, opencc-jieba and libopencc* deb packages for one architecture along with a SHA256SUMS file. If your target is not in that list, budget for a build environment rather than a download.

OpenCC compared with script-based conversion

The obvious alternative is a hand-maintained mapping table in your own code, or one of the community reimplementations that wrap a similar table. The difference is not speed, it is curation. OpenCC ships reviewed entries, including what the README calls "rigorously reviewed one-to-many Simplified-to-Traditional entries", and it separates Simplified-Traditional mapping from variant mapping. A local table usually merges those two concerns, which is exactly how 裏 and 裡 get flattened into one form.

The second difference is the config layer. With OpenCC you select behaviour by config name, and the README points to DESIGN_PRINCIPLES.md and doc/regional-phrase-criteria.md for how regional phrases are admitted. A hand-rolled table has no admission criteria, so every disputed term becomes a code review argument. OpenCC moves that argument into a documented standard.

The trade-off is that you inherit someone else's standard. If your product deliberately uses a variant that OpenCC's criteria exclude, you will be fighting the dictionaries rather than benefiting from them. The README does note that dictionaries can be modified and extended, so the escape hatch exists, but at that point you own the maintenance of your fork.

Maintenance, releases and licence

The repository is not archived, and the last push was on 2026-09-20. Recent releases are ver.1.4.0 on 2026-07-01, ver.1.4.1 on 2026-07-12 and ver.1.4.2 on 2026-08-22. The 1.4.2 notes describe a substantial speed-up of the conversion hot path, dictionary corrections, and the big-endian legacy dictionary fix, while keeping the C++ ABI at SOVERSION 1.4. For a C++ consumer that means an upgrade from 1.4.0 or 1.4.1 does not require relinking.

Upgrade cost is dominated by dictionary changes rather than API changes. A dictionary correction can change output for text you already converted, so re-running your own corpus is the only way to know what moved. The project provides a test target in the Makefile that configures a Debug build with ENABLE_GTEST and runs ctest, which is useful if you build from source and want to run the project's own tests before adopting a release.

Licensing is Apache-2.0, per the repository LICENSE file and the license field in package.json. Apache-2.0 is a permissive licence with an explicit patent grant and notice requirements. If you redistribute OpenCC or a modified dictionary set, read the licence text and your organisation's policy; this article is not legal advice.

Editorial conclusion

Adopt OpenCC when you need repeatable, offline conversion between Traditional and Simplified Chinese, regional wording for Mainland China, Taiwan and Hong Kong, or the limited Japanese Shinjitai and Kyujitai mapping, and when you are willing to pin a dictionary version and re-run your own test corpus on upgrade. Do not adopt it for Mandarin to Cantonese, Southern Min or Japanese translation, or for Japanese text beyond the limited Shinjitai and Kyujitai scope: the README states those are not supported. Before rollout, verify that the config you picked matches the region your content targets, and check whether your platform has a prebuilt npm binary or will compile the native addon with Bazel.

Frequently asked questions

What does OpenCC do?

It converts between Traditional Chinese, Simplified Chinese, Japanese Shinjitai and regional wording for Mainland China, Taiwan and Hong Kong, using reviewed dictionaries rather than a language model. It provides the dictionaries, a reusable library, conversion tools and dictionary generation tools.

How do I install OpenCC?

The README lists package manager routes including brew install opencc on macOS, winget install opencc on Windows, pip install opencc for the Python API and CLI, and npm install -g opencc for the Node.js CLI. Debian, Ubuntu, Fedora, Arch Linux and Bazel packages are also listed, and prebuilt Windows and Debian/Ubuntu binaries are published with each release.

How do I use OpenCC from the command line?

The README's npm CLI example runs opencc -c s2t.json -i input.txt -o output.txt, where -c selects the conversion config and -i and -o are the input and output files. The native CLI adds plugins, --inspect, --segmentation and --ambiguities, which the npm CLI does not support.

What is OpenCC?

Open Chinese Convert is an open source project for high-quality conversion between Traditional Chinese, Simplified Chinese, Japanese Shinjitai and regional wording, released under Apache-2.0. Its conversion is deterministic and dictionary-based, with no large language model involved, so it runs offline.

Official sources

  1. BYVoid/OpenCC on GitHub
  2. License: Apache-2.0
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/byvoid-opencc.svg)](https://hysenlabs.com/projects/byvoid-opencc)