# jieba-rs ships a C API and a proc macro that its own project file never mentions

> A Rust port of the Jieba Chinese word segmenter, published as a cargo workspace with four members, an embedded compressed dictionary, and two optional keyword extractors. Every performance claim in the project file is a link to somebody's blog rather than a number.

**messense/jieba-rs** — The Jieba Chinese Word Segmentation Implemented in Rust

- Repository: https://github.com/messense/jieba-rs
- Stars: 987 · Forks: 67
- Language: Rust
- License: MIT
- Published: 2026-09-16 · Updated: 2026-09-16 · Language: en
- Canonical page: https://hysenlabs.com/projects/messense-jieba-rs

## A four member workspace, and two of the members are undocumented

The published crate is one member of a cargo workspace, and the manifest is more informative than the project file.

```toml
[workspace]
resolver = "3"
members = ["capi", "jieba", "jieba-macros", "examples/weicheng"]
```

The main crate is jieba. The other three are capi, a C interface; jieba-macros, a procedural macro crate; and examples/weicheng, a single example promoted to a workspace member so that it is compiled and versioned with the library rather than left as a loose file.

The project file mentions none of that. It shows one dependency line, one example, three optional features and a list of bindings in other languages, which is enough to use the crate and not enough to know what the repository contains.

So a C developer will find a C API and a fixed string dependency to match it, and a Rust developer who hits a macro will find a second crate in the tree that nothing in the file explains. Neither is hidden, and neither is announced.

The manifest also fixes the version in one place. The workspace package version is 0.11.0 and the internal dependencies point at each other by path with matching versions, so the two crates in this repository cannot drift apart.

## The dictionary is compressed into the binary and SIMD is chosen at run time

Two dependencies explain two things about how this works.

The first is a compression crate, include-flate, which embeds its compressed data directly in the build. That is how the dictionary gets into the binary: the embedded dictionary is on by default through the default-dict feature, and the mechanism for that is a compressed blob included at compile time rather than a file loaded at run time. There is no data directory to ship and no lookup path at startup, at the cost of a binary that carries the dictionary whether or not you segment anything.

The second is bytecount, requested with the runtime-dispatch-simd feature. That name is the whole point: the SIMD implementation is selected when the program runs, not when it is compiled. The practical effect is that the fast path is available on a machine that supports it without a rebuild, and that a machine without it takes a slower path with no error and no warning.

The rest of the dependency list reads like a keyword extractor's shopping list, which matches the optional features: an ordered float type, a fast hash map, and an expect-test snapshot crate for assertions. Rayon is there for the parallel path, and a codspeed-compatible criterion adapter for the benchmarks.

Those two dependencies are the reason a drop-in swap is not as cheap as swapping a string, and the reason a machine that lacks SIMD silently runs a different code path.

## Keyword extraction is two features you have to ask for

Segmentation is the default. Keyword extraction is not, and the two extractors are separate opt-ins.

```toml
[dependencies]
jieba-rs = { version = "0.11", features = ["tfidf", "textrank"] }
```

The tfidf feature enables the TF-IDF keywords extractor and the textrank feature enables the TextRank one, while default-dict enables the embedded dictionary and is on by default. So the minimal dependency gives you a segmenter, and anything that ranks words in a document costs a feature flag.

That split is worth knowing because the natural next question after segmenting a document is which words in it matter, and the answer is not in the box. A second crate for the embedding models would be the obvious alternative; here the ranking is lexical, computed from the segmenter's own output, which is a different and much cheaper thing.

The benchmark instruction reflects the same structure. To measure everything, not just the default path:

```bash
cargo bench --all-features
```

Nothing in the project file states what any of this costs in time or memory. There is no size for the embedded dictionary, no throughput figure, and no comparison table. The only numbers anywhere in the repository live in blog posts.

## Every performance number is a link, and the links do not measure the same thing

The performance section of the project file is five links and zero numbers.

Two are the author's own: a record of making the segmenter 2.4 times faster, in Chinese and in English. Three are by another author, on being 33 percent faster than cppjieba, again in English, Simplified Chinese and Traditional Chinese.

So the comparison with the C++ implementation everyone actually benchmarks against is not in the repository, and the headline figures come from two different sets of posts by two different people, comparing against two different baselines. A 2.4 times speedup and a 33 percent lead over cppjieba are not the same claim, and the file does not say which version of which code produced either.

What the repository does contain is the machinery to check: a criterion-compatible benchmark adapter, a CI workflow badge, a coverage badge, and the instruction to run the suite with all features enabled. A comparison is reproducible by whoever wants to do the work.

The honest reading is that the performance story is maintained in a blog and the code is maintained in a repository, and the two drift apart on their own schedule.

## Nine bindings live in nine other repositories

The bindings list is the longest section in the project file, and it is entirely outbound links.

Node has a binding through the napi-rs project, and so does a separate WebAssembly pair: one for the web and one framed for web and Node. PHP, Python and R each have one. There is a Chinese tokenizer for the Tantivy search engine, plus a separate adapter that bridges Tantivy and this segmenter, so the search-engine route has two projects rather than one. And there is an Emacs binding that also ships a minor mode extension.

The practical consequence is about versions. A JavaScript developer who reads the install line for version 0.11 and then installs the Node binding gets whatever version that package pins, which is not necessarily 0.11, and there is nothing in this repository that keeps them in step.

The same applies to the two WebAssembly bindings, which are the ones people reach for when there is no Rust in the stack, and to the two Tantivy projects, where it is not obvious from the list which one to use.

Read the list as a map of who has built what, not as a compatibility table. The upstream project maintains the segmenter; the language ports are maintained by whoever wrote the port.

## The whole usage example is six lines, and one argument is unnamed

The example in the project file is the entire usage documentation.

```rust
use jieba_rs::Jieba;

fn main() {
    let jieba = Jieba::new();
    let words = jieba.cut("我们中出了一个叛徒", false);
    assert_eq!(words, vec!["我们", "中", "出", "了", "一个", "叛徒"]);
}
```

One sentence in, six tokens out, asserted rather than printed. The sentence is chosen well for a segmenter, since it mixes a single character with a multi character word, and the expected output is given exactly, which means you can tell whether your build agrees with the published one before you write anything else.

The gap is the second argument to cut. It is a boolean, the example passes false, and nothing in the file says what it controls or which value to pass instead. Anyone who wants the other setting has to read the source or the documentation site, and the assertion above only covers the one path the example takes.

There is no second example for the keyword extractors, none for the C API and none for the WebAssembly route. The six lines above are the whole contract.

## Edition 2024 in the manifest and Rust 2015 advice in the install note

Two details in the installation section are worth reading together with the manifest.

The manifest sets the workspace edition to 2024, and the workspace resolver to version 3, which is the modern setting. The install note, by contrast, tells Rust 2015 users that they must add an extern crate declaration for the library to their crate root. That advice is a leftover from an era when it was still required, and it is the only place in the project file that mentions an edition at all.

So the stated floor for using this crate is now two full editions behind the edition it is built with. Nobody following that note on a current toolchain would be harmed by it, and nobody would notice the advice was stale, which is a decent description of most compatibility notes.

The install instruction itself is the minimum possible surface. One line under dependencies, then a single statement that you are good to go.

## Conclusion

Use jieba-rs when you are segmenting Chinese text inside a Rust program and want one dependency rather than a Python process. Read three things first. The manifest is more than the library: there is a C API crate and a proc macro crate in the same workspace, neither described in the project file, so check whether one of them is the version you actually want. The embedded dictionary is on by default and compiled into the binary, which is convenient and is also a fixed cost you cannot turn off without losing segmentation. And the performance comparisons are all external, in blog posts in four languages, with figures that measure different things. Take them as links to read, not as numbers to quote.

## FAQ

### What is jieba-rs?

A Rust implementation of the Jieba Chinese word segmentation algorithm, published on crates.io as jieba-rs under the MIT license. It is developed as a cargo workspace whose members are the main jieba crate, a C API crate, a proc macro crate and one example.

### How do I use jieba-rs in a Rust project?

Add jieba-rs = "0.11" under [dependencies], construct Jieba::new(), and call cut with your text. The example passes a sentence and gets six tokens back. If you are on Rust 2015 you also need an extern crate declaration for it in the crate root, though the crate itself is built on edition 2024.

### Does jieba-rs include keyword extraction?

Not by default. The tfidf and textrank features enable the TF-IDF and TextRank keyword extractors, while the embedded dictionary is enabled by default through default-dict. Run cargo bench --all-features to measure the whole set, since the default build is segmentation only.

## Sources

- [Issues](https://github.com/messense/jieba-rs/issues)
- [License: MIT](https://github.com/messense/jieba-rs/blob/main/LICENSE)
- [messense/jieba-rs on GitHub](https://github.com/messense/jieba-rs)
- [README](https://github.com/messense/jieba-rs/blob/main/README.md)
- [Releases](https://github.com/messense/jieba-rs/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/messense-jieba-rs
