Jieba segments Chinese text with four modes from precise to paddle
结巴中文分词
At a glance
- What is it?
- Jieba is a Python Chinese word segmentation module built on prefix dictionaries, dynamic programming and HMM unknown word detection. Four cutting modes and a 2024 last push shape the adoption call.
- Who is it for?
- Python teams doing Chinese text analysis, search indexing or keyword extraction should adopt Jieba for its documented modes and tunable dictionaries, after confirming the pinned paddle dependency fits. Teams needing a maintained deep learning first tokenizer or current release cadence should look elsewhere.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Probably not. The repository last received commits 25 months ago, on August 21, 2024.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Chinese text without spaces meets the module named stutter
"Jieba" is Chinese for "to stutter," and the module turns that joke into a job description: cutting unspaced Chinese character strings into words. The tagline aims high, built to be the best Python Chinese word segmentation module, with code compatible with both Python 2 and 3. Setup metadata at version 0.42.1 names Sun, Junyi as author, MIT as license, and classifiers covering simplified and traditional Chinese, OS independence and text processing topics. Repository layout stays small: a jieba package directory, test and extra_dict folders, setup.py, MANIFEST.in, Changelog and LICENSE. Default branch is master. Audience is Python developers analyzing Chinese text, indexing it for search, or extracting keywords, not developers working on spaced alphabetic languages where the problem never arises.
Four cutting modes trade precision against recall
One cutter cannot serve analysis and search equally, so Jieba ships four modes with stated tradeoffs. Precise mode cuts the sentence as accurately as possible and suits text analysis. Full mode scans out every word that could form, runs very fast, and cannot resolve ambiguity. Search engine mode starts from precise mode and re-cuts long words to raise recall, which suits search segmentation. Paddle mode trains sequence labeling with a bidirectional GRU network inside the PaddlePaddle deep learning framework and also supports part of speech tagging. Paddle mode needs paddlepaddle-tiny installed first, supports Jieba v0.40 and above, and loads lazily through the enable_paddle interface. Traditional Chinese segmentation and custom dictionaries round out the feature list.
A prefix dictionary graph plus HMM for unknown words
Segmentation rests on two documented mechanisms. A prefix dictionary scan builds every possible word formation in the sentence into a directed acyclic graph, then dynamic programming finds the maximum probability path, meaning the cut combination with the highest word frequency score. Unknown words fall to a second mechanism: an HMM model of character word forming ability decoded with the Viterbi algorithm. The page proves the point with a concrete readout where the string Hangyan was absent from the dictionary yet Viterbi still surfaced it as one token. Input strings may be unicode, UTF-8 or GBK, with one warning attached: GBK input is discouraged because it can decode into UTF-8 in unpredictable ways. That warning should be read as a preprocessing contract.
Cut functions expose four switches and two return shapes
Daily use runs through a small family of entry points. The jieba.cut method takes four inputs: the string to segment, the cut_all switch for full mode, the HMM switch for unknown word detection, and the use_paddle switch for paddle mode segmentation. The jieba.cut_for_search method takes two: the string and the HMM switch, and its fine granularity suits building inverted indexes for search engines. Both cut methods return an iterable generator for for loop consumption. The jieba.lcut and jieba.lcut_for_search variants return a plain list instead. A jieba.Tokenizer built with dictionary set to DEFAULT_DICT creates an independent segmenter for mixed dictionary setups, while jieba.dt remains the default instance behind every global function.
pip install jiebaInstallation is one command, and the module is referenced with a plain import. A first segmentation then reads naturally:
import jieba
seg_list = jieba.cut("我来到北京清华大学", cut_all=False)
print("Default Mode: " + "/ ".join(seg_list)) # 精确模式Expect the precise cut of the sample sentence with Tsinghua University kept whole. Flip cut_all to True for the full mode readout, or call cut_for_search when feeding an inverted index.
Custom dictionaries and frequency surgery fix domain words
Generic dictionaries fail on product names, place names and jargon, so Jieba lets developers load their own words. The jieba.load_userdict call takes a file object or a dictionary path, and the format mirrors dict.txt: one word per line, word plus optional frequency plus optional tag separated by spaces in fixed order, with UTF-8 encoding required for path inputs. Omitted frequencies get an auto computed value that keeps the word splittable. The page shows the payoff directly: before loading, a title about innovation office and cloud computing shatters into fragments, and after loading, innovation office and cloud computing survive as units. Programs can also call add_word and del_word for runtime edits, or suggest_freq to tune one word up or down. One caution is stated: auto computed frequencies can fail silently when HMM unknown word discovery is active. Restricted filesystems get relief through the tmp_dir and cache_file attributes of the segmenter.
>>> print('/'.join(jieba.cut('如果放到post中将出错。', HMM=False)))
如果/放到/post/中将/出错/。
>>> jieba.suggest_freq(('中', '将'), True)
494
>>> print('/'.join(jieba.cut('如果放到post中将出错。', HMM=False)))
如果/放到/post/中/将/出错/。
>>> print('/'.join(jieba.cut('「台中」正确应该不会被切开', HMM=False)))
「/台/中/」/正确/应该/不会/被/切开
>>> jieba.suggest_freq('台中', True)
69
>>> print('/'.join(jieba.cut('「台中」正确应该不会被切开', HMM=False)))
「/台中/」/正确/应该/不会/被/切开Each suggest_freq call returns the tuned frequency and the next cut keeps the intended unit unsplit. Rehearse this on domain terms before trusting default output.
Keyword extraction and tagging cover TF-IDF, TextRank and ictclas
Beyond cutting, Jieba extracts keywords and tags parts of speech. The jieba.analyse.extract_tags call takes the sentence, topK defaulting to 20 heaviest TF-IDF terms, a withWeight switch defaulting to False, and an allowPOS filter defaulting to empty. Builders can point IDF and stop word corpora at custom files through set_idf_path and set_stop_words, with sample corpora and usage scripts linked under extra_dict and test. TextRank extraction shares the interface with topK 20 and a default POS filter of ns, n, vn and v, building a graph over a fixed co-occurrence window defaulting to 5 and scoring nodes with PageRank on an undirected weighted graph, following the cited TextRank paper. Part of speech tagging uses jieba.posseg.POSTokenizer with ictclas compatible marks under the default jieba.posseg.dt instance, while paddle mode adds its own set of 24 lowercase POS tags plus PER, LOC, ORG and TIME name tags:
>>> import jieba
>>> import jieba.posseg as pseg
>>> words = pseg.cut("我爱北京天安门") #jieba默认模式
>>> jieba.enable_paddle() #启动paddle模式。 0.40版之后开始支持,早期版本不支持
>>> words = pseg.cut("我爱北京天安门",use_paddle=True) #paddle模式
>>> for word, flag in words:
... print('%s %s' % (word, flag))
...
我 r
爱 v
北京 ns
天安门 nsThe loop prints each word with its tag, showing pronoun, verb and place names in order. Compare default and paddle taggings on the same sentence before choosing a mode.
Pinned paddle versions and a 2024 push bound production use
Two limits deserve plain language. Paddle mode chains the reader to a dated pin: paddlepaddle-tiny at exactly 1.6.1, with older Jieba releases told to upgrade first. A pinned tiny framework from that era complicates fresh environments and security review, and lazy loading means the first paddle call pays the import cost. Freshness is the second limit. The last push landed on 2024-08-21 and the newest releases read v0.42.1, v0.42 and v0.41 from January 2020, so callers inherit six years of unchanged release lines and should state the last push date in any adoption memo rather than implying current maintenance. Against the default dictionary plus HMM pipeline, paddle mode stands as the in house alternative: a neural sequence labeler instead of frequency paths. Choose dictionary mode for transparency and zero extra installs, paddle mode when tagging accuracy justifies the dated dependency.
Editorial conclusion
Python teams doing Chinese text analysis, search indexing or keyword extraction should adopt Jieba for its documented modes and tunable dictionaries, after confirming the pinned paddle dependency fits. Teams needing a maintained deep learning first tokenizer or current release cadence should look elsewhere. Before building on it, install the exact release, test custom dictionary behavior on domain text, and note the last push of 2024-08-21.
Frequently asked questions
How do you use Jieba in a Python program?
Install with pip install jieba, reference it with import jieba, then call jieba.cut for generators or jieba.lcut for lists, choosing precise, full, search or paddle mode for the task.
What is Jieba?
Jieba, Chinese for to stutter, is a Chinese text segmentation module built to be the best Python Chinese word segmentation component, supporting four cutting modes, traditional Chinese and custom dictionaries.
What is Jieba in Python?
It is a Python 2 and 3 compatible module exposing cut, cut_for_search, lcut, Tokenizer, analyse keyword extraction and posseg tagging, installable with pip install jieba.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/fxsjy-jieba)