Model or dataset
dongrixinyu/JioNLP avatar
dongrixinyu/JioNLP

JioNLP: the Chinese NLP toolbox whose tables read like a feature catalog

中文 NLP 预处理、解析工具包,准确、高效、易用 A Chinese NLP Preprocessing & Parsing Package www.jionlp.com

3,872 stars442 forksPythonApache-2.0

At a glance

What is it?
JioNLP is dongrixinyu's Apache-2.0 Python toolkit for Chinese NLP preprocessing and parsing, installable with pip install jionlp, covering text cleaning, information extraction, data augmentation, parsing gadgets like license plate, ID card and time semantics parsing, plus NER, classification and model acceleration baselines. Its feature tables rate each function with stars, its requirements are four packages, and the Chinese-language wiki documents every entry.
Who is it for?
Use JioNLP when a Chinese NLP pipeline needs the unglamorous layers done well, text cleaning before training, extraction of money, phone, ID and location strings after, augmentation for scarce labeled data, and rule-based parsing of China-specific formats like license plates, ID numbers and lunar calendar dates, since these are the catalog's starred core. It is a rules-and-resources toolkit, not a model zoo, so pair it with your own models.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 64 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The gadget catalog, star-rated

The gadgets table is the project's core inventory, and its star column is an editorial layer most libraries skip. The starred entries define what the author considers best of class, license plate parsing, time semantic parsing producing timestamps and durations, key phrase extraction, stopword filtering, sentence splitting by punctuation, location parsing identifying province, city, county, township and village from an address string, news location recognition across domestic provinces and foreign countries, ID card parsing recovering province, city, county, birth date, gender and checksum, pinyin conversion returning initials, finals and tones, and character radical analysis returning radicals, structure, four-corner codes, decomposition and Wubi codes. Unstarred but present, lunar and solar date conversion in both directions, traditional to simplified conversion with per-character and maximum matching modes, money number to Chinese characters, new word discovery, idiom solitaire, and phone number carrier and location parsing.

Data augmentation for Chinese text

The augmentation table lists five methods tuned for Chinese text. Back translation, starred, uses the major cloud platforms' machine translation APIs to paraphrase. Homophone substitution, starred, replaces words sharing pronunciation, a transformation specific to Chinese where sound and character diverge. Adjacent character swapping randomly exchanges nearby characters, random add and delete inserts or removes a character without changing meaning, and NER entity substitution, starred, swaps entities from a dictionary so sequence labeling and classification data grow without breaking semantics. Each function is a callable transformation over text or token lists, so an augmentation pass is a pipeline stage rather than a research project.

Regular extraction and the clean pipeline

The regex section covers the extraction layer every Chinese text pipeline needs. clean_text, starred, removes abnormal and redundant characters, HTML tags, bracketed content, URLs, emails and phone numbers, and converts full-width letters and digits to half-width, the normalization pass that precedes everything. Extraction functions return positions and types, email with domain, money amounts with parsing, WeChat IDs, phone numbers mobile and landline, ID cards feeding into parse_id_card for details, QQ numbers under strict and loose rules, URLs, IP addresses, bracketed content across the full Chinese bracket set, and license plates. The removal twins, remove_email, remove_url, remove_phone_number and remove_ip_address, scrub the same patterns out, the redaction side of the same rules.

Four dependencies, one of them the tokenizer

Installation is one command:

code
$ pip install jionlp

and the first-use pattern from the README:

code
>>> import jionlp as jio
>>> print(jio.__version__)  # 查看 jionlp 的版本
>>> dir(jio)
>>> print(jio.extract_parentheses.__doc__)

The requirements file is four lines, numpy, jiojio, requests and zipfile36. The jiojio dependency is the author's own Chinese word segmentation engine, so the tokenization that stopword filtering and augmentation consume comes from a sibling project rather than jieba or pkuseg, and zipfile36 handles the dictionary resources the package ships. The setup.py extracts the version number from the README's version badge with a regex rather than a version file, packages everything except the test module, and carries the Apache 2.0 license with the author's contact. The lean dependency set is a deliberate portability choice, installable into most environments without conflict.

MELLM, and the LLM-era additions

The news section's most recent named addition is MELLM, Mutual Evaluation of Large Language Models, an automatic evaluation algorithm for LLMs without human supervision, tested across several LLMs and datasets with results published. Running it means downloading norm_score.json and max_score.json from a Baidu netdisk with the password jmbo, cloning the repository, and running python test_mellm.py from the test directory, with the test file documenting the download path if errors appear. The surrounding article list, on the pipeline to end-to-end shift, why not to trust LLM benchmarks, and ChatGPT's principles, marks the author's engagement with the LLM era from a classical NLP vantage point, and the FFIO project is linked as a separate release.

Discovery by Ctrl-F, and the wiki behind the tables

The introduction instructs users to scroll the page and search with Ctrl-F for specific features, an admission that the README is a catalog to be searched rather than a narrative to be read, and the help gadget formalizes this, if you do not know what JioNLP has, type keywords at the command line prompt to search the functions. Every table row links into the project wiki's per-feature documentation pages, so the README is the index, the wiki is the manual, and the import example closes the loop, import jionlp as jio, print the version, run dir on the module, print a function's docstring. A WeChat public account under the same name publishes AI news, and time semantic parsing customization is offered through a direct WeChat contact.

Editorial conclusion

Use JioNLP when a Chinese NLP pipeline needs the unglamorous layers done well, text cleaning before training, extraction of money, phone, ID and location strings after, augmentation for scarce labeled data, and rule-based parsing of China-specific formats like license plates, ID numbers and lunar calendar dates, since these are the catalog's starred core. It is a rules-and-resources toolkit, not a model zoo, so pair it with your own models. Before adopting, note the jiojio dependency provides the word segmentation, use the help gadget's keyword search to find functions, and consult the wiki pages each table links, since the README is the index rather than the manual.

Frequently asked questions

What is JioNLP?

JioNLP is an Apache-2.0 Python toolkit for Chinese NLP preprocessing and parsing, installed with pip install jionlp. It provides text cleaning, information extraction for money, phone numbers, ID cards and URLs, data augmentation methods, and rule-based parsing gadgets for license plates, addresses, ID numbers, time semantics and lunar calendar dates.

How do you use JioNLP?

Import jionlp as jio and call the functions directly, for example jio.parse_time for time semantic parsing, jio.parse_location for address parsing, jio.clean_text for normalization, and jio.extract_parentheses for bracketed content. The help gadget searches available functions by keyword, dir(jio) lists the module, and each function's docstring prints usage.

Does JioNLP include data augmentation?

Yes, five methods are included, back translation through cloud machine translation APIs, homophone substitution, adjacent character swapping, random character addition and deletion, and NER entity replacement from dictionaries, each a callable transformation applicable to sequence labeling and text classification datasets.

Official sources

  1. dongrixinyu/JioNLP on GitHub
  2. Issues
  3. License: Apache-2.0
  4. Project website
  5. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/dongrixinyu-jionlp.svg)](https://hysenlabs.com/projects/dongrixinyu-jionlp)