HanLP gives three routes to every task, and the annotation standard you choose decides what your tokens mean
中文分词 词性标注 命名实体识别 依存句法分析 成分句法分析 语义依存分析 语义角色标注 指代消解 风格转换 语义相似度 新词发现 关键词短语提取 自动摘要 文本分类聚类 拼音简繁转换 自然语言处理
At a glance
- What is it?
- HanLP is an Apache-2.0 Python toolkit for Chinese and multilingual NLP, built on PyTorch and TensorFlow 2.x, covering 130 languages and ten joint tasks. Its own table is the most useful artifact here, because it shows which tasks have local tutorials and which annotation standard each model emits.
- Who is it for?
- Pick HanLP if your work is Chinese or multilingual text and you want parsing depth that a segmenter alone will not give you, and pick the multitask route when you need several tasks at once, since it is the one that shares a model across them. Do not pick it because a task appears in the description; AMR and coreference have gaps in the table, and the TensorFlow path disappears on Intel macOS with Python 3.13.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 16 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Every task in the table has three columns, and they are not interchangeable
The capability table is organized by task, and each row carries the same three routes before the model and annotation links. Tokenization, part-of-speech tagging, named entity recognition, dependency parsing, constituency parsing, semantic dependency parsing, semantic role labeling, abstract meaning representation and coreference each have a RESTful tutorial, a multitask tutorial and a single-task tutorial, followed by a pretrained model page and a set of annotation standards. Consequence: RESTful sends your text to a hosted service, multitask loads the joint model that handles ten tasks at once across 130 languages, and single-task loads one dedicated model for the one task you care about. Picking a route changes where your text goes, how much you download, and how many tasks you can run in a single pass, so choosing late means rework.
The annotation column is where two HanLP tokenizations stop being interchangeable
Each row ends with the standards that model can emit, and they are not the same granularity. Tokenization comes in a coarse MSR segmentation and a fine CTB segmentation. Part-of-speech tagging offers CTB, PKU and 863. Named entity recognition offers PKU, MSRA and OntoNotes, so a model trained on OntoNotes brings English-derived entity conventions into a Chinese pipeline. Dependency parsing offers SD, UD and PMT, constituency parsing uses the Chinese Tree Bank, semantic dependency parsing uses CSDP, and semantic role labeling uses the Chinese Proposition Bank. Consequence: swapping one HanLP model for another can change where token boundaries fall without changing any of your code, which silently invalidates character offsets, span indices and any downstream lookup keyed on token position.
AMR and coreference are the two holes in the capability table
Not every task is filled in, and the gaps are visible in the table itself. For abstract meaning representation, the multitask column reads 暂无, meaning none, while a single-task tutorial and a model page are present. For coreference, the multitask and single-task columns both read 暂无, leaving the RESTful route as the only tutorialed path in that row. The AMR dependency is real too, and setup.py shows it as the amr extra pulling penman==1.2.1, networkx>=2.5.1 and perin-parser>=0.0.12, with perin-parser pinned to a floor rather than a range. Consequence: if abstract meaning representation or coreference is on your roadmap, the local model story for those two tasks is thinner than the description suggests, and for coreference the table points you at a hosted service rather than a model you run yourself.
The tf extra is a matrix of narrow version windows, and one of them excludes Intel Macs
The dual engine claim depends on an extra named tf, and its contents are a set of conditional pins rather than a single range. Python below 3.12 gets tensorflow>=2.6.0,<2.14. Python 3.12 and 3.13 get tensorflow>=2.16,<2.17 with tf-keras>=2.16,<2.17. Python 3.13 and 3.14 get tensorflow>=2.21,<2.22 with tf-keras>=2.21,<2.22, and that last pair carries the condition sys_platform != "darwin" or platform_machine != "x86_64", so it does not apply to Intel Macs. transformers<4.55 is pinned for Python below 3.14 with an inline comment that TF is deprecated in Transformers and no longer maintained. Consequence: an Intel Mac on Python 3.13 satisfies no branch of the tf extra, so the TensorFlow half of the toolkit quietly drops out and you are left with PyTorch alone, and the extras_require entry named full is the union of every extra, so asking for everything pulls the entire matrix in at once.
Python 3.6 and 3.7 on macOS and Windows get hardcoded pins that exist because a build failed
The first thing setup.py computes is an empty EXTRAS list, then fills it only when sys.platform is darwin or win32. On Python 3.6 the list becomes tokenizers==0.10.3, an exact pin. On Python 3.7 it becomes safetensors<0.5, with the comment that it failed to build safetensors. The condition is scoped to those two platforms, so Linux installations never enter this branch. Consequence: on a Mac or a Windows box running an end-of-life interpreter, pip resolves to a tokenizers build that no current release targets, so a fresh environment there is a source build rather than a wheel install, and the comment above the pin records that this was a workaround for a compilation failure rather than a tested combination. Check your interpreter before you rely on the extra.
The default branch is doc-zh, so a bare clone lands in documentation
The repository's default branch is doc-zh, not master, and the language switch at the top of the README sends English readers to tree/master and Japanese readers to doc-ja. The tutorials referenced from the table are notebook files under plugins/hanlp_demo/hanlp_demo/zh/, and the run-in-browser link encodes the branch and path directly, pointing at the doc-zh branch and a zh tutorial notebook. Consequence: git clone without a branch argument hands you the documentation tree rather than the code branch you meant to work from, and the tutorial path encodes a plugins layout that you will not find on master. The top-level entries confirm where the code lives, with hanlp/, tests/, docs/, plugins/ and a setup.py that is the real entry point for installation.
v1.8.6, v2.1.0 and v2.1.1 share one tag namespace and two of them share a day
The release tags do not sort into a clean line. v1.8.6 is labelled 常规维护, meaning routine maintenance, and is dated 2024-12-29. v2.1.0, labelled English Support, is dated 2024-12-29 as well, minutes later. v2.1.1, labelled Ancient Chinese Support, is dated 2025-01-13. The most recent push to the repository is dated 2026-09-15, which puts the newest tag around twenty months behind the working tree. Consequence: the same tag namespace carries both a 1.x maintenance line and the 2.1 feature line, so a tag name alone does not tell you which major version of the API you get, and the version pip installs comes from hanlp/version.py read by setup.py rather than from the tag you thought you were selecting. Anchor on the package version, not the tag.
Editorial conclusion
Pick HanLP if your work is Chinese or multilingual text and you want parsing depth that a segmenter alone will not give you, and pick the multitask route when you need several tasks at once, since it is the one that shares a model across them. Do not pick it because a task appears in the description; AMR and coreference have gaps in the table, and the TensorFlow path disappears on Intel macOS with Python 3.13. Before you install, read setup.py rather than the docs page, because the version windows and the macOS and Windows pins live there, and decide which annotation standard your downstream offsets depend on before you generate a single token.
Frequently asked questions
What is NLP used for in HanLP?
HanLP's own task list runs from word segmentation, part-of-speech tagging and named entity recognition through dependency and constituency parsing, semantic role labeling and coreference resolution. It also covers style transfer, semantic similarity, new word discovery, keyword and phrase extraction, summarization, text classification and clustering, and simplified to traditional Chinese conversion.
What are the four types of NLP in HanLP?
The documentation does not divide NLP into four types. What it states for HanLP 2.1 is support for 130 languages including simplified and traditional Chinese, English, Japanese, Russian, French and German, across ten joint tasks plus a range of single tasks.
What are the 5 stages of NLP in HanLP?
No five-stage pipeline is described. HanLP organizes its capabilities as joint multitask models and single-task models, each reachable through a RESTful tutorial, a multitask tutorial or a single-task tutorial, with a pretrained model page and an annotation standard named for every task.
What is multilingual NLP and how does HanLP do it?
HanLP 2.1 covers 130 languages on a multilingual corpus and offers multitask models that handle ten joint tasks at once, alongside dedicated single-task models. Each task links to its own pretrained model page, and the default branch of the repository is doc-zh, where the zh tutorial notebooks live.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/hankcs-hanlp)