HIT-SCIR LTP: a Chinese NLP pipeline for segmentation, tagging and parsing
Language Technology Platform
At a glance
- What is it?
- LTP 4 splits into a Rust legacy model and a PyTorch multi-task model, both driven through one Pipeline API. The install is three pip packages, the model download needs Hugging Face access, and the two model families do not cover the same tasks.
- Who is it for?
- Adopt LTP if you need several Chinese analysis layers over the same text and want to choose between a fast legacy model and a fuller deep model. Do not adopt it if you need languages other than Chinese, or if you need semantic role labeling, dependency parsing or semantic dependency parsing from the fast path, because the legacy model covers only segmentation, POS and NER.
- Can I use it commercially?
- Not without permission. GitHub finds no licence file in the repository, and without a licence all rights are reserved by default: you may read the code but not reuse it. Check the README, or ask the authors, before using it.
- Is it still maintained?
- Activity is slowing. The repository last received commits 6 months ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 3, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What LTP does and who reaches for it
LTP is a Chinese natural language processing toolkit from the Harbin Institute of Technology Social Computing and Information Retrieval lab. The README describes it as a set of tools for segmenting, part-of-speech tagging and parsing Chinese text, and the accompanying N-LTP paper lists six fundamental tasks: Chinese word segmentation, part-of-speech tagging, named entity recognition, dependency parsing, semantic dependency parsing and semantic role labeling.
The audience is narrow and specific. If your input is Chinese and you need more than one analysis layer over the same sentence, LTP is built for that. The paper frames the design against toolkits such as Stanza that use an independent model per task, and says N-LTP instead uses a multi-task framework with a shared pre-trained model so that knowledge is shared across related Chinese tasks. That is the pitch: one pass, several annotations, one model family.
If your text is English or multilingual, this is the wrong shelf. Nothing in the README claims coverage outside Chinese, and the task list is defined around Chinese lexical analysis.
The two-model split inside LTP 4.2.0
Version 4.2.0 restructured the project into two parts, and the distinction matters more than any other detail on this page.
The legacy model is a perceptron-based algorithm rewritten in Rust. The release notes state that its accuracy is comparable to LTP 3 and that it is 3.55 times faster than LTP v3, with a 17.17 times speedup when multithreading is enabled. It supports only three tasks: segmentation, part-of-speech tagging and named entity recognition.
The deep learning model is the PyTorch implementation and supports all six tasks. It is the one you load when you ask for srl, dep, sdp or sdpg in a pipeline call.
So the choice is not a tuning knob. It is a capability boundary. A pipeline that requests semantic dependency parsing cannot be served by the legacy model, and the README's own commented example shows a related trap: requesting NER without POS fails, because NER needs the part-of-speech result. Tasks have dependencies on each other, and the pipeline enforces them.
The same release also moved decoding for segmentation and for Eisner dependency and semantic dependency parsing into Rust, and added training scripts and examples so users can train on private data. Training configuration for the deep model uses hydra.
Installing LTP and running a first pipeline
The README gives two install routes. Both pull three packages: ltp, ltp-core and ltp-extension. PyTorch and Transformers are installed first as dependencies. The first route uses the Tsinghua mirror explicitly on each command.
pip install -i https://pypi.tuna.tsinghua.edu.cn/simple torch transformers
pip install -i https://pypi.tuna.tsinghua.edu.cn/simple ltp ltp-core ltp-extensionThe second route sets the mirror globally, then installs the same packages without the -i flag.
pip config set global.index-url https://pypi.tuna.tsinghua.edu.cn/simple
pip install torch transformers
pip install ltp ltp-core ltp-extensionThe README notes that if you hit errors you should retry the LTP install and file a GitHub issue if it still fails. That is the whole troubleshooting guidance.
Once installed, the README's quickstart loads the small deep model by name and runs several tasks at once. The model is fetched from Hugging Face by default, and the README warns that a proxy may be needed.
import torch
from ltp import LTP
ltp = LTP("LTP/small")
if torch.cuda.is_available():
ltp.to("cuda")
output = ltp.pipeline(["他叫汤姆去拿外衣。"], tasks=["cws", "pos", "ner", "srl", "dep", "sdp", "sdpg"])
print(output.cws)
print(output.pos)
print(output.sdp)The result object supports attribute access, index access and dictionary-style access, so output.cws, output[0] and output['cws'] all reach the same field. Instead of a model name you can pass a local directory, provided it contains config.json and the other model files.
For the fast path, load the legacy model and expect a tuple. The README shows the working call and the failing one side by side.
ltp = LTP("LTP/legacy")
cws, pos, ner = ltp.pipeline(["他叫汤姆去拿外衣。"], tasks=["cws", "pos", "ner"]).to_tuple()
print(cws, pos, ner)The commented-out line in the README requests only cws and ner and is marked as an error, because NER requires the POS result. If your first pipeline call throws, check the task list before you check the model.
The README also documents custom vocabulary through add_word and add_words, which take a word and a frequency. That is the documented way to push domain terms into segmentation without retraining.
Where LTP stops being the right tool
The task split is the sharpest limitation. If you want speed and you also want dependency or semantic parsing, the legacy model cannot give you both. You either accept the deep model's cost or you drop those tasks. The release notes present the Rust rewrite as an answer to users' demand for inference speed, but that answer covers three tasks out of six.
Model acquisition is the second constraint. The README's own comment on the default load says Hugging Face download may need a proxy. In environments where that host is unreachable, the named-model shortcut fails, and you are pushed to the local-directory form, which means obtaining the files by some other route first. The README does not describe an offline bundle or a mirror for the models.
The README does not document rollback, and it does not publish accuracy numbers for the deep models. The speed figures belong to the legacy model, and they are quoted against LTP v3, not against the deep model in the same release. Do not read them as a comparison between LTP/small and LTP/legacy.
Finally, the API surface changed. The 4.2.0 notes label the switch to the Pipeline API a breaking change, so code written against earlier LTP 4 releases will not simply run. The release notes give no migration guide beyond pointing at the GitHub quickstart section.
LTP against Stanza and the per-task approach
The N-LTP paper names Stanza directly as the contrasting design. Stanza adopts an independent model for each task. LTP uses a shared pre-trained model across tasks, with knowledge distillation in which a single-task model teaches the multi-task model.
The practical difference is what you maintain. With independent models you can swap one task's model without touching the others, and a failure in one task stays local. With a shared backbone, one load serves several tasks and the annotations are produced together, which is the stated reason for the Pipeline API: the release notes mention that semantic dependency parsing and semantic dependency graph parsing overlap heavily and that reusing work speeds up inference.
That reuse is also coupling. A shared model means a shared dependency on the same pre-trained weights, the same download, and the same memory footprint, and it means task ordering constraints like the NER-needs-POS rule. If your work is one task only, say segmentation alone, the multi-task design buys you nothing and the legacy model is the more direct fit.
Maintenance, packaging and licence status
The repository is not archived, and the last push was on 2026-03-11. The most recent tagged release is v4.2.0 from 2022-08-15. That gap is worth stating plainly: the codebase has moved since the last tag, but the release notes stop at 4.2.0, so anything added after that date is not described in the README or the release notes.
The packaging is split across ecosystems. The Cargo workspace lists rust/ltp, rust/ltp-cffi and python/extension as members, so the Rust models, the C FFI layer and the Python extension are built from one workspace. The Makefile shows the build path: pip wheel for python/core and python/interface, maturin for the extension, and cbindgen targets that generate bindings/c/ltp.h and compile a C example against libltp. There is also a train_legacy target that runs the Rust examples for cws and pos with train, eval and predict subcommands over files under data/examples.
Upgrade cost concentrates in the API break at 4.2.0 and in the model download. The licence is not stated in the README or the file listing, so check the repository's licence file before you depend on it; this is not legal advice.
Editorial conclusion
Adopt LTP if you need several Chinese analysis layers over the same text and want to choose between a fast legacy model and a fuller deep model. Do not adopt it if you need languages other than Chinese, or if you need semantic role labeling, dependency parsing or semantic dependency parsing from the fast path, because the legacy model covers only segmentation, POS and NER. Before committing, verify which model directory your code loads, whether the Hugging Face download works from your network, and whether the 4.2.0 Pipeline API is the interface your existing code expects, since the release notes call the move away from the earlier API a breaking change.
Frequently asked questions
How do I install HIT-SCIR LTP?
Install PyTorch and Transformers first, then ltp, ltp-core and ltp-extension. The README offers a Tsinghua mirror route using the -i flag on each command and a global mirror route using pip config set global.index-url.
Which tasks does the LTP legacy model support?
The legacy model covers segmentation, part-of-speech tagging and named entity recognition only. The deep learning model is the one that supports all six tasks, including semantic role labeling, dependency parsing and semantic dependency parsing.
Why does my LTP pipeline call fail when I request only cws and ner?
NER depends on the part-of-speech result. The README's example marks tasks=["cws", "ner"] as an error and shows the working call with tasks=["cws", "pos", "ner"].
Does HIT-SCIR LTP need a proxy to download models?
The README's quickstart comment on LTP("LTP/small") says the default Hugging Face download may need a proxy. You can instead pass a local model directory that contains config.json and the other model files.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/hit-scir-ltp)