pycorrector: a Python toolkit for Chinese spelling and grammar correction
pycorrector is a toolkit for text error correction. 文本纠错,实现了Kenlm,T5,MacBERT,ChatGLM3,Qwen2.5等模型应用在纠错场景,开箱即用。
At a glance
- What is it?
- pycorrector bundles Kenlm, MacBERT, T5, ERNIE, ChatGLM and Qwen correction models behind one pip package. It is aimed at Chinese text, and the model you pick decides both the error types you can fix and the hardware you need.
- Who is it for?
- Adopt pycorrector if your input is Chinese text and you need spelling, phonetic, shape-similar or grammar correction with an off-the-shelf model you can swap. Do not adopt it if you need a general-purpose spell checker for English, or if you cannot run PyTorch or PaddlePaddle models and would be limited to the CPU-only Kenlm path, whose SIGHAN-2015 F1 in the project's own table is 0.3147.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 67 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Which Chinese errors pycorrector is built to catch
The README splits Chinese text errors into phonetic similarity, shape similarity, grammar and proper-noun errors, and says the project focuses on those four. That framing matters because the corrector is not a general spell checker. A pinyin input method produces errors that sound right and look wrong, an OCR pipeline produces errors that look right and sound wrong, and a search query can contain either. The README makes exactly this point: pinyin input and speech recognition review care about phonetic errors, Wubi input and OCR review care about shape errors, and search query correction cares about all of them.
The intended user is a Python developer who already has Chinese text and wants a correction layer without training a model from scratch. The repository ships example directories for each model family (examples/kenlm, examples/macbert, examples/t5, examples/ernie_csc, examples/gpt, examples/seq2seq, examples/deepcontext, examples/mucgec_bart), so the practical choice is which directory to open, not which architecture to design. The project description calls this out-of-the-box usage, and the examples layout supports that claim.
One interface, several very different correction mechanisms
pycorrector is a collection rather than a single model. The Kenlm path trains an NGram language model over Chinese and combines it with rules and a confusion set; the README describes it as fast, extensible and average in quality. The DeepContext path is a PyTorch reimplementation of a Stanford NLC-style model. The Seq2Seq path is a ConvSeq2Seq model that placed third in the NLPCC-2018 Chinese grammatical error correction shared task using a single model. T5, ERNIE_CSC and MacBERT4CSC are fine-tuned encoder or encoder-decoder models, with MacBERT marked as recommended in the feature list because it adds error detection and correction networks suited to Chinese spelling correction. The GPT path fine-tunes ChatGLM and LLaMA variants on Chinese CSC and grammatical error data.
The distinction that actually changes your integration is length alignment. The evaluation table separates CSC models, which handle phonetic, shape and grammatical errors where the corrected text has the same length as the input, from CTC models, which also handle extra or missing characters. If your users drop or duplicate a character, a CSC-only model cannot represent the fix. The Qwen2.5 and Qwen3-based correctors added in v1.1.0 and v1.1.2 are described as supporting extra characters, missing characters, wrong characters, word order and grammar.
Installing pycorrector and running a first correction
The package is on PyPI and the Dockerfile installs it with pip3, so a plain pip install is the documented path. The requirements file lists jieba, pypinyin, transformers, datasets, numpy, pandas, six, loguru and pyahocorasick; setup.py declares Python 3.6 or newer while the README badge says 3.8 or newer, so treat 3.8 as the practical floor.
pip install pycorrectorFor a container, the repository's Dockerfile shows the sequence it uses: a Python 3.8 base image, a CPU-only PyTorch wheel, then the requirements, then pycorrector itself.
FROM nikolaik/python-nodejs:python3.8-nodejs20
WORKDIR /app
COPY . .
RUN pip3 install torch --index-url https://download.pytorch.org/whl/cpu --no-cache-dir
RUN pip3 install -r requirements.txt -i https://pypi.tuna.tsinghua.edu.cn/simple
RUN pip3 install pycorrector -i https://pypi.tuna.tsinghua.edu.cn/simpleThe README points at examples/macbert/gradio_demo.py as a runnable demo, started with a single command. That is the fastest way to see a correction end to end before wiring anything into your own code.
python examples/macbert/gradio_demo.pyExpect the first run to download model weights from HuggingFace, since the feature list links each model to a shibing624 or third-party repository. The ERNIE_CSC example is the exception in dependency terms: it is implemented on PaddlePaddle, not PyTorch.
Reading the evaluation table before you pick a model
The repository publishes an evaluation script at examples/evaluate_models/evaluate_models.py and reports F1 at strict sentence level, meaning a corrected sentence counts as right only if it matches the reference exactly. Three test sets are used: SIGHAN-2015, EC-LAW and MCSC. The GPU listed is a Tesla V100 with 32 GB of memory.
The per-dataset spread is the part worth studying. Mengzi-T5-CSC scores 0.7758 on SIGHAN-2015 but 0.1039 on MCSC, and ERNIE-CSC scores 0.8383 on SIGHAN-2015 against 0.1318 on MCSC. A high average can therefore hide a model that collapses on your domain. EC-LAW sits between the two for both models, around 0.32 to 0.34. Kenlm-CSC is the only CPU entry, at 9 QPS, against 214 QPS for Mengzi-T5-CSC and 114 for ERNIE-CSC on GPU.
Sentence-level strict matching is also a harsh metric. If a model fixes the target error but rewrites one unrelated character, the sentence is scored wrong. A model with a modest F1 here may still be useful in a suggestion UI where a human confirms each change.
Where pycorrector is the wrong tool
The project is Chinese-specific. The classifier in setup.py lists Simplified and Traditional Chinese as the natural languages, and every model in the feature list is trained on Chinese correction data. If your text is English, this is not the library you want, and the README offers no English path.
The second limitation is that model quality varies by dataset far more than a single average suggests. A team that reads only the Avg column and deploys a model without testing on its own corpus is taking a real risk, and the repository gives you the script to avoid that. Third, the split between CSC and CTC models is a hard constraint, not a tuning knob: a length-aligned model cannot delete a duplicated character or insert a missing one.
Finally, the heavier paths carry infrastructure cost. The GPU entries in the evaluation table assume a 32 GB V100, and the GPT-family correctors are fine-tuned 6B and larger models. If you cannot serve that, the realistic options narrow to Kenlm on CPU or a smaller fine-tuned encoder, with the accuracy that implies.
How pycorrector differs from a general-purpose spell checker
A general-purpose spell checker such as a Hunspell-style dictionary checker works on token boundaries and edit distance against a word list. It has no notion of pinyin, so it cannot know that a wrong character was typed because it sounds like the right one, and it cannot score a whole sentence for fluency. pycorrector's Kenlm path is the closest thing to that class of tool, but it scores character sequences with an NGram language model plus a confusion set rather than checking words against a dictionary.
The larger difference is the model-based paths. MacBERT4CSC, T5, ERNIE_CSC and the Qwen and ChatGLM correctors are fine-tuned neural models that read the whole sentence and can propose substitutions a dictionary would never list, including grammar and word-order fixes. That is also why they need PyTorch or PaddlePaddle and, for the GPT family, a GPU. If your requirement is a fast, dependency-light checker on Latin script, a dictionary checker is the better fit and pycorrector is not a substitute.
Maintenance, releases and the Apache-2.0 licence
The repository is not archived, and the last push was on 2026-07-25, the same day as the 1.1.4 release. The release notes for 1.1.4 describe a change to ProperCorrector: replacing a per-segment full scan of the proper-noun dictionary with a multi-level inverted index keyed on word length and position, removing a redundant Trie, and fixing OOV stroke false positives. The stated result is roughly an 80x speedup on a 40,000-entry proper-noun dictionary. That is a performance claim from the release notes, not an independent measurement, and the same notes are the only place it appears.
Upgrade cost is mostly model-side. Releases 1.1.0, 1.1.2 and 1.1.4 added Qwen2.5, Qwen3 and ProperCorrector work respectively, so moving between minor versions can mean downloading new weights rather than only replacing a wheel. The package pins nothing tightly: requirements.txt lists transformers without a version bound, which means a breaking transformers release can reach you on a fresh install. Pinning transformers in your own lockfile is the practical mitigation.
On licensing, the project is Apache-2.0, which permits commercial use and modification with attribution and notice requirements. That covers the pycorrector code. The fine-tuned model weights are hosted separately on HuggingFace and ModelScope under their own terms, and the ERNIE_CSC example depends on PaddlePaddle, so check each model card and each upstream framework licence before shipping. This is a description of what the repository states, not legal advice.
Editorial conclusion
Adopt pycorrector if your input is Chinese text and you need spelling, phonetic, shape-similar or grammar correction with an off-the-shelf model you can swap. Do not adopt it if you need a general-purpose spell checker for English, or if you cannot run PyTorch or PaddlePaddle models and would be limited to the CPU-only Kenlm path, whose SIGHAN-2015 F1 in the project's own table is 0.3147. Before committing, run examples/evaluate_models/evaluate_models.py against your own labelled data, because the published averages hide large per-dataset swings.
Frequently asked questions
Is there a spell checker available for Python?
pycorrector is one, but it is a Chinese text corrector rather than a general English spell checker. The README lists Kenlm, MacBERT, T5, ERNIE, ChatGLM and Qwen-based models, all applied to Chinese spelling, phonetic, shape-similar and grammar errors.
Is there a spell checker available for Python for Chinese text?
pycorrector's feature list covers Kenlm, DeepContext, ConvSeq2Seq, T5, ERNIE_CSC, MacBERT4CSC, MuCGECBart, NaSGECBart and ChatGLM or LLaMA correctors, and the README says the project focuses on phonetic, shape, grammar and proper-noun errors in Chinese.
Is there a Python spell checker that handles phonetic and shape-similar errors?
The README states that pinyin input and speech recognition review care about phonetic errors while Wubi input and OCR review care about shape errors, and that pycorrector focuses on phonetic, shape, grammar and proper-noun error types. The evaluation table separates CSC models, which handle length-aligned phonetic, shape and grammar errors, from CTC models that also handle extra or missing characters.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/shibing624-pycorrector)