PERT: BERT pre-training by shuffling instead of masking
PERT: Pre-training BERT with Permuted Language Model
At a glance
- What is it?
- ymcui/PERT releases Chinese and English BERT models pre-trained with a permuted language model, where tokens are reordered rather than replaced with [MASK]. The repository ships weights and documentation only, and the README warns that PERT improves some NLU tasks and does worse on others.
- Who is it for?
- Use PERT if you run Chinese or English NLU at encoder scale and want to try a checkpoint whose pre-training avoided the [MASK] artifact, since it loads as a drop-in BERT replacement through hfl/chinese-pert-base or hfl/english-pert-base with no code changes. Use the MRC variants only if your task is extractive question answering in Chinese.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 164 days ago.
- What is it written in?
- GitHub does not report a main language for this repository.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Shuffling instead of masking
PERT is a pre-training method for BERT from the HFL lab, published on arXiv as 2203.06906 by Yiming Cui, Ziqing Yang and Ting Liu. The motivation in the README starts from an observation about reading: a certain amount of scrambled word order does not stop you understanding a sentence. The question the authors ask is whether semantic knowledge can be learned from that scrambled text.
BERT's masked language model picks some tokens, replaces them with a [MASK] symbol, and trains the model to recover the original token. PERT instead permutes the order of tokens in the input and trains the model to predict the original position of each token. No [MASK] symbol is introduced at all, because nothing was removed.
The README's worked example makes the difference concrete. Given a Chinese sentence about word order not affecting reading, BERT masks positions 7, 10 and 13 and learns to predict the tokens that belong there. PERT swaps tokens at positions 2 and 3 and at positions 13 and 14, then learns position 2 maps to 3, 3 maps to 2, 13 maps to 14 and 14 maps to 13.
Why this matters is the pretrain-finetune mismatch. Downstream tasks never contain [MASK], so a model trained to expect it spends capacity on an artifact. Removing the symbol removes the artifact, at the cost of a harder training objective.
What is released
Six models are published, in Chinese and English, at two sizes. The README gives the architectures: PERT-large is 24 layers, 1024 hidden, 16 heads and 330M parameters, and PERT-base is 12 layers, 768 hidden, 12 heads and 110M parameters. Those are the standard BERT-large and BERT-base shapes.
The Chinese models are trained on the EXT corpus, which the README describes as Chinese Wikipedia plus other encyclopedias, news and question answering data, totalling 5.4B tokens and roughly 20G on disk, the same corpus used for MacBERT. The English models are trained on Wikipedia plus BookCorpus, matching the original BERT recipe, and are uncased.
Two extra models are fine-tuned rather than pre-trained. Chinese-PERT-base-MRC and Chinese-PERT-large-MRC were released on 2022-05-07 after fine-tuning on several reading comprehension datasets, with an interactive Hugging Face demo. Those are the ones to reach for if your task is extractive question answering in Chinese.
One detail worth knowing: the README states that the config and vocab files match Google's original BERT-base Chinese exactly, and the English ones match BERT-uncased. That is what makes drop-in loading work.
Loading PERT with transformers
Because the body of the model is still a BERT architecture, PERT loads through the standard BERT classes. The README is explicit that every model in this directory is loaded with BertTokenizer and BertModel, and that the MRC models use BertForQuestionAnswering.
from transformers import BertTokenizer, BertModel
tokenizer = BertTokenizer.from_pretrained("MODEL_NAME")
model = BertModel.from_pretrained("MODEL_NAME")The MODEL_NAME values given in the README are hfl/chinese-pert-large, hfl/chinese-pert-base, hfl/chinese-pert-large-mrc, hfl/chinese-pert-base-mrc, hfl/english-pert-large and hfl/english-pert-base. The base models are listed at 0.4G and the large models at 1.2G.
That is the entire integration. If your pipeline already loads a BERT checkpoint from Hugging Face, swapping in PERT is a string change, and any code written against BertModel keeps working. There is no custom class to import and no new tokeniser to learn.
The flip side is that this compatibility is structural rather than a feature. PERT is BERT with a different pre-training objective, so anything you were doing with BERT you can do here, and anything BERT cannot do, PERT will not do either.
Where the weights live
There are two distribution channels, and they are not equivalent.
The original release is TensorFlow 1.15 checkpoints hosted on Baidu Pan, with the extraction passwords printed in the README next to each link. Extracting the Chinese-PERT-base archive yields pert_model.ckpt, pert_model.meta, pert_model.index, pert_config.json and vocab.txt. TensorFlow 1.15 is a long-retired version, and a Chinese file-hosting service with a manual password is an awkward dependency for a team outside that ecosystem.
The practical route is the second channel. PyTorch and TensorFlow 2 versions are published on Hugging Face under the hfl organisation, which is what the loading example above uses. The README describes downloading through the Files and versions tab on each model page.
Anyone who wants to inspect the original checkpoint files or verify they match will need the Baidu links, and should expect the friction that comes with them.
The README admits the results are mixed
This is the part to read before adopting the model. The README states plainly that PERT achieves performance improvements on some Chinese and English NLU tasks, and that it performs worse on others, and asks users to use it as appropriate.
That is an unusually direct admission, and it is consistent with the design. A permuted objective is a different inductive bias, not a strictly better one. A task that depends heavily on exact word order, such as some forms of sequence labelling, is plausibly the loser when the pre-training signal spent its capacity on recovering order.
The README's baseline section lists results on ten Chinese tasks, including extractive reading comprehension on CMRC 2018 for Simplified Chinese and DRCD for Traditional Chinese, with the note that values outside parentheses are maximum and values inside are average across runs. It says only part of the results are listed and refers to the paper for full detail and analysis.
Since the paper is on arXiv and the plots in it are, per the README, wrong in at least one figure, with the README's own diagram flagged as authoritative until the paper is updated, treat the repository as the reference for how the method works and the paper as the reference for the numbers.
A 2022 encoder in a 2026 stack
The timeline in the README is short and tells its own story. Models were announced on 2022-02-17, released on 2022-02-24, the technical report appeared on 2022-03-15, the MRC models on 2022-05-07, and the last news entry is from 2023-03-28 pointing at the lab's Chinese LLaMA and Alpaca work.
So PERT is a 2022 encoder, and the lab's attention moved to decoder models in 2023. The last push to the repository was on 2026-04-19, but no new model has been announced in the news section since 2022.
That is not a defect, it is a position. A 110M or 330M encoder is still the right tool for classification, sequence labelling and extractive question answering at high throughput, and it runs on a single GPU. It is the wrong tool for generation, instruction following or anything that needs world knowledge, and no amount of fine-tuning changes that.
Choose it as a BERT replacement in an existing pipeline, not as a step into current model practice.
MacBERT and whole word masking as alternatives
The README links to the same lab's other Chinese models, and the closest comparison is MacBERT, since both were built to reduce the pretrain-finetune gap and both use the same EXT corpus.
MacBERT's approach is called MLM as correction: it still masks, but replaces the masked token with a similar real word rather than a [MASK] symbol, so the input stays natural while the objective stays close to BERT's. PERT removes the masking step entirely and instead permutes token order, changing the objective as well as the input.
Chinese-BERT-wwm is the more conservative option, applying whole word masking so that all subword pieces of a Chinese word are masked together, which matters for a language without spaces. It keeps [MASK] and is the safest baseline.
In practice these are three answers to the same question, and the README's own warning applies to all of them: test on your task rather than picking by reputation. If you want the smallest departure from standard BERT, take whole word masking. If you want masking without the artifact symbol, take MacBERT. If you want no masking at all, PERT is the experiment.
Licence, code and upkeep
PERT is Apache-2.0, with the LICENSE file at the root. That is permissive and allows commercial use of the weights, subject to whatever terms apply to the corpora they were trained on, which the README does not discuss.
The repository contains no code. The root listing is README.md, README_EN.md, LICENSE, .github/ and pics/. There are no training scripts, no data preparation scripts and no fine-tuning examples, and there are no releases or tags.
So this is a model release, not a software project. You can use the weights and you cannot reproduce the pre-training from this repository, and there is nothing here to upgrade or maintain. If you need the training procedure, the arXiv paper is the only source.
That shapes the maintenance question completely. There is no dependency to track, no upstream to follow, and no expectation of fixes. The risk is not that PERT breaks, it is that a 2022 checkpoint quietly becomes the oldest thing in your stack while nobody revisits the decision.
Editorial conclusion
Use PERT if you run Chinese or English NLU at encoder scale and want to try a checkpoint whose pre-training avoided the [MASK] artifact, since it loads as a drop-in BERT replacement through hfl/chinese-pert-base or hfl/english-pert-base with no code changes. Use the MRC variants only if your task is extractive question answering in Chinese. Do not use it if you need generation or instruction following, or if you expect the best result on every task, because the README states PERT is worse than BERT on some tasks. Do not come here looking for training code, since the repository ships weights and documentation only. Verify on your own dataset before switching, and compare against MacBERT or whole word masking on the same split rather than trusting the published table.
Frequently asked questions
How does PERT differ from BERT?
BERT replaces tokens with a [MASK] symbol and predicts the original token. PERT permutes token order instead and predicts each token's original position, so no [MASK] symbol is introduced during pre-training.
How do I load PERT in Python?
The README says to use BertTokenizer and BertModel from the transformers library, with MODEL_NAME values such as hfl/chinese-pert-base or hfl/english-pert-large. The MRC models use BertForQuestionAnswering instead.
Does PERT include training code?
No. The repository contains only README files, a LICENSE and images, with no training or fine-tuning scripts and no releases, so the weights can be used but the pre-training cannot be reproduced from here.
Is PERT better than BERT on every task?
No. The README states PERT improves some Chinese and English NLU tasks and performs worse on others, and asks users to choose accordingly.
What model sizes are available for PERT?
PERT-large is 24 layers, 1024 hidden, 16 heads and 330M parameters, and PERT-base is 12 layers, 768 hidden, 12 heads and 110M parameters, in both Chinese and English plus Chinese MRC variants.
Where are the PERT weights hosted?
PyTorch and TensorFlow 2 versions are on Hugging Face under the hfl organisation, while the original TensorFlow 1.15 checkpoints are on Baidu Pan with passwords given in the README.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/ymcui-pert)