# Chinese-BERT-wwm: Whole Word Masking Pre-Trained Models for Chinese NLP

> Chinese-BERT-wwm is a family of pre-trained Chinese language models from Harbin Institute of Technology and iFLYTEK Research that applies whole word masking to Chinese BERT, publishing eight variants covering base, large, and compact sizes. It corrects a specific flaw in Google's original Chinese BERT where character-level tokenization allowed masked characters to be trivially inferred from their neighbors within the same word.

**ymcui/Chinese-BERT-wwm** — Pre-Training with Whole Word Masking for Chinese BERT（中文BERT-wwm系列模型）

- Repository: https://github.com/ymcui/Chinese-BERT-wwm
- Website: https://ieeexplore.ieee.org/document/9599397
- Stars: 10,232 · Forks: 1,379
- Language: Python
- License: Apache-2.0
- Published: 2026-09-16 · Updated: 2026-09-16 · Language: en
- Canonical page: https://hysenlabs.com/projects/ymcui-chinese-bert-wwm

## What Whole Word Masking Fixes in Chinese BERT Pre-Training

Google's original Chinese BERT tokenizes text at the character level. Chinese has no natural whitespace between words, so the original model treats each character as an independent token. During pre-training, BERT randomly selects 15% of tokens for masking. In character-level Chinese, adjacent characters that form a single word can end up with some masked and others visible. For example, in the two-character word meaning "language," masking only the first character while leaving the second visible gives the model a strong contextual hint that effectively defeats the purpose of the mask. The model learns to predict characters from adjacent characters rather than from broader sentence semantics.

Whole Word Masking addresses this. The README describes the mechanism: before constructing training samples, Chinese text is segmented into complete words using the LTP (Language Technology Platform) toolkit from Harbin Institute of Technology. Once segmentation identifies a word boundary, masking applies to all characters in that word together. If any character of a word is selected for masking, the rest of the word's characters are also masked in the same training sample.

The README provides a concrete example with the sentence containing the phrase meaning "use a language model to predict the probability of the next word." Under character-level masking, a word like "model" (two characters) might have one character masked while the other remains visible. Under whole word masking, both characters are masked simultaneously, requiring the model to reconstruct the full word from the surrounding sentence. The README also clarifies that "masking" here covers all three BERT masking strategies: replacement with the [MASK] token, retaining the original character, and randomly substituting a different character.

## Eight Model Variants: Base, Large, and the Compact RBT Series

The repository releases eight model variants. They divide into three groups based on architecture size and training corpus.

Base models (12 transformer layers, 768 hidden units, 12 attention heads, approximately 110 million parameters): BERT-wwm trains on Chinese Wikipedia only. BERT-wwm-ext and RoBERTa-wwm-ext train on the larger extended corpus. The RoBERTa variant uses the RoBERTa pre-training approach (longer training, dynamic masking) applied to the same Chinese data.

Large model (24 transformer layers, 1024 hidden units, 16 attention heads, approximately 330 million parameters): RoBERTa-wwm-ext-large. This is the highest-capacity model in the series.

Compact RBT series (all using the extended corpus): RBT3, RBT4, and RBT6 reduce depth to 3, 4, and 6 layers respectively, each keeping the base model width. RBTL3 uses the large model's width (1024 hidden units) with only 3 layers. These are intended for deployments where the full base model latency is prohibitive.

The naming convention is consistent: "ext" marks the larger training corpus, "large" marks the larger architecture, and the RBT prefix with a number marks the compact layer-reduced variants. For new projects, the README consistently points toward the ext variants as the recommended starting point over models trained only on Chinese Wikipedia.

## How to Download the Models and Access the Repository

All eight models are available from the HuggingFace model hub under the `hfl` organization. The specific model identifiers from the download table include `hfl/chinese-bert-wwm`, `hfl/chinese-bert-wwm-ext`, `hfl/chinese-roberta-wwm-ext`, `hfl/chinese-roberta-wwm-ext-large`, `hfl/rbt3`, `hfl/rbt4`, `hfl/rbt6`, and `hfl/rbtl3`. The README recommends Baidu Netdisk for users in mainland China and HuggingFace for users outside. Base model archives are approximately 400 MB each.

To download PyTorch weights directly, the README instructs users to navigate to the model's HuggingFace page, select the "Files and versions" tab, and download the individual checkpoint files. TensorFlow weights come as zip archives. All models support TensorFlow 2 through the Transformers library (integrated December 2019) and PaddleHub through PaddleNLP (March 2020).

To clone the repository itself and access the comparison data files in the `data/` directory:

```bash
git clone https://github.com/ymcui/Chinese-BERT-wwm
```

The repository's `data/` directory contains input and output examples for downstream task evaluation. The actual model weights are not stored in the repository and must be downloaded separately from HuggingFace or Baidu Netdisk.

## Training Corpus and Architecture Specifications

The extended corpus used in BERT-wwm-ext and all subsequent ext variants totals 5.4 billion words. It combines Chinese Wikipedia (both simplified and traditional characters) with encyclopedia entries, news articles, and question-and-answer data. The original BERT-wwm uses only Chinese Wikipedia, which is smaller. This corpus difference is the primary reason the ext variants are generally preferred.

The base architecture follows the standard BERT-base configuration precisely: 12 transformer layers, 768 hidden units, 12 attention heads, and roughly 110 million parameters. The large architecture doubles depth to 24 layers and widens the hidden dimension to 1024 units with 16 heads, reaching approximately 330 million parameters.

The RBT compact models reduce transformer layer count while preserving other hyperparameters. RBT3 has 3 layers and base-model width. RBTL3 also has 3 layers but large-model width (1024 hidden units), creating a parameter-asymmetric option for cases where high representation capacity per layer matters more than depth. The README does not quantify latency improvements of compact models over the full base or large variants, so practitioners should benchmark their own serving infrastructure.

## A Hard Limit: Open-Source Weights Do Not Include the MLM Head

The README states this explicitly in the download section. The published model weights do not include the masked language modeling (MLM) output layer, which maps hidden states to vocabulary-size logits. Any use case that requires predicting masked tokens directly is affected: language model scoring, perplexity evaluation, domain-adaptive pre-training that starts from the HFL checkpoint, and data augmentation workflows that generate plausible token substitutions.

For downstream fine-tuning tasks such as named entity recognition, text classification, reading comprehension in SQuAD format, and natural language inference, the MLM head is not needed. Those tasks attach their own output layers to the encoder's hidden states and train on labeled data. The absent MLM head is not an obstacle for these use cases, which represent the majority of applied Chinese NLP work.

Anyone who needs to run secondary pre-training on a domain corpus (medical, legal, financial) and wants to continue from the HFL checkpoint must either add a fresh randomly-initialized MLM head or obtain the task-complete weights through other means. The README advises that secondary pre-training proceeds "like any other downstream task," but this requires code to instantiate the MLM head, which the repository does not provide as a ready-to-run script in the visible portion of the codebase.

## Google BERT-base-chinese: When the Standard Alternative Is Sufficient

The direct alternative is Google's `bert-base-chinese` on HuggingFace, which uses the same 12-layer base architecture but pre-trains with character-level masking on Chinese Wikipedia. Its main advantage for practitioners is familiarity: most published Chinese NLP benchmark results and tutorials report numbers against `bert-base-chinese`, which simplifies comparison and code reuse. Teams with existing pipelines built for `bert-base-chinese` can switch to an HFL model with a one-line identifier change, but they must re-run all downstream evaluation to confirm that results hold.

The practical case for the HFL models is strongest when the task involves word-level understanding. Named entity recognition, where multi-character entities must be identified as units, and reading comprehension, where answer spans often align with word boundaries, are the cases where whole word masking most directly affects pre-training signal quality. For sentence-pair tasks like natural language inference where character-level co-occurrence statistics are often sufficient, the difference between the two masking strategies is smaller.

The research is peer-reviewed and published as "Pre-Training with Whole Word Masking for Chinese BERT" in IEEE/ACM Transactions on Audio, Speech, and Language Processing (document identifier 9599397), authored by Yiming Cui, Wanxiang Che, Ting Liu, Bing Qin, and Ziqing Yang from HFL.

## Licence, Publication Record, and Repository Status

All models and associated code are released under the Apache-2.0 licence, which permits commercial use, redistribution, and modification with attribution. No additional patent or royalty requirements apply beyond standard Apache-2.0 terms. This makes the models usable in production applications without per-deployment licensing concerns.

The HFL (Harbin Institute of Technology and iFLYTEK Research Joint Laboratory) has published multiple related models in the same repository namespace, including MacBERT, Chinese-ELECTRA, Chinese-XLNet, and LERT. The Chinese-BERT-wwm work preceded most of these and represents the foundational whole-word-masking contribution from that lab. The paper citation is stable and peer-reviewed, which matters for teams who need to cite their pre-training basis in technical reports or regulatory filings.

The repository is not archived. The last recorded push was on 2026-04-19.

## Conclusion

Teams working on Chinese NLP for classification, named entity recognition, or reading comprehension should start with BERT-wwm-ext rather than BERT-wwm: the larger extended corpus covering 5.4 billion words offers a stronger starting point. Anyone who needs masked language modeling weights must perform secondary pre-training on their own data, since the README states explicitly that the MLM task head is absent from the open-source release. Before fine-tuning, confirm that the target framework matches the weight format you download, and check the HuggingFace model page for `hfl/chinese-bert-wwm-ext` to verify the available checkpoint files.

## FAQ

### What is the difference between Chinese-BERT-wwm and BERT-wwm-ext?

BERT-wwm trains only on Chinese Wikipedia, while BERT-wwm-ext uses an extended corpus combining Wikipedia with encyclopedia, news, and question-and-answer data totaling 5.4 billion words. The README consistently points to BERT-wwm-ext as the stronger starting point for most downstream tasks.

### Do Chinese-BERT-wwm models include weights for the masked language modeling task?

No. The README explicitly states that the open-source release does not include MLM task weights. Users who need to run language model scoring, evaluate perplexity, or perform domain-adaptive secondary pre-training must add a fresh MLM head and train it on their own data.

### Which frameworks support loading Chinese-BERT-wwm models?

All models support TensorFlow 2 and PyTorch through the Hugging Face Transformers library, integrated since December 2019. PaddleHub integration via PaddleNLP has been available since March 2020. PyTorch weights are downloadable directly from the HuggingFace model pages under the hfl organization.

## Sources

- [Issues](https://github.com/ymcui/Chinese-BERT-wwm/issues)
- [License: Apache-2.0](https://github.com/ymcui/Chinese-BERT-wwm/blob/master/LICENSE)
- [Project website](https://ieeexplore.ieee.org/document/9599397)
- [README](https://github.com/ymcui/Chinese-BERT-wwm/blob/master/README.md)
- [ymcui/Chinese-BERT-wwm on GitHub](https://github.com/ymcui/Chinese-BERT-wwm)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/ymcui-chinese-bert-wwm
