# MacBERT: Chinese BERT Pre-trained with MLM as Correction

> MacBERT is a Chinese pre-trained language model that modifies BERT's masked language modelling objective by substituting similar words for [MASK] tokens rather than masking them, reducing the mismatch between pre-training and fine-tuning. The model is available on HuggingFace in base and large sizes and loads as a drop-in replacement for standard BERT.

**ymcui/MacBERT** — Revisiting Pre-trained Models for Chinese Natural Language Processing (MacBERT)

- Repository: https://github.com/ymcui/MacBERT
- Website: https://www.aclweb.org/anthology/2020.findings-emnlp.58/
- Stars: 722 · Forks: 61
- Language: Unknown
- License: Apache-2.0
- Published: 2026-09-17 · Updated: 2026-09-17 · Language: en
- Canonical page: https://hysenlabs.com/projects/ymcui-macbert

## The Pre-training Mismatch MacBERT Was Built to Address

Standard masked language modelling introduces a [MASK] token at pre-training time, but that token never appears during fine-tuning or inference. The gap between the two phases means the model sees a distribution of inputs at pre-training that does not match what it will see later. MacBERT, published in Findings of EMNLP 2020 by researchers from the Harbin Institute of Technology and iFLYTEK Research, addresses this by replacing the [MASK] token with a similar word instead of a special masking symbol.

The target audience is teams building Chinese NLP applications on top of a pre-trained encoder: sentence classification, extractive reading comprehension, natural language inference, or sentence-pair matching. MacBERT's architecture is identical to BERT, so any fine-tuning code that loads a BERT model works with MacBERT without modification. The README states this directly: the model can be used as a drop-in replacement for BERT without changing existing code.

## MLM as Correction: Similar Words Instead of [MASK]

In MacBERT's pre-training setup, tokens selected for masking receive a similar word drawn from the Synonyms toolkit (Wang and Hu, 2017), which uses word2vec similarity to find near-synonyms. When no similar word can be found, a random word is substituted instead. The [MASK] token itself never appears. This is what the paper calls MLM as correction, or Mac.

MacBERT combines this substitution with two techniques from prior work: Whole Word Masking (wwm) and N-gram masking. With wwm, entire words (after Chinese word segmentation) are masked together rather than individual characters. With N-gram masking, consecutive words are masked as a unit. When an N-gram is selected, each constituent word gets its own similar-word replacement independently. The README shows a worked example comparing the four masking approaches on an English sentence, illustrating how the correction approach produces outputs that look like natural text with plausible substitutions rather than sentences full of special tokens.

A practical consequence is that the model's input distribution during pre-training stays closer to natural Chinese text. The trade-off is that the training target becomes harder: the model must recover the original token from a sentence that already looks grammatically plausible, rather than filling in an obvious mask.

## Loading MacBERT with HuggingFace Transformers

MacBERT is distributed through HuggingFace under the hfl organisation. The base model identifier is hfl/chinese-macbert-base and the large model is hfl/chinese-macbert-large. The README explicitly instructs users to load the model with BertTokenizer and BertModel, not with AutoTokenizer or AutoModel. The quick-load example from the README uses MODEL_NAME as a placeholder:

```python
tokenizer = BertTokenizer.from_pretrained("MODEL_NAME")
model = BertModel.from_pretrained("MODEL_NAME")
```

Replace MODEL_NAME with hfl/chinese-macbert-base for the base model or hfl/chinese-macbert-large for the large variant, as the README's model name table specifies. The HuggingFace download for the base model is 383 MB; the large model is 1.2 GB.

For teams without reliable access to HuggingFace, the README also lists Baidu Pan download links for both models in TensorFlow format. The base model uses password 61ga and the large model uses password zejf, as listed in the README's download table.

Because MacBERT shares BERT's architecture, it integrates with any HuggingFace pipeline, Trainer, or fine-tuning script that accepts a BertModel. No code changes beyond the model identifier are required.

## Two Model Sizes and Six Evaluation Benchmarks

MacBERT-base has 12 transformer layers, a hidden size of 768, 12 attention heads, and 102 million parameters. MacBERT-large has 24 layers, a hidden size of 1024, 16 attention heads, and 324 million parameters.

The README reports results on six Chinese NLP benchmarks. CMRC 2018 is a Simplified Chinese extractive reading comprehension dataset from Harbin Institute of Technology and iFLYTEK. DRCD is a Traditional Chinese reading comprehension dataset from Delta Research Center. XNLI covers natural language inference. ChnSentiCorp measures sentiment classification. LCQMC and BQ Corpus both test sentence-pair matching. On CMRC 2018 test, the README reports MacBERT-base reaching 73.2 EM and 89.5 F1 (average over 10 runs: 72.4 / 89.2), and MacBERT-large reaching 74.8 EM and 90.7 F1 (average: 73.2 / 90.1).

The README gives results as both the maximum over 10 independent runs and the average, which is a more stable indicator. Reporting only the maximum of repeated runs can overstate performance on small test sets, and the gap between maximum and average in these tables is sometimes several points.

## What MacBERT Does Not Include

The repository contains only the README, an English README, and image assets. There is no fine-tuning code, no pre-training script, and no example training loop. Teams that need working training or fine-tuning code must source it from the HuggingFace Transformers documentation or other community resources.

The repository has no GitHub releases. Updates are tracked through commits, and the most recent news entry in the README dates to 2023.

The BertTokenizer constraint is a real integration note. Using AutoModel or AutoTokenizer, which are common in HuggingFace tutorials, may cause incorrect loading behaviour with MacBERT. The README marks this as a notice in the quick-load section.

MacBERT was trained on Simplified Chinese corpora. For Traditional Chinese tasks, the README's DRCD benchmark shows results, but notes that models like ERNIE explicitly remove Traditional Chinese characters and should not be used for Traditional Chinese data. MacBERT does not have this restriction, but the degree to which it handles Traditional Chinese in practice depends on the pre-training corpus composition, which the README does not detail in full.

## MacBERT Against Chinese RoBERTa-wwm-ext

The HFL lab, which developed MacBERT, also released Chinese RoBERTa-wwm-ext (hfl/chinese-roberta-wwm-ext). Both models are available on HuggingFace and load the same way. The difference is in the pre-training objective: RoBERTa-wwm-ext uses standard masked language modelling with whole word masking and an extended pre-training dataset, while MacBERT adds the similar-word substitution on top.

For teams choosing between the two, the README's benchmark tables show that MacBERT-base and RoBERTa-wwm-ext perform similarly on CMRC 2018 test (MacBERT-base: 73.2 EM, RoBERTa-wwm-ext: 72.6 EM). The differences are small and task-dependent; the right choice depends on which benchmark is closest to the target task. RoBERTa-wwm-ext has a longer history of community use for Chinese BERT fine-tuning, which may mean more available fine-tuning examples. MacBERT's advantage is theoretically cleaner pre-training inputs at a modest overhead in vocabulary lookup complexity during training.

## Maintenance Status and Apache-2.0 License

The last push to the MacBERT repository was on 2026-04-19. The repository is not archived. The README's news section covers announcements through 2023, with no entries since then. The repository has no GitHub releases.

The Apache-2.0 license permits commercial use, modification, and redistribution with attribution. The model weights themselves are distributed through HuggingFace under the same license, so teams building products on top of MacBERT can use the weights in commercial applications with attribution to the original authors. The paper citation format is included in the README's citation section for teams that need to reference the work academically.

## Conclusion

MacBERT suits teams fine-tuning a pre-trained model for Chinese reading comprehension, sentiment classification, natural language inference, or sentence-pair matching, particularly when they want to load a model through HuggingFace without modifying existing BERT-based code. It does not suit teams that need training code, a custom pre-training pipeline, or a model with Traditional Chinese character exclusions stripped out; verify that your task is covered by the six evaluation benchmarks in the README before committing. The last push to the repository was on 2026-04-19, under the Apache-2.0 license.

## FAQ

### How do I load MacBERT in Python using HuggingFace Transformers?

The README instructs loading with BertTokenizer and BertModel, not AutoTokenizer or AutoModel. Use BertTokenizer.from_pretrained("hfl/chinese-macbert-base") and BertModel.from_pretrained("hfl/chinese-macbert-base"), substituting hfl/chinese-macbert-large for the larger model.

### What is the difference between MacBERT-base and MacBERT-large?

MacBERT-base has 12 layers, 768 hidden dimensions, and 102 million parameters (383 MB). MacBERT-large has 24 layers, 1024 hidden dimensions, and 324 million parameters (1.2 GB). The README reports higher benchmark scores for the large model on CMRC 2018 and DRCD.

### Does MacBERT support Traditional Chinese text?

The README reports results on DRCD, a Traditional Chinese reading comprehension benchmark, and notes that some models like ERNIE remove Traditional Chinese characters and are unsuitable for Traditional Chinese data. MacBERT does not carry that restriction, though the README does not detail the Traditional Chinese portion of its pre-training corpus.

## Sources

- [Issues](https://github.com/ymcui/MacBERT/issues)
- [License: Apache-2.0](https://github.com/ymcui/MacBERT/blob/master/LICENSE)
- [Project website](https://www.aclweb.org/anthology/2020.findings-emnlp.58/)
- [README](https://github.com/ymcui/MacBERT/blob/master/README.md)
- [ymcui/MacBERT on GitHub](https://github.com/ymcui/MacBERT)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/ymcui-macbert
