# Chinese-Mixtral: a Chinese-adapted Mixtral 8x7B MoE model and its quantized deployment path

> ymcui/Chinese-Mixtral continues Chinese incremental pretraining on top of Mixtral-8x7B-v0.1 and releases both base and instruction models, plus LoRA and GGUF variants. The interesting part is the deployment story: the README says llama.cpp quantization runs in as little as 16 GB of memory.

**ymcui/Chinese-Mixtral** — 中文Mixtral混合专家大模型（Chinese Mixtral MoE LLMs）

- Repository: https://github.com/ymcui/Chinese-Mixtral
- Website: https://arxiv.org/abs/2403.01851
- Stars: 612 · Forks: 43
- Language: Python
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/ymcui-chinese-mixtral

## The gap Chinese-Mixtral fills: Mixtral speaks English well, Chinese less so

Mixtral-8x7B-v0.1 is a sparse mixture-of-experts model released by Mistral.ai. Its Chinese ability is not the reason it was built. Chinese-Mixtral takes that checkpoint and performs Chinese incremental training on large-scale unlabeled Chinese data, producing a base model, then applies instruction fine-tuning to produce Chinese-Mixtral-Instruct. The intended user is someone who wants Mixtral's architecture and context length but needs Chinese generation and instruction following, and who would rather download a checkpoint than run a full pretraining job. The repository also publishes the pretraining and fine-tuning scripts, so the audience extends to teams that want to reproduce or extend the adaptation rather than just consume it. Two variants ship for different jobs: the base model for continuation, the Instruct model for question answering, writing and chat. The README is explicit that the base model needs no input template while the Instruct model requires the Mixtral-Instruct template, which is the kind of detail that decides whether your outputs look coherent or garbled.

## Inside the sparse MoE design: 8 experts, 2 active, 13B of work per token

The architecture differs from the dense LLaMA-style models that dominated earlier Chinese adaptation projects. Each FFN layer holds 8 separate experts, implemented as fully connected layers. A gating value selects the best 2 for activation. The selection happens per token, not per sequence, so different positions in the same input can route to different expert pairs. The README states the total parameter count is about 46.7B while roughly 13B parameters are active during inference. That gap is the whole point of the design: you store a large checkpoint and pay for a smaller computation. It also explains the memory profile. Weights still have to be resident or swapped in, which is why the GGUF quantization path matters so much for local use. The repository also notes native support for a 32K context, with the README claiming measured support up to 128K. Treat the 128K figure as a claim from the project rather than a guarantee; long-context behavior degrades in ways that a single number cannot capture, and the repository does not publish a per-length accuracy curve.

## Choosing between the base, Instruct, LoRA and GGUF downloads

The model download section splits artifacts by how much work you are willing to do. The full model is roughly 87 GB and needs no merge step. The LoRA version is about 2.4 GB but cannot be used alone: it must be merged with the original Mixtral-8x7B-v0.1 to become a full model, and the repository points to a wiki page for the merge procedure. The GGUF version targets llama.cpp and similar tools and is aimed at users who only need inference. That three-way split is a sensible concession to bandwidth, but it also means the repository carries three different support surfaces. A bug report about a LoRA merge failure and a bug report about GGUF quantization are not the same problem, and the README does not present a single troubleshooting path that covers all three. The selection table is the most useful part of the documentation: it pairs each variant with a scenario, and it states plainly that if you want chat interaction you should pick the Instruct model.

## Installing the Python dependencies and running a first generation

The repository ships a requirements.txt at the top level. It pins peft to 0.7.1, transformers to 4.37.2, sentencepiece to 0.1.99 and bitsandbytes to 0.42.0, requires torch at 2.0.1 or newer, and lists datasets and safetensors without version pins. Install it into a fresh environment before touching any model code.

```bash
pip install -r requirements.txt
```

After installation, the practical first step is a local inference run against a downloaded checkpoint. The README points to llama.cpp for quantized deployment and states that the minimum requirement is 16 GB of memory or VRAM. The GGUF artifacts added in the 2024-03-27 news entry include 1-bit, 2-bit and 3-bit quantizations, which is what makes that 16 GB figure plausible for a model with roughly 46.7B total parameters. The exact llama.cpp invocation is not reproduced in the README, so check the llama.cpp documentation for the current command-line flags rather than guessing at them.

If you are using the LoRA weights instead of the full or GGUF model, the merge step is mandatory. The README links to a wiki page titled model_conversion_zh for the procedure. Do not attempt to load the 2.4 GB LoRA directory as a standalone model; it will not work, and the README says so directly.

## Where Chinese-Mixtral is the wrong tool

The 16 GB floor is the first constraint, and it is a floor for quantized inference only. Full-precision serving of a 46.7B-parameter checkpoint needs substantially more, and the repository does not publish a table of memory requirements per quantization level. If your deployment target is a laptop with 8 GB of unified memory, or an edge device, this is not the model for you. The second constraint is the training corpus. The README describes the incremental training data as large-scale unlabeled Chinese data and does not publish its composition, provenance or licensing. If your compliance process requires knowing where the training text came from, the repository does not answer that question, and the Apache-2.0 license on the code does not settle the question of the data. The third constraint is the release cadence. The latest release listed is v1.2 from 2024-03-26, and the last push to the repository was on 2026-04-19. That is a long quiet stretch for a model repository, and the README's news section ends in April 2024. Anyone expecting a stream of checkpoint refreshes should look elsewhere. Finally, the project is a research artifact with a paper behind it, not a serving product; there is no documented rollback procedure, no versioned API contract, and no deprecation policy.

## How it compares with dense Chinese LLaMA adaptations

The obvious alternative in the same family is Chinese-LLaMA-Alpaca-2, from the same maintainer. The difference is architectural rather than cosmetic. Chinese-LLaMA-Alpaca-2 adapts a dense LLaMA-2 model; Chinese-Mixtral adapts a sparse MoE model. Dense models activate all parameters for every token, so inference cost scales with total size. The MoE design activates roughly 13B of about 46.7B parameters per token, which is why the README can claim local deployment at 16 GB with quantization. The trade-off is memory residency and tooling maturity: MoE support in inference stacks arrived later, and the README's own list of supported ecosystems (transformers, llama.cpp, text-generation-webui, LangChain, privateGPT, vLLM) exists partly to document which tools actually handle the architecture. A second alternative is to skip adaptation entirely and prompt the original Mixtral-8x7B-v0.1 in Chinese. That avoids a download and a merge, but it gives up the incremental Chinese training and the instruction tuning, which is the entire contribution of this repository. The repository also lists a comparison against Alpaca and GPT-4 rating in its examples directory, though the README does not reproduce those numbers.

## Licence, maintenance and what an upgrade actually costs

The repository is licensed Apache-2.0. That covers the code and scripts in this repository. It does not automatically relicense the base Mixtral-8x7B-v0.1 checkpoint, which carries its own terms on Hugging Face, and it does not resolve the question of the Chinese training data. If you are shipping a product, verify the upstream model licence separately; the Apache-2.0 badge on this repository is not a blanket clearance. On maintenance: the last push was on 2026-04-19, which is within six months of the current date, so the repository is not dormant in the strict sense. But the release history tells a different story. v1.0 arrived 2024-01-29, v1.1 on 2024-03-05, and v1.2 on 2024-03-26, with nothing after that. The news section's last entry is 2024-04-30 and it announces a different project, Chinese-LLaMA-Alpaca-3. Upgrading between model versions is not documented as a migration path. There is no changelog describing weight compatibility, no statement about whether a LoRA trained against v1.0 applies to v1.1 or v1.2, and no rollback guidance. In practice an upgrade means re-downloading a checkpoint, re-quantizing if you use GGUF, and re-validating your outputs. Plan for that cost rather than assuming an in-place swap.

## Conclusion

Adopt Chinese-Mixtral if you need a Chinese-capable MoE model you can pull down as GGUF and run on a single machine without merging anything, or if you want the LoRA weights and training scripts to build your own variant. Do not adopt it if you need a small model, a permissively licensed training corpus, or a project with documented upgrade paths; the README does not cover version migration. Before committing, check the Hugging Face repository for the exact GGUF quantization you intend to use, read the model card for the prompt template the Instruct variant expects, and confirm on the llama.cpp side that your build supports the Mixtral architecture.

## FAQ

### What is Chinese-Mixtral?

It is a Chinese adaptation of Mixtral-8x7B-v0.1. The repository releases a base model built through Chinese incremental training and an Instruct model built through further instruction fine-tuning, along with LoRA and GGUF variants.

### What license does Chinese-Mixtral use?

The repository is licensed Apache-2.0. That covers the code in the repository; the upstream Mixtral-8x7B-v0.1 checkpoint carries its own terms, so check those separately if you plan to ship a product.

### How much memory does Chinese-Mixtral need to run?

The README states that quantized inference with llama.cpp requires a minimum of 16 GB of memory or VRAM. The repository does not publish a per-quantization memory table, so treat that as a floor rather than a sizing guide.

### Can I use the LoRA version on its own?

No. The README states the LoRA version cannot be used alone and must be merged with the original Mixtral-8x7B-v0.1 to produce a full model, with the merge procedure documented on the project wiki.

### Which variant should I pick for chat?

The README recommends the Instruct model for chat interaction, and notes that it requires the Mixtral-Instruct input template while the base model does not use a template.

## Sources

- [License: Apache-2.0](https://github.com/ymcui/Chinese-Mixtral/blob/main/LICENSE)
- [Project website](https://arxiv.org/abs/2403.01851)
- [README](https://github.com/ymcui/Chinese-Mixtral/blob/main/README.md)
- [Releases](https://github.com/ymcui/Chinese-Mixtral/releases)
- [ymcui/Chinese-Mixtral on GitHub](https://github.com/ymcui/Chinese-Mixtral)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/ymcui-chinese-mixtral
