# Chinese-LLaMA-Alpaca: LoRA Weights, a Merged Model, and CPU Inference

> The ymcui/Chinese-LLaMA-Alpaca repository publishes LoRA patches that add a Chinese vocabulary to Meta's LLaMA, plus instruction-tuned Alpaca variants. It is a merge-and-quantize workflow, not a drop-in model download, and its own README now points users to the third generation.

**ymcui/Chinese-LLaMA-Alpaca** — 中文LLaMA&Alpaca大语言模型+本地CPU/GPU训练部署 (Chinese LLaMA & Alpaca LLMs)

- Repository: https://github.com/ymcui/Chinese-LLaMA-Alpaca
- Website: https://github.com/ymcui/Chinese-LLaMA-Alpaca/wiki
- Stars: 18,938 · Forks: 1,837
- Language: Python
- License: Apache-2.0
- Published: 2026-09-09 · Updated: 2026-09-09 · Language: en
- Canonical page: https://hysenlabs.com/projects/ymcui-chinese-llama-alpaca

## What the Chinese-LLaMA-Alpaca repository actually publishes

This project targets a specific problem: the original LLaMA tokenizer handles Chinese text poorly, so the same sentence costs more tokens and the model's Chinese semantic understanding is weaker. The repository's answer is to extend the original LLaMA vocabulary with Chinese tokens, continue pretraining on Chinese text, and then instruction-tune the result on Chinese instruction data. The README states that the released models are 7B, 13B and 33B, each in a base version plus Plus and Pro variants, and that the Pro series from v5.0 was aimed at longer and better answers. The audience is Chinese NLP researchers and engineers who want to run or fine-tune these weights locally, including on a laptop CPU. One detail matters more than the rest: what is published here are LoRA weights, described in the README as a patch on top of the original LLaMA model. They cannot be used alone. The README is explicit that the Chinese LLaMA and Alpaca LoRA models must be paired with the original LLaMA model and merged before they produce a complete checkpoint. That single constraint shapes the entire setup path.

## How the vocabulary extension and instruction tuning fit together

The pipeline has three stages. First, the Chinese vocabulary is added to the base LLaMA tokenizer and the model is pretrained further on Chinese text, producing the Chinese LLaMA base model. Second, Chinese instruction data is used to fine-tune that base into Chinese Alpaca, which is the instruction-following variant. Third, the training scripts and inference scripts ship in the repository so a user can continue from there. The README contrasts the two families directly: Chinese LLaMA is a base model for text continuation, while Chinese Alpaca is the instruction model for question answering, writing and multi-turn chat. The tokenizer difference is visible in the vocabulary sizes the README lists, 49953 for Chinese LLaMA and 49954 for Chinese Alpaca, the extra entry being a pad token. That distinction is not cosmetic. The README warns that Chinese Alpaca requires an input template, and that using it without the template degrades results, while Chinese LLaMA should not be used for instruction following at all. In the Hugging Face inference script, the switch is a flag: Chinese LLaMA runs without extra arguments, and Chinese Alpaca needs --with_prompt. In llama.cpp the equivalent choice is -p for the base model and -ins for instruction and chat mode.

## Installing the dependencies and merging a LoRA into a full model

The repository pins its Python dependencies in requirements.txt. These versions are old, so a fresh virtual environment is the sane starting point rather than installing into an existing one.

```bash
pip install -r requirements.txt
```

The file pins torch==1.13.1, transformers==4.30.0, sentencepiece==0.1.97 and peft from a specific git commit. If your CUDA build of torch differs, resolve that before running anything else, because the merge script loads the base model through transformers and peft. The merge step is the part users skip and then wonder why nothing loads. The README's merge section describes combining the downloaded LoRA weights with the original LLaMA checkpoint to reconstruct a full model. The exact script name and arguments are in that section of the README and the wiki; the repository does not restate them in the top-level file. What the README does state clearly is the order: download the LoRA weights, obtain the original LLaMA model, merge, and only then quantize. Skipping the merge leaves you with an adapter that most runtimes will not load as a standalone model.

## Quantizing and running on a CPU with llama.cpp

After merging, the README's local inference section covers quantization and deployment on a personal computer. The repository ships ready-made example files under examples/, including q4_7b-13b, q8_7b-13b-p7b, q8_13b-p7b-p13b and f16-p7b-p13b-33b, which correspond to quantization levels and model sizes. Those directories exist to be copied into a llama.cpp checkout, not executed from this repository. The README also notes that llama.cpp supports an 8K context without modifying the model, with the method discussed in the project's discussions area, and that a pull request adds 4K+ context support under transformers. For the instruction-tuned models, the README specifies the -ins parameter to start instruction and chat mode in llama.cpp, and for text-generation-webui it notes that --cpu runs the model without a graphics card. The README's model selection table also says text-generation-webui is not suited to chat mode for the base Chinese LLaMA model, and that LlamaChat expects you to pick "LLaMA" or "Alpaca" when loading depending on which you merged.

## Where this project is the wrong tool

The most consequential limitation is licensing, and it is not buried. The README states that Meta's LLaMA model prohibits commercial use and that no official weights were open sourced, which is precisely why this project distributes LoRA patches instead of full checkpoints. The repository's own licence file is Apache-2.0, but that covers the code and the LoRA weights, not the base model you must supply. Anyone planning a commercial deployment needs to treat the base model's terms as the binding constraint, and the README does not offer a path around it. The second limitation is that this is a first-generation project. The README's news section carries an entry dated 2024/04/30 announcing Chinese-LLaMA-Alpaca-3 and recommending that all first and second generation users upgrade. The last push to this repository was on 2026-04-19, and the most recent release listed is v5.0 from 2023-07-19, so the model line itself has not moved in years. Third, the dependency pins are dated: transformers 4.30.0 and torch 1.13.1 will conflict with newer CUDA stacks and newer peft releases, and the README does not document a rollback path for a failed merge. If your goal is a modern Chinese instruction model, this repository is the wrong starting point.

## How it differs from Chinese-LLaMA-Alpaca-2 and -3

The obvious alternative is the project's own successor, Chinese-LLaMA-Alpaca-2, which the README links at the top and in its news entries, covering the LLaMA-2 generation and releasing Chinese-LLaMA-2-13B and Chinese-Alpaca-2-13B in its v2.0 release. The third generation, Chinese-LLaMA-Alpaca-3, builds on Llama-3 and ships Llama-3-Chinese-8B and Llama-3-Chinese-8B-Instruct. The difference in approach is not just a newer base model. The successor projects inherit the same vocabulary-extension and instruction-tuning method, so the workflow of merge-then-quantize carries over, but they target bases whose licence terms and tokenizer behaviour differ from the original LLaMA. There is also a separate multimodal line, Visual-Chinese-LLaMA-Alpaca, for visual question answering and dialogue. Choosing between them is mostly a question of which base model you can legally obtain and which generation your tooling supports. The README itself frames the upgrade as a recommendation rather than an option, which is a fair signal about where maintenance effort goes.

## Maintenance, licence scope and the upgrade cost

The last push to this repository was on 2026-04-19, but that date should not be read as model development. The newest release in the list is v5.0 from 2023-07-19, and the README's own news entries stop pointing forward to new work in this repository, instead directing readers to the third generation project. Treat this as a frozen, documented artifact rather than something receiving model updates. The practical upgrade cost is the merge pipeline itself: moving to Chinese-LLaMA-Alpaca-2 or -3 means re-downloading a different base model, re-running the merge with a new LoRA, and re-quantizing, because the example files under examples/ are tied to specific model sizes and quantization levels. On licensing, the repository's Apache-2.0 covers the code and the published LoRA weights, while the README states plainly that the original LLaMA model forbids commercial use and that the base weights must come from elsewhere. This is a factual constraint about the base model, not legal advice; if commercial use matters, the base model's terms are the thing to read, and the README does not claim to resolve them.

## Conclusion

Adopt this repository if you specifically need the first-generation Chinese LLaMA and Alpaca weights, are willing to obtain the original Meta LLaMA checkpoint yourself, and can run the merge and quantization scripts on a machine with enough RAM for a 7B or 13B model. Do not adopt it if you want a model you can load straight from a hub, if you need commercial-use clarity, or if you are starting a new project in 2026, because the README's own news entry recommends that all first and second generation users upgrade to Chinese-LLaMA-Alpaca-3. Before committing, verify three things: that the base LLaMA weights are legally available to you, that the pinned transformers==4.30.0 and torch==1.13.1 in requirements.txt match your environment, and that the LoRA size you pick (7B, 13B or 33B) fits the RAM of the machine that will run the merge.

## FAQ

### Can Chinese-LLaMA-Alpaca be used on its own without the original LLaMA model?

No. The README states that the published Chinese LLaMA and Alpaca LoRA models cannot be used alone and must be paired with the original LLaMA model, then merged to reconstruct a complete checkpoint.

### Which variant should I pick, Chinese LLaMA or Chinese Alpaca?

The README's comparison table says Chinese LLaMA is a base model suited to text continuation, while Chinese Alpaca is instruction-tuned for question answering, writing and multi-turn chat. Chinese Alpaca requires an input template, and using it without one degrades output.

### How do I run Chinese-LLaMA-Alpaca on a CPU?

The README describes quantizing the merged model and deploying it locally, with example quantization files under examples/ for llama.cpp. For the instruction-tuned models, llama.cpp uses the -ins parameter to start instruction and chat mode.

## Sources

- [License: Apache-2.0](https://github.com/ymcui/Chinese-LLaMA-Alpaca/blob/main/LICENSE)
- [Project website](https://github.com/ymcui/Chinese-LLaMA-Alpaca/wiki)
- [README](https://github.com/ymcui/Chinese-LLaMA-Alpaca/blob/main/README.md)
- [Releases](https://github.com/ymcui/Chinese-LLaMA-Alpaca/releases)
- [ymcui/Chinese-LLaMA-Alpaca on GitHub](https://github.com/ymcui/Chinese-LLaMA-Alpaca)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/ymcui-chinese-llama-alpaca
