Model or dataset
ymcui/Chinese-LLaMA-Alpaca-3 avatar
ymcui/Chinese-LLaMA-Alpaca-3

Chinese-LLaMA-Alpaca-3: Chinese Llama 3 Models, LoRA Training Scripts and Local Deployment

中文羊驼大模型三期项目 (Chinese Llama-3 LLMs) developed from Meta Llama 3

1,981 stars169 forksPythonApache-2.0

At a glance

What is it?
The third Chinese LLaMA project from ymcui adapts Meta Llama 3 with large-scale Chinese continued pretraining and instruction tuning. This review covers what the released models are, how the training and quantization scripts fit together, and where the 8K context and Apache-2.0 licence leave limits.
Who is it for?
Adopt Chinese-LLaMA-Alpaca-3 if you need a Chinese-capable 8B model you can run locally, quantize, or continue tuning with the released LoRA and instruction-tuning scripts, and if you are willing to pull weights from Hugging Face or ModelScope and follow the Llama-3-Instruct template yourself.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 164 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The gap Chinese-LLaMA-Alpaca-3 fills between Meta Llama 3 and Chinese workloads

Meta Llama 3 ships strong English and multilingual behaviour, but the project's premise is that Chinese semantic understanding and instruction following still improve with targeted continued pretraining. Chinese-LLaMA-Alpaca-3 is the third iteration of that idea, following the first and second Chinese LLaMA and Alpaca projects. It releases two families: an 8B base model (Llama-3-Chinese-8B) for text continuation, and instruction-tuned models (Llama-3-Chinese-8B-Instruct) for question answering, writing and chat. The audience is practitioners who want a Chinese-capable 8B model they can download, quantize and run on their own CPU or GPU rather than call an API. The repository also publishes pretraining and instruction-tuning scripts, so the same audience can continue training on private data. Three instruction datasets are released alongside the weights: alpaca_zh_51k, stem_zh_instruction and ruozhiba_gpt4 (4o/4T). That combination of weights plus scripts plus data is the reason to look at this project rather than a single checkpoint on a model hub. It is also the reason the repository is larger than a model card: you are getting a reproducible recipe, not just a binary.

Architecture choices: 128K vocabulary, 8K context and GQA

The README is unusually explicit about what was not changed. Llama 3 expanded the tokenizer from roughly 32K to 128K entries and switched to BPE. The project states that its own encoding-efficiency tests on Wikipedia data put Llama-3's tokenizer at about 95% of the efficiency of the Chinese LLaMA-2 vocabulary, which it had previously expanded by hand. On that basis it did not extend the vocabulary further, citing earlier work on Chinese Mixtral. For anyone who has fought with a modified tokenizer, that is a real convenience: the released models keep the original 128,256-entry vocabulary, so standard Llama-3 tooling does not need vocabulary surgery. Context length moves from 4K in the second generation to 8K natively, and the README notes that PI, NTK or YaRN can extend it further. Grouped-query attention, previously used in larger Llama-2 variants, is now present at the 8B scale, which the README frames as an efficiency improvement. The trade-off is visible in the same table: 8K is the ceiling unless you do the extension work yourself, and the project does not ship a long-context checkpoint. Keeping the stock tokenizer also means you cannot claim a tokenizer advantage over Llama 3 itself; the gain has to come from the training data and the tuning stages.

Base versus Instruct, and why v3 is the default instruction model

The model-selection table is the most useful page in the repository. Llama-3-Chinese-8B is a causal language model trained on roughly 120GB of unlabelled Chinese text with LoRA plus full embedding and LM-head updates; it takes no prompt template and is meant for continuation. The Instruct models are trained on about 5 million instruction examples and require the Llama-3-Instruct template, which the README stresses is incompatible with Llama-2-chat. Within the Instruct line, the three versions differ in initialization: v1 starts from the Chinese base model, v2 fine-tunes Meta-Llama-3-8B-Instruct directly on the 5M instruction set, and v3 is a fusion of inst-v1, inst-v2 and inst-meta followed by roughly 5K more instruction examples. The README's recommendation is blunt: if there is no specific preference, use Instruct-v3. The reported Chinese-ability scores in the table increase across the three (49.3 / 51.5 for v1, 51 for v2, with v3 described as a significant improvement over v1 and v2 on downstream tasks). Treat those as self-reported numbers from the project, not third-party evaluations. Note the training-method column: v3 is a merge, so its behaviour is not explained by any single training run, which matters if you plan to reproduce or audit it.

Installing dependencies and running a first quantized inference

The repository keeps its Python dependencies in requirements.txt at the top level. The pinned set is small: peft==0.7.1, torch>=2.0.1, transformers>=4.40.0, bitsandbytes, datasets and safetensors. Installing it is a single pip command, and the versions matter because the scripts rely on PEFT for LoRA loading.

bash
pip install -r requirements.txt

The README does not walk through a Python inference snippet in the sections available here; it points instead to the quantization and deployment tutorial section and to the supported ecosystem, listing transformers, llama.cpp, text-generation-webui, vLLM and Ollama. The practical first step is therefore to pick a runtime from that list and download the matching weights from Hugging Face (hfl) or ModelScope. If you stay inside transformers, the dependency set above is what the project expects. One detail to carry into any runtime: the Instruct models need the Llama-3-Instruct template, and the README states it is not compatible with the Llama-2-chat template. Feeding an Instruct model a Llama-2-style prompt will produce worse output, and that failure looks like a model quality problem when it is a formatting problem.

bash
pip install transformers>=4.40.0 torch>=2.0.1

The second block is only a reminder that the two core packages must satisfy the minimums in requirements.txt; installing an older transformers will not match what the training and inference scripts were written against.

Where the released scripts stop short

Two limitations stand out. First, the repository is a model-and-script release, not a serving framework. It provides pretraining and instruction-tuning scripts and a quantization and local-deployment tutorial, but the README's ecosystem list hands off to llama.cpp, text-generation-webui, vLLM and Ollama for actual serving. If you need an OpenAI-compatible endpoint with batching, continuous batching and metrics, you are assembling that from a third-party runtime, not from this repository. Second, the release cadence is old. The latest release is v3.0, dated 2024-05-30, and the last push to the repository was on 2026-04-19. Anyone evaluating this in 2026 should read the model cards and the release notes rather than assume the weights reflect newer Llama 3 variants or newer tuning recipes. The 8K context ceiling is a third constraint: the README names PI, NTK and YaRN as extension methods but does not ship an extended checkpoint, so long-document Chinese work is your project, not the repository's. None of these are defects in what was shipped; they are boundaries of the scope. The practical consequence is that you should treat the repository as a source of weights and recipes, and budget separately for the serving layer and any context extension.

Chinese Mixtral and the rest of the ymcui family as alternatives

The README itself points to the closest alternatives. Chinese-Mixtral is the sibling project that adapts Mixtral with Chinese data, and the vocabulary decision in this repository is explicitly traced to experiments done there. The difference is architectural: Mixtral is a mixture-of-experts model, so per-token compute is lower than a dense 8B model but the memory footprint for the full expert set is larger, which changes what hardware you need for local inference. The second-generation Chinese-LLaMA-Alpaca-2 is the other alternative, and it is the one to pick if you already have a 4K-context pipeline tuned to the Llama-2 chat template; the third generation breaks that template. If your requirement is a Chinese model with an expanded, Chinese-optimized vocabulary, the second-generation line is where that design lives, since this project deliberately kept Llama 3's original 128K BPE vocabulary. Choosing between them is mostly a question of which tokenizer and prompt format your existing code already assumes. There is also a Mixtral-based line and a multimodal Visual-Chinese-LLaMA-Alpaca project in the same family, so if your workload is not pure text, the third generation is not the branch to look at.

Licence, upgrade cost and what a fork inherits

The repository is Apache-2.0, which covers the code, the scripts and the released data. The model weights are a separate matter: they are derived from Meta Llama 3, so the Meta Llama 3 licence and its acceptable-use terms travel with the weights. The README does not restate those terms in the sections available here, so read them at the source before shipping anything. Upgrade cost is mostly about the template break. Moving from the second generation to the third means rewriting prompt construction, because Llama-3-Instruct is not compatible with Llama-2-chat, and re-validating any evaluation harness that assumed the older format. Moving between Instruct v1, v2 and v3 is cheaper: they share the template, so a swap is a weights swap, but v3 is a fusion model, which means you cannot reproduce it from a single training run using only the v1 or v2 recipe. If auditability of the training path matters more than the reported score, v2 is the simpler artefact to reason about. The released instruction datasets are a separate asset: if you fine-tune on them, your derivative work inherits both the repository's Apache-2.0 terms and whatever the underlying data sources require.

Frequently asked questions about Chinese Llama 3 models

The questions collected below are the ones a new user is most likely to hit before downloading anything: what role Meta Llama 3 plays in the project, whether the llama and alpaca naming implies a merged model, which checkpoint to pick, whether local CPU or GPU deployment is supported, and how much context the models actually handle. Each answer is drawn from the README and the repository files. Two of the questions come from search data and are answered as the README frames them; the rest are the practical questions that follow from the model-selection table. Where the README is silent, for example on the exact Meta Llama 3 licence terms, the answer says so rather than guessing.

Editorial conclusion

Adopt Chinese-LLaMA-Alpaca-3 if you need a Chinese-capable 8B model you can run locally, quantize, or continue tuning with the released LoRA and instruction-tuning scripts, and if you are willing to pull weights from Hugging Face or ModelScope and follow the Llama-3-Instruct template yourself. Do not adopt it if you need long-context Chinese work beyond 8K without adding PI, NTK or YaRN yourself, or if you want a maintained framework with recent releases: the latest release is v3.0 from 2024-05-30 and the last push was on 2026-04-19. Before committing, verify the model variant you download (base versus instruct v1, v2 or v3), confirm the licence and Meta Llama 3 terms you inherit, and check that your inference stack reads the Llama-3 chat template the README points to.

Frequently asked questions

What is Llama 3 used for in Chinese-LLaMA-Alpaca-3?

The project uses Meta Llama 3 as the starting point for continued pretraining on roughly 120GB of Chinese text and instruction tuning on about 5 million examples, producing base and instruct models for Chinese text continuation and chat.

Is Chinese-LLaMA-Alpaca-3 a mix of llama and alpaca models?

The naming follows the project's own convention rather than a merged animal model: the base checkpoint is called Llama-3-Chinese and the instruction-tuned checkpoint is called Llama-3-Chinese-Instruct, continuing the naming from the first and second Chinese LLaMA and Alpaca projects.

Which Chinese-LLaMA-Alpaca-3 model should I download?

The README recommends the Instruct version for chat interaction and states that Instruct-v3 should be preferred unless there is a specific reason to choose otherwise; the base model is for text continuation and needs no prompt template.

Does Chinese-LLaMA-Alpaca-3 support local CPU or GPU deployment?

The README lists quantization and local deployment on a personal computer's CPU or GPU as one of the project's main offerings, and names transformers, llama.cpp, text-generation-webui, vLLM and Ollama as supported runtimes.

What context length does Chinese-LLaMA-Alpaca-3 support?

The models use an 8K context window, up from 4K in the second generation, and the README notes that PI, NTK or YaRN can be applied to extend it further.

Official sources

  1. Issues
  2. License: Apache-2.0
  3. README
  4. Releases
  5. ymcui/Chinese-LLaMA-Alpaca-3 on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/ymcui-chinese-llama-alpaca-3.svg)](https://hysenlabs.com/projects/ymcui-chinese-llama-alpaca-3)