# Baby-Llama2-Chinese: a 24GB single-GPU path from raw Chinese corpus to a chat model

> DLLXW/baby-llama2-chinese is a Python repository that walks through pretraining and SFT for a small Chinese Llama2, with a 63.4B-token corpus, a ChatGLM2 tokenizer, and a unified inference script for old checkpoints.

**DLLXW/baby-llama2-chinese** — 用于从头预训练+SFT一个小参数量的中文LLaMa2的仓库；24G单卡即可运行得到一个具备简单中文问答能力的chat-llama2.

- Repository: https://github.com/DLLXW/baby-llama2-chinese
- Stars: 2,923 · Forks: 354
- Language: Python
- License: MIT
- Published: 2026-09-24 · Updated: 2026-09-24 · Language: en
- Canonical page: https://hysenlabs.com/projects/dllxw-baby-llama2-chinese

## What Baby-Llama2-Chinese is for

Most Chinese LLM tutorials start from a downloaded checkpoint and a LoRA script. This repository starts earlier. Its stated goal is a 500M to 1B parameter Llama2-Chinese built from scratch, plus a codebase that covers pretraining, SFT instruction tuning, a reward model and reinforcement learning as one flow. The README says the reward model and RL parts are still to be done, so what exists today is the pretraining and SFT half.

The audience is narrow and specific: people who want to run the full pipeline themselves rather than call an API, and who have one 24GB GPU rather than a cluster. The README is blunt about the author's own hardware, four 3090s, and notes that 63.4B tokens with a 300M parameter model is the limit of what was trained without DeepSpeed or Megatron. That honesty is useful. It tells you the target scale before you download anything.

The repository also positions itself as teaching material. The README says it will compile a set of LLM learning resources, and that part is marked as in progress. Treat the code as the deliverable and the write-up as partial.

## The ChatGLM2 tokenizer and the uint16 trick

The tokenizer choice is the most interesting design decision in the repo. Rather than train a custom vocabulary, the project uses the ChatGLM2-6B tokenizer, with a vocabulary size of 64793. The README explains why: Llama's official vocabulary contains only around 700 Chinese entries, which is the stated reason for its weak Chinese ability. Swapping in a Chinese-heavy vocabulary is the fix.

The second reason is storage. 64793 sits just under 65535, the top of the uint16 range, so each token costs two bytes instead of the four an int32 would need. The README calls this out as halving storage when the corpus is large. That is a real constraint doing real work: the pretraining data is 63.4B tokens and 118GB on disk, and the format described is a single flat np.uint16 array written to a .bin file.

The trade-off is that you inherit someone else's vocabulary and its quirks. The release notes mention a fix for ChatGLM tokenizer compatibility with newer Transformers versions, which is exactly the kind of maintenance cost you take on when you depend on another project's tokenizer files. The repository ships a chatglm_tokenizer/ directory, so the tokenizer is vendored rather than fetched, but the surrounding transformers library still moves.

## Data flow: from Baidu Netdisk to a .bin file

The pipeline has two preprocessing stages before any training happens. First, data cleaning lives in the data_clean directory, with functions for short-text filtering, Minhash and Simhash deduplication, format conversion and merging multiple datasets. The README gives measured results on the budubaike dataset: 5634898 rows before, 3605212 after short-text filtering and format conversion, then down to 2736033 after Minhash deduplication, or 3548779 after Simhash. The recorded times are 552 seconds, 4 hours and 23 minutes respectively. The README recommends Minhash for quality and notes it is the slower option, and recommends parquet for storage.

Second, data_process.py tokenizes each sample, appends an <eos> token to separate samples, and concatenates everything into one uint16 array saved as pretrain_data.bin in ./data. The README notes that mmap is an option if the corpus is too large for memory.

Then pretrain.py trains on that binary file. The README's example uses torchrun with four processes on four 3090s, and says the output lands in out/pretrain. After that, sft_data_process.py turns the alpaca-zh and bell SFT corpora into sft_data.csv under ./sft_data, and sft.py produces a model in out/sft. The README says new SFT corpora can be added by extending the script. Note that the corpus itself is distributed through Baidu Netdisk links with extraction codes, not through a package registry, which is the first practical hurdle for anyone outside that ecosystem.

## Installing Baby-Llama2-Chinese and running a first inference

There is no package on PyPI and no install command in the README. You clone the repository and install requirements.txt, which pins torch==2.0.1, numpy==1.23.5, scikit-learn==1.3.0, tqdm==4.64.1, pytest==7.4.0, Requests==2.31.0, and constrains sentencepiece to >=0.1.99,<0.3 and transformers to >=4.33.2,<5. jieba and pandas are unpinned.

```bash
pip install -r requirements.txt
```

After that you can run the preprocessing scripts. The README's sequence for the pretraining path is data_process.py, then pretrain.py via torchrun, then sft_data_process.py, then sft.py. The tokenized corpus has to be downloaded separately from the Baidu Netdisk link and placed in ./data before data_process.py will find it, and the data_path_list inside data_process.py has to be edited to match what you downloaded.

```bash
python data_process.py
```

The command reads the corpus paths you listed and writes pretrain_data.bin into ./data. If the paths are wrong or the files are missing, this is where it fails, before any GPU work starts.

The fastest way to see something work without training is infer.py, added on 2026-08-13. It loads pretrained or SFT weights and, per the README, infers model dimension, layer count, vocabulary size and KV head count from the checkpoint, so you no longer edit hardcoded values in eval.py. It also handles DDP and torch.compile prefixes and causal attention masks from older saves.

```bash
python infer.py \
  --checkpoint /path/to/sft_model.pth \
  --mode sft \
  --prompt "世界上最大的动物是什么？"
```

For a pretrained checkpoint you switch --mode to pretrain and give a continuation prompt. With no --prompt the script drops into interactive mode. Device selection defaults to CUDA, then Apple Silicon MPS, then CPU; other flags are listed by python infer.py --help.

## Where the repository is thin

The reward model and reinforcement learning stages promised in the introduction do not exist yet, and the README says so. Anyone arriving for a full RLHF pipeline will find the second half missing.

Distributed training is also narrower than the project vision implies. The vision mentions DeepSpeed and Megatron, but the author's own note says the 63.4B-token, 300M-parameter run did not use them. What the README documents is torchrun with --standalone and --nproc_per_node, which is data-parallel training, not the sharded pipeline and tensor parallelism that DeepSpeed and Megatron provide. If your plan depends on training something much larger than 300M parameters, this repository is not the tool, and the README does not claim otherwise.

The low-quality-text filtering section of the README is literally marked as to be supplemented. Deduplication and short-text filtering are documented with numbers; the filtering heuristics are not. That matters because corpus quality drives the result more than most hyperparameters, and you would be trusting code you have to read rather than documentation you can follow.

Finally, the maintenance signal is mixed. The last push was on 2026-08-13, which is recent, and the repository is not archived. But there are no retrieved releases, so there is no versioned artifact to pin against. The recent commits fixed a single-GPU sampler, a sequence-length off-by-one, DDP gradient accumulation synchronization and multi-epoch shuffle, which suggests the training code had correctness bugs until recently. If you trained on an older copy, that is worth knowing.

## How it compares to nanoGPT-style and LoRA-based projects

The natural alternative is karpathy/llama2.c, which the README itself cites as the reference for custom tokenizer construction. llama2.c is smaller in scope: it trains a Llama2-style model and its emphasis is a compact, readable implementation in C with a Python training script, and it targets English text. Baby-Llama2-Chinese differs in three concrete ways. It vendors a Chinese-optimized tokenizer instead of training one, it ships a 63.4B-token pre-tokenized corpus in uint16 form, and it includes an SFT stage with alpaca-zh and bell data. If your goal is to understand the minimal training loop, llama2.c is the cleaner read. If your goal is a Chinese chat model you trained yourself, the corpus and tokenizer here save you the two most tedious weeks.

The other comparison is LoRA fine-tuning of an existing Chinese checkpoint. That path skips pretraining entirely, costs far less compute, and gives a better model per GPU-hour in almost every practical case. Baby-Llama2-Chinese is the wrong tool if you just want a working Chinese assistant. It is the right tool if you want to watch the whole process, including the parts that a fine-tuning script hides: tokenization, binary packing, sampler correctness, gradient accumulation across processes.

## Licence and upgrade cost

The repository is MIT licensed, and the LICENSE file sits at the top level. MIT is permissive, so reuse and modification are straightforward as far as the code goes.

The code licence is not the whole picture. The README points at pretraining corpora from Wikipedia, BaiduBaiKe, C4_zh, WuDaoCorpora and shibing624/medical, and at a tokenizer derived from ChatGLM2-6B. Those sources carry their own terms, which the README does not restate, and the repository does not bundle them. If you plan to release a model trained on this pipeline, the corpus terms are the thing to check, not the MIT header. This is a description of what the README does and does not say, not legal advice.

Upgrade cost is mostly dependency drift. requirements.txt pins torch==2.0.1 and constrains transformers to >=4.33.2,<5, and the release notes already record one compatibility fix for the ChatGLM tokenizer under newer Transformers. The infer.py script was written specifically to load historical weights across DDP, torch.compile and attention-mask format changes, which tells you that checkpoint compatibility has been a recurring cost. Expect to spend time on that boundary rather than on the training loop itself.

## Frequently asked questions

The questions below cover the points a new user hits first: hardware, the corpus download, and how to run a trained model without editing eval.py.

## Conclusion

Adopt it if you want to see the whole pretrain-to-SFT loop on one 24GB card and are willing to fetch a 118GB corpus from Baidu Netdisk. Skip it if you need a reward model, RLHF, or DeepSpeed and Megatron scaling today, since the README lists those as pending. Before committing, open pretrain.py and confirm the max_seq_len, dim, n_layers, n_heads and batch_size values fit your card, and check whether your tokenizer files load under the transformers version pinned in requirements.txt.

## FAQ

### What GPU do I need to run Baby-Llama2-Chinese?

The repository description states it runs on a single 24GB card, and the pretraining example in the README uses torchrun with four 3090s. The README advises lowering batch_size if you run out of VRAM, and adjusting max_seq_len, dim, n_layers and n_heads in pretrain.py to change model size.

### Where do I get the pretraining corpus for Baby-Llama2-Chinese?

The README links a Baby-llama2-chinese Corpus on Baidu Netdisk with extraction code 6unr, described as tokenizer-processed pretraining data totaling 63.4B tokens and 118GB. You download it, place it in ./data, and edit data_path_list in data_process.py to match.

### How do I run inference with Baby-Llama2-Chinese without editing eval.py?

Use infer.py, added in the 2026-08-13 update. It accepts --checkpoint, --mode (sft or pretrain) and --prompt, and infers model dimension, layer count, vocabulary size and KV head count from the checkpoint. With no --prompt it enters interactive mode.

### Does Baby-Llama2-Chinese include reward model training and RLHF?

No. The README lists the reward model and reinforcement learning stages as pending work, while pretraining and SFT are implemented in pretrain.py and sft.py.

## Sources

- [DLLXW/baby-llama2-chinese on GitHub](https://github.com/DLLXW/baby-llama2-chinese)
- [Issues](https://github.com/DLLXW/baby-llama2-chinese/issues)
- [License: MIT](https://github.com/DLLXW/baby-llama2-chinese/blob/main/LICENSE)
- [README](https://github.com/DLLXW/baby-llama2-chinese/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/dllxw-baby-llama2-chinese
