# llms-from-scratch-cn: Building GLM4, Llama3 and RWKV from Zero in Notebooks

> Datawhale's Chinese-language tutorial walks through GPT-style pretraining, then rebuilds GLM3, Llama3 and RWKV architectures notebook by notebook. It is a teaching repository, not a model release, and several chapters are still marked as forthcoming.

**datawhalechina/llms-from-scratch-cn** — 仅需Python基础，从0构建大语言模型；从0逐步构建GLM4\Llama3\RWKV6， 深入理解大模型原理

- Repository: https://github.com/datawhalechina/llms-from-scratch-cn
- Stars: 4,381 · Forks: 600
- Language: Jupyter Notebook
- License: NOASSERTION
- Published: 2026-09-09 · Updated: 2026-09-09 · Language: en
- Canonical page: https://hysenlabs.com/projects/datawhalechina-llms-from-scratch-cn

## The gap llms-from-scratch-cn fills: architecture, not fine-tuning

Most LLM material available in Chinese covers fine-tuning and deployment: how to load a checkpoint, how to run LoRA, how to serve a model behind an API. The README states the project's focus explicitly, saying that against a background of abundant fine-tuning and deployment tutorials the maintainers concentrate on architecture implementation. That is the actual problem being solved. If you want to know why a rotary embedding is applied the way it is, or how a gated linear unit sits inside a transformer block, fine-tuning guides will not tell you.

The audience is narrow and stated plainly: readers with Python basics. The README claims that even with only a PyTorch foundation you can complete the model construction, and it frames the whole thing as educational, aimed at training and developing small but functional models rather than reproducing ChatGPT-scale systems. That framing matters. Nothing here produces a competitive model. It produces understanding of the code path that a competitive model uses.

## Two parallel tracks: Codes and Translated_Book

The repository splits its introductory material into two paths, and choosing the wrong one wastes time. The README says that readers who want a quick start should use the notebooks under Codes, described as concise code that gets you moving, while readers who want detail should use Translated_Book, which provides more related knowledge. So the same chapter exists twice at different depths.

The Codes path is organized by chapter. Chapter 2 covers text data processing with ch02.ipynb, dataloader.ipynb and exercise-solutions.ipynb. Chapter 3 covers attention with ch03.ipynb and multihead-attention.ipynb. Chapter 4 builds the GPT model itself, with ch04.ipynb and a standalone gpt.py. Chapter 5 handles pretraining on unlabeled data, with ch05.ipynb plus train.py and generate.py. Appendix A introduces PyTorch, including a DDP-script.py for distributed training, and appendix D adds extra training features. Each chapter folder also carries an exercise-solutions notebook, which is the part most tutorials omit.

## The architecture track: GLM3, Llama3 and five RWKV generations

The second half of the project lives in Model_Architecture_Discussions, and this is where it diverges from the upstream English book it credits. The README lists ChatGLM3, Llama3 and RWKV V2 through V5 as separate notebooks, each with a named contributor, and states that the directory contains each model's configuration files, training scripts and core code.

The RWKV entries are the distinctive part. RWKV replaces attention with a recurrent formulation, so reading five successive versions side by side shows how the architecture changed between generations. A single tutorial on the final version would hide that. The Llama3 notebook is credited to an external contributor and the ChatGLM3 notebook is described as loading model weights, which suggests it inspects a real checkpoint rather than training from zero. The README does not spell out what each notebook requires in terms of downloaded weights or GPU memory, and that is a real gap: you find out by opening the notebook.

## Opening the first chapter without an install step

There is no package to install and no release to download. The README gives no pip command, no setup script and no environment file. The unit of work is the notebook, so you obtain the repository and open a file. The README's only stated prerequisite is Python basics plus a PyTorch foundation, and the notebooks themselves are the deliverable.

Start with the chapter the README assigns to text data processing. Open Codes/ch02/01_main-chapter-code/ch02.ipynb in your notebook environment and run the cells in order. You should see raw text being tokenized and batched, which is the input pipeline every later chapter depends on.

Once chapter 2 runs cleanly, the natural next step is the model definition itself rather than another notebook: open Codes/ch04/01_main-chapter-code/gpt.py. That file holds the GPT implementation that ch04.ipynb walks through, and reading the module directly is faster than stepping through notebook cells when you already know what a transformer block looks like. The README does not list a requirements.txt or pinned versions anywhere, so dependency conflicts with a current PyTorch build are something you resolve yourself.

## Where the repository stops short

Three of the eight chapters are marked forthcoming in the README's chapter table: chapter 6 on fine-tuning for text classification, chapter 7 on fine-tuning from human feedback, and chapter 8 on using large language models in practice. The repository's last push was on 2026-03-26, so this is not an abandoned project, but the table is the authoritative statement of what exists and it lists those three as not yet published.

That gap has a practical consequence. The project teaches you to build and pretrain a model, and it teaches you how several production architectures are assembled, but the alignment and deployment material is absent. If your goal is to take a base model and make it follow instructions, the parts you need are the parts that are missing.

A second limitation is scale. The README describes the approach as educational and aimed at small but functional models. Nothing in the repository claims the notebooks train anything comparable to a released checkpoint, and no hardware requirements are documented. Budget your own GPU time accordingly; the repository will not tell you what that budget is.

## How it relates to rasbt/LLMs-from-scratch

The README states that the foundational section is based on rasbt/LLMs-from-scratch and thanks that author directly. The relationship is a translation and extension, not a fork in the software sense. The English original is a book with companion notebooks covering the GPT pipeline from tokenization to fine-tuning.

The difference in approach is what each project spends its pages on. The upstream book follows one architecture, the GPT-style decoder, from data through pretraining to classification and instruction fine-tuning. This repository keeps that pipeline as its first section and then adds a second section the original does not have: side-by-side reimplementations of ChatGLM3, Llama3 and RWKV V2 through V5. If you want one coherent path through a single architecture with the fine-tuning chapters complete, the English original is the better fit today. If you want to compare how different families assemble their blocks, and you read Chinese, this repository covers ground the original does not.

## Licence, maintenance and the cost of keeping up

The README displays an Apache 2.0 badge linking to LICENSE.txt, but the repository metadata reports the licence as NOASSERTION, meaning GitHub could not classify the file automatically. Those two signals disagree, and the README does not explain the discrepancy. Read LICENSE.txt yourself before reusing code, and treat the badge as a claim rather than a verified fact. Nothing here is legal advice.

On maintenance: the last push was on 2026-03-26, and the repository is not archived. There are no releases, so there is no version to pin and no changelog to track. Changes arrive as commits to notebook files, which means a chapter you ran last month can change under you without a version number moving.

Upgrade cost is mostly your own environment. Because the project pins nothing, keeping the notebooks running against a current PyTorch release is your responsibility, and a breaking change in PyTorch surfaces as a failing cell rather than a dependency conflict you can read in a lockfile. The architecture notebooks add a second cost: model weights for GLM3 and Llama3 have to be obtained separately, and the README does not document where from or how large they are.

## Conclusion

Adopt llms-from-scratch-cn if you already write PyTorch and want to read attention, tokenization and RWKV recurrence as runnable code rather than equations. Skip it if you need a production fine-tuning pipeline or a finished text: chapters 6, 7 and 8 are listed as forthcoming, and the repository has no releases. Before committing time, open Codes/ch04/01_main-chapter-code/gpt.py and one Model_Architecture_Discussions notebook to confirm the code style and the GPU assumptions match your setup, and read LICENSE.txt because the README badge points at Apache 2.0 while the repository carries a NOASSERTION licence.

## FAQ

### How are LLMs built from scratch in llms-from-scratch-cn?

The project walks through text processing, attention and the GPT model in chapter notebooks under Codes, then pretrains on unlabeled data in chapter 5 using train.py and generate.py. A separate section rebuilds GLM3, Llama3 and RWKV architectures notebook by notebook.

### How do I learn LLMs from scratch with llms-from-scratch-cn?

The README states that readers wanting a quick start should use the notebooks under Codes, while readers wanting more detail should use Translated_Book. The only stated prerequisite is Python basics with a PyTorch foundation.

### How do I build an LLM from scratch using llms-from-scratch-cn?

Work through the chapter notebooks in order, starting with ch02.ipynb for text data and reaching ch04.ipynb and gpt.py for the model itself. Chapter 5 covers pretraining with train.py and generate.py.

## Sources

- [datawhalechina/llms-from-scratch-cn on GitHub](https://github.com/datawhalechina/llms-from-scratch-cn)
- [Issues](https://github.com/datawhalechina/llms-from-scratch-cn/issues)
- [README](https://github.com/datawhalechina/llms-from-scratch-cn/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/datawhalechina-llms-from-scratch-cn
