diy-llm: a Chinese-language CS336 course with six graded LLM assignments
Covers pre-training data, Tokenizer, Transformer, MoE,distributed training, Scaling Laws, inference & alignment .6 progressive code assignments for full-stack LLM learning | 涵盖预训练数据、分词器、Transformer、MoE、分布式训练、缩放定律、推理与对齐,6 项渐进代码作业,掌握 LLM 全栈知识
At a glance
- What is it?
- Datawhale's diy-llm is a VitePress course that rebuilds Stanford CS336 for Chinese learners, pairing 16 chapters of theory with six progressive coding assignments that run from a hand-written tokenizer to GRPO alignment. It is a curriculum, not a model you download and run.
- Who is it for?
- diy-llm fits a Chinese-reading engineer who already knows PyTorch and wants a structured path through the full LLM stack, from BPE to GRPO, with graded code to write along the way. It does not fit someone looking for a downloadable model to run at home, and it does not fit a reader who needs English as the primary language, since the Chinese docs are the default and the English tree is secondary.
- Can I use it commercially?
- Not without permission. GitHub finds no licence file in the repository, and without a licence all rights are reserved by default: you may read the code but not reuse it. Check the README, or ask the authors, before using it.
- Is it still maintained?
- Yes. The repository last received commits 19 days ago.
- What is it written in?
- Mainly Jupyter Notebook, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 27, 2026, and from our analysis. They are not legal advice.
DEEP OPEN-SOURCE ANALYSIS
What diy-llm actually is, and who it is written for
The repository describes itself as a Chinese-language rebuild of Stanford CS336, and the framing is explicit: the authors say they want it to be more than a translated version of the original. The stated prerequisites are Python and software engineering ability, PyTorch and neural network fundamentals, linear algebra, probability, and calculus. CUDA knowledge is listed as optional, with the note that the project includes introductory material for readers who do not have it.
The intended audience is narrow in a useful way. This is for someone who wants to write the components of a language model rather than call an API for one. The README lists sixteen chapters, and the table maps each chapter to a companion assignment, so the course is designed to be worked through in order rather than browsed. The stated localization goal is concrete: the authors say they will use domestic models such as Qwen and DeepSeek in examples, on the grounds that network conditions and available compute differ from the original course's setting.
What it is not: a runnable model, a serving stack, or a fine-tuning toolkit. Nothing in the README describes downloading weights and chatting with them. The deliverable is your own code and your own understanding.
The chapter-to-assignment pipeline
The structure is a matrix. Sixteen chapters in `docs/zh/` cover pre-training data, tokenizers, PyTorch and resource accounting, architecture and training details, mixture-of-experts, GPU optimization, high-performance GPU programming, distributed training, scaling laws, inference, data engineering, evaluation, the basic training flow, reinforcement learning with verifiable rewards, multimodal models, and an extension chapter.
Six assignments in `coursework/` carry the code. Assignment 1 asks you to implement a tokenizer, a model architecture, and an optimizer, then train a minimal language model. Assignment 2 covers profiling and benchmarking, a Triton implementation of FlashAttention-2, and distributed training code. Assignment 3 has you fit a scaling law. Assignment 4 turns Common Crawl data into a pre-training dataset with filtering and deduplication. Assignment 5 applies SFT and reinforcement learning, with GRPO named, to math problems. Assignment 6 uses `lm-evaluation-harness` and `evalscope` for multi-dimensional evaluation covering language understanding, commonsense reasoning, code, and math.
The chapter table shows assignment 1 attached to both chapter 2 and chapter 4, and assignment 2 attached to chapters 6, 7, and 8. That overlap is the design: a single piece of code is revisited as more theory lands. Chapter 1, on W&B experiment tracking and hyperparameter search, is marked as still being written, and chapter 16 is marked as updating. Everything else in the table is marked complete.
Cloning the repository and building the docs locally
The README's quick start is deliberately thin. It gives the clone command and a note that dependencies are installed per assignment.
git clone https://github.com/datawhalechina/diy-llm.git
cd diy-llmAfter that, the README points you at `docs/zh/` for reading and `coursework/` for practice. It does not list a requirements file at the repository root, and it does not name a Python version. If you want the site running locally rather than reading on the hosted page, the repository root has a `package.json` with VitePress scripts and pins `[email protected]` as the package manager.
pnpm install
pnpm devThe `dev` script runs `vitepress dev docs`, so the site is served from the `docs` directory rather than the repository root. The same file defines `build` and `preview` scripts. The dev dependencies include `markdown-it-mathjax3` for math rendering, `vitepress-plugin-mermaid` and `mermaid` for diagrams, and `vitepress-sidebar`. None of that is needed to do the assignments; it is only needed to render the book yourself.
For a first real use, the honest starting point is the coursework directory rather than the docs. The README says to install dependencies according to the specific assignment, which means the assignment folder is where the environment actually gets defined. Check there before assuming anything about Python or CUDA versions.
Where the course stops short
The dependency situation is the first real friction. The README's install section is two lines and a comment, and the repository root has no requirements file listed among the top-level entries. The environment for each assignment lives inside that assignment's directory, and the README does not enumerate what those environments contain. Anyone expecting a single `pip install` to prepare the whole course will be disappointed, and anyone without a GPU will need to check assignment by assignment whether the work is feasible on CPU.
The second limitation is language. The repository badges mark Chinese as the default, and the directory layout puts Chinese docs at `docs/zh/` with English at `docs/en/`. The chapter table in the README lists only Chinese paths. An English tree exists, but the README does not describe its coverage, so a reader who needs English cannot tell from the README how complete it is.
The third is scope. Chapters 1 and 16 are marked as incomplete or in progress. Chapter 1 covers experiment tracking and hyperparameter search, which is not optional infrastructure in practice; skipping it means running training without the tooling the course assumes. If your goal is a polished end-to-end walkthrough, the course is not finished at the edges.
Finally, the licence is not stated in the README. The repository has no licence identifier in the facts provided, so anyone planning to reuse the code or text in another project has to check the repository directly rather than assume a permissive default.
How it differs from nanoGPT and Karpathy-style walkthroughs
The obvious comparison is nanoGPT, the compact GPT training repository that many people use as a first hands-on language model. The difference is in what is left out. nanoGPT is a working training script you can run and modify; it gets you to a trained model quickly and leaves most of the surrounding stack as an exercise for the reader.
diy-llm inverts that. It is a course with a reading path, and the code is the assessment rather than the product. Where nanoGPT is one artifact, diy-llm is six assignments spread across data processing, kernel writing, distributed training, scaling-law fitting, alignment, and evaluation. Assignment 4, converting Common Crawl into a filtered and deduplicated pre-training set, has no counterpart in a minimal training repository. Assignment 2's Triton FlashAttention-2 implementation is kernel work, not model work.
The trade-off is time. You can run nanoGPT in an afternoon. Working through diy-llm's six assignments means writing a tokenizer, a transformer, an optimizer, a kernel, a distributed training loop, a scaling-law fit, a data pipeline, an alignment run, and an evaluation harness. The course also assumes Chinese as the reading language, which nanoGPT does not. If your goal is a model that trains tonight, diy-llm is the wrong tool. If your goal is to be able to explain every component, the assignment structure is the point.
Maintenance, releases, and licence status
The repository is not archived, and the last push was on 2026-09-08. Two releases exist: V0.1 on 2026-06-10 and V0.2 on 2026-08-17. The README also points to a PDF build of the course notes, distributed through the releases page, with a Datawhale watermark added deliberately so that reposted copies can be identified. That PDF is the offline reading path, and it tracks the releases rather than the `main` branch.
Upgrade cost is low for a reader and moderate for a contributor. Reading the hosted site or the PDF means no upgrade work at all; you get whatever the release contains. Building the site locally means tracking the VitePress toolchain in `package.json`, which pins `[email protected]` and lists `vitepress` at `^1.6.4` alongside a set of plugins for math, Mermaid diagrams, image viewing, and sidebars. Those are ordinary dev-dependency updates, but a broken plugin will break the build rather than the content.
On licensing, the README does not state a licence for the repository. That matters more than usual for a course, because the code in `coursework/` is meant to be copied into your own work, and the chapters are meant to be read and cited. The README's own note about watermarking the PDF shows the authors are thinking about reuse. Until a licence file is confirmed, treat redistribution of the text or code as something to verify rather than assume.
Editorial conclusion
diy-llm fits a Chinese-reading engineer who already knows PyTorch and wants a structured path through the full LLM stack, from BPE to GRPO, with graded code to write along the way. It does not fit someone looking for a downloadable model to run at home, and it does not fit a reader who needs English as the primary language, since the Chinese docs are the default and the English tree is secondary. Before committing, open the coursework directory for the assignment you care about and check whether its dependencies and GPU requirements are documented, because the README's quick start only clones the repository and says to install dependencies per assignment. Also check the licence file, since the repository's licence is not stated in the facts provided.
Frequently asked questions
Can I build my own LLM model with diy-llm?
The course is built around writing the components yourself: assignment 1 asks you to implement a tokenizer, a model architecture, and an optimizer, then train a minimal language model. Later assignments extend that to systems optimization, data processing, alignment, and evaluation. It teaches the construction rather than shipping a finished model.
How much does it cost to build your own LLM with diy-llm?
No cost figures are stated, and the repository itself is free to clone. The README does say the course accounts for the compute resources readers are likely to have and provides localized solutions, but no pricing or hardware budget is given.
Is diy-llm a free LLM model I can run at home?
No. diy-llm is a course repository containing documentation and six coding assignments, not a model with weights you download and run. The README describes it as a systematic learning path for large language models.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/datawhalechina-diy-llm)
Community notes