nndl/llm-beginner: a Chinese LLM course that grades your code
《大模型与智能体》电子书与 6 个编程任务:Transformer、mini-GPT、SFT/DPO、RAG、工具调用与编程智能体。
At a glance
- What is it?
- The nndl/llm-beginner repository pairs a 17-chapter Chinese ebook with six Python tasks, from hand-written attention to a coding agent, and ships self-check scripts that import your implementations by signature.
- Who is it for?
- Adopt nndl/llm-beginner if you read Chinese and want to write attention, a mini-GPT, LoRA, DPO, RAG and a ReAct loop yourself rather than call an API. Skip it if you need a runnable reference implementation, since the repository leaves src/ to the learner, or if you cannot read Chinese, because the prose, task READMEs and tutor prompts are Chinese-first.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 24 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 17, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What nndl/llm-beginner actually hands you
This is a course, not a library. The repository holds the preprint PDF of 《大模型与智能体》 by 邱锡鹏 (17 chapters split into shared foundations, large models, agents, and boundaries and the future), plus six task directories that walk from Transformer to a mini coding agent. The README states the book and the exercises can be studied independently, and that the exercises assume Python and deep-learning background. The intended reader is someone who wants to implement the machinery rather than consume it: task 1 asks for scaled dot-product attention, multi-head attention and a full encoder block written by hand, with matplotlib heatmaps of attention weights over a Chinese sentiment classification dataset. The suggested pace is explicit, two weeks for task 1 up to five or six weeks for task 6, and the README notes those are learning-rhythm suggestions that can be adjusted. If you are looking for a pip-installable package, this is the wrong repository: the README says the implementation code is completed by the learner in each task's src/ directory.
How the six tasks are wired together
The progression is Transformer, mini-GPT, SFT and DPO, RAG, tool-calling agent, coding agent, and each stage feeds the next. Task 1 warms up causal masking for task 2, which asks for a hand-written BPE tokenizer, RoPE and a KV cache on top of a decoder-only model, explicitly extending the nanoGPT walkthrough in the second edition of the practice book by adding those three pieces. Task 3 asks for hand-written LoRA plus SFT and DPO. Task 4 builds retrieval, reranking and generation. Task 5 implements a ReAct loop with tool calls and error recovery. Task 6 produces an agent that edits code, runs tests and iterates. Every task directory has the same four files: requirements.txt, data/download.py, eval/run.py and eval/tutor_prompt.md, and the shared runner _eval_harness.py sits at the repository root. The model ecosystem is Qwen throughout, and the README says Chinese is the default language, switching to English only where English data is clearly better, such as some small-model pretraining corpora.
Installing one task and running its self-check
Each task is independent, so you can install dependencies per task or share one environment. Python 3.10 or newer is required, with 3.11 or 3.12 recommended. From the repository root, install task 1's dependencies:
pip install -r task-1-transformer/requirements.txtIf Hugging Face is unreliable from your network, the README suggests setting a mirror before downloading; the download scripts also prompt for this when files are missing:
export HF_ENDPOINT=https://hf-mirror.com
# Windows PowerShell: $env:HF_ENDPOINT = "https://hf-mirror.com"Then move into the task, fetch its data, write your implementation under src/, and run the self-check:
cd task-1-transformer
python data/download.py
# write your implementation in src/ per the task README's implementation contract
python eval/run.pyThe download command for task 1 pulls ChnSentiCorp, a Chinese sentiment classification dataset. Task 2's download takes an optional --dataset flag with poetry, tinystories or skypile and defaults to poetry, a roughly 49KB Tang-poetry file meant to validate the pipeline in about five minutes. Task 4 accepts --skip-models to fetch only the PDF and verify gold_qa, and task 6 accepts --with-swebench to additionally pull sampled SWE-bench Lite metadata. When the run finishes, eval/run.py prints each item and writes structured results to eval/result.json in UTF-8. Items report [通过], [跳过] or [失败]; a skip means a prerequisite such as a model, checkpoint or dataset is missing and is not an error. The README warns that eval/run.py must be run inside the repository because it imports _eval_harness.py from the root, so copying a single task directory elsewhere breaks the import.
The self-check grades contracts, not your experiments
This is the most interesting design decision in the repository and also the easiest to misread. The self-check imports your code by the class and function signatures listed in each task README's implementation-convention table, which means the signatures are a hard interface, not a suggestion. Miss one and the harness cannot score you. What it verifies is narrow: the README describes it as checking key contracts such as attention numerical correctness, recall and task success rate, and states plainly that it is a lower-bound check on whether things run correctly, not a substitute for the comparisons and ablations in each task's experiment section. So a green result.json does not mean you have a good model; it means your components satisfy the stated contracts. The three-state output is sensible for a course, since a missing checkpoint should not look like a bug, but it also means a learner can accumulate skips and read a mostly-skipped report as success. Read the per-item output, not just the file.
Where this course will frustrate you
The repository contains no reference implementations of the six tasks. The README says the code is written by the learner in src/, and the tutor_prompt.md files are prompts you paste into Claude, Qwen or DeepSeek alongside your code to get a review organized around that task's check items. That is a deliberate pedagogy, and it means you cannot diff your work against a canonical answer inside this repository. The second constraint is language. The book, the task READMEs and the tutor prompts are Chinese-first, with English reserved for data where it is clearly better, so a reader without Chinese cannot use most of the material. Third, hardware: the README says resource requirements vary by task and that memory use depends on model, precision, sequence length and batch size, with the advice to run a small configuration first. Task 2's advanced tier, a SkyPile-150B subset at roughly 1GB or more, is flagged as GPU-recommended, while TinyStories at around 100MB is described as runnable on CPU. Task 6 budgets five to six weeks, which is a real commitment. If you want a working mini-GPT to read today rather than write over three weeks, this is the wrong tool.
How it differs from nanoGPT and the Annotated Transformer
The obvious alternatives are the projects this course cites as references. Karpathy's nanoGPT is a complete, runnable GPT training script; you clone it and train. The Annotated Transformer is a line-by-line walkthrough of the original paper with working code attached. nndl/llm-beginner inverts that relationship: it names both as reading, then asks you to produce the artifact yourself and checks it against contracts. The difference matters when you get stuck. With nanoGPT you can read the answer; here you get a tutor prompt and a failing item in result.json. The other difference is coverage. nanoGPT stops at pretraining, while this course continues through LoRA, DPO, RAG, a ReAct tool loop and a coding agent, and it adds RoPE and a KV cache on top of the nanoGPT material, which the README notes the practice book only discusses without implementing. If your goal is to understand why a KV cache exists, writing one under a self-check is a stronger exercise than reading about it. If your goal is a trained checkpoint this week, nanoGPT is the shorter path.
Maintenance, licence and what a fork inherits
The repository is not archived, and the last push was on 2026-09-06, which is recent. A book-pdf release was published the same day, containing the full PDF, and the README notes the manuscript is in pre-publication preparation with content updated as revisions land. That is the upgrade model: you pull the repository and re-read the affected chapter or task README, since there is no package to version and no changelog beyond the releases list. The licence is MIT, which is permissive and places few obligations on reuse; the repository also ships a LICENSE file at the root. Note that the licence covers the repository's code and materials, and the README describes the book as a pre-publication electronic draft, so if you plan to redistribute or translate the text, check the LICENSE file and the release page rather than assuming the MIT grant settles every question. This is a description of what the repository states, not legal advice.
Editorial conclusion
Adopt nndl/llm-beginner if you read Chinese and want to write attention, a mini-GPT, LoRA, DPO, RAG and a ReAct loop yourself rather than call an API. Skip it if you need a runnable reference implementation, since the repository leaves src/ to the learner, or if you cannot read Chinese, because the prose, task READMEs and tutor prompts are Chinese-first. Before starting, open one task's README and confirm the class and function signatures in its implementation-convention table, then run python data/download.py and python eval/run.py from inside the repository so _eval_harness.py stays importable.
Frequently asked questions
What is nndl/llm-beginner for beginners?
It is a Chinese ebook plus six progressive Python tasks covering Transformer, mini-GPT, SFT and DPO, RAG, tool-calling agents and a coding agent. The README states the book and the exercises can be studied independently, and that the exercises require Python and deep-learning background.
How long does nndl/llm-beginner take to work through?
The README gives per-task estimates: two weeks for the Transformer task, three weeks for mini-GPT, two to three weeks for SFT and DPO, two weeks each for RAG and the tool agent, and five to six weeks for the coding agent. It notes these are learning-rhythm suggestions that can be adjusted to your background and experiment scale.
How difficult is nndl/llm-beginner?
The README states the exercises need Python and deep-learning foundations, and each task lists prerequisite abilities in its own description. The self-check only verifies key contracts such as attention numerical correctness, recall and task success rate, so passing it is a lower bound rather than proof of a good model.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/nndl-llm-beginner)