Model or dataset
bbruceyuan/LLMs-Zero-to-Hero avatar
bbruceyuan/LLMs-Zero-to-Hero

LLMs-Zero-to-Hero: A Chinese-Language Hands-On LLM Tutorial Series

从无名小卒到大模型(LLM)大英雄~ 欢迎关注后续!!!

2,297 stars158 forksJupyter NotebookApache-2.0

At a glance

What is it?
LLMs-Zero-to-Hero is a Jupyter Notebook repository that walks through building large language models entirely from scratch, from a basic transformer to mixture-of-experts and GRPO, with paired Bilibili and YouTube video explanations written in Chinese.
Who is it for?
LLMs-Zero-to-Hero is a good fit for Chinese-speaking practitioners who want a code-first path through modern LLM internals, from a basic transformer up to MoE, DeepSeek MLA, and GRPO, with paired video explanations. The minimum hardware the README states is an RTX 3090 or 4090.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 46 days ago.
What is it written in?
Mainly Jupyter Notebook, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What LLMs-Zero-to-Hero Is and Who It Is For

LLMs-Zero-to-Hero is a tutorial repository designed to take a practitioner from no knowledge of LLM internals to a working implementation of the key components in modern large models. The README describes the goal as going from novice to master. Every concept is implemented from scratch in Jupyter Notebooks, following the code-first, explain-while-building approach that Andrej Karpathy established with his Neural Networks: Zero to Hero series. The README explicitly credits Karpathy as the inspiration.

The primary audience is Chinese-speaking practitioners. The README text, the companion blog at yuanchaofa.com, and the video explanations on Bilibili are all in Chinese. The curriculum goes beyond Karpathy's original series by covering more recent LLM components: mixture-of-experts, the DeepSeek MLA attention mechanism, and GRPO for reinforcement learning from human feedback. This makes the repository more relevant to someone who wants to understand architectures used in production LLMs today, not just the original GPT design.

A companion repository, Hands-On Large Language Models CN, is mentioned for practitioners who find LLMs-Zero-to-Hero too steep. It uses the Hugging Face transformers library as a starting point and is described as an entry point to the ecosystem before tackling full from-scratch implementations.

How the Repository Is Organised

The repository uses a chapter-based structure. Each chapter directory contains a README.md and supporting source files. The src/ directory mirrors the chapters and also contains a video/ subdirectory with notebooks used specifically for the recorded video walkthroughs:

code
├── chapter01
│   ├── README.md
│   ├── ...
├── chapter02
│   ├── README.md
│   ├── train.py
│   ├── ...
├── src/
│   ├── hero/
│   ├── chapter01/
│   ├── chapter02/
│   ├── video/
├── README.md

The src/hero/ directory is described in the README as the location for fully self-developed LLM implementations, the endpoint of the curriculum rather than an intermediate step. Hands-on-Agents/ is a top-level directory distinct from the chapter-based content.

The primary learning path is to follow the chapter directories in order, reading each chapter's README alongside its notebooks. The video/ subdirectory contains the notebooks that were used for the Bilibili and YouTube recordings, which may differ slightly from the chapter notebooks because they are optimised for live demonstration.

Curriculum: From nanoGPT to DeepSeek MLA and GRPO

The curriculum is divided into areas. The foundation section starts with a dense transformer model (nanoGPT) built entirely from scratch. The corresponding notebook is at src/video/build_gpt.ipynb. The second foundational notebook covers mixture-of-experts, tracing the evolution from simple MoE to sparse MoE and then to the shared-expert sparse MoE used in DeepSeek (src/video/build_moe_model.ipynb).

The DeepSeek MLA algorithm receives two dedicated entries. Part 1 covers the attention mechanism without matrix absorption; Part 2 covers the matrix-absorption variant. Both are paired with blog posts on bruceyuan.com and Bilibili and YouTube videos. A section on activation functions traces the path from ReLU and GELU to swishGLU.

The GRPO section (Group Relative Policy Optimization) implements the algorithm from scratch for Agentic RAG. Its notebook is in the Hands-On Large Language Models CN companion repository rather than in this one, reflecting the split between this repository and its sibling project.

The planned curriculum extends further into areas not yet published: pre-training from scratch, supervised fine-tuning, DPO, RLHF, a Code-LLM for generating Python, and model deployment with inference optimization and quantization. The README marks several of these as forthcoming.

Running the Notebooks and Hardware Requirements

The repository does not include a package install script. To work through the material, clone the repository and open the relevant notebook for the chapter you are on. Each notebook is self-contained for its chapter's task.

The README states the minimum hardware is an RTX 3090 or 4090, which means these notebooks require a GPU for the training steps. The README includes a link to Featurize, a GPU rental platform, noting that 24 hours of RTX 4090 access is available for 9.9 yuan at the time the README was written. This is a disclosure from the author, not an endorsement by this review.

A reference implementation called BitBrain (bit-brain in the repository list) is described in the README as a fully trained miniLLM that can be accessed as a demo. It represents the endpoint of the self-developed model track and is hosted separately.

For practitioners who do not have GPU access and want to follow along with concepts rather than run training jobs, the companion blog at yuanchaofa.com provides written explanations for each chapter without requiring a local compute environment.

Comparison with Karpathy's Zero to Hero

The README cites Andrej Karpathy's Neural Networks: Zero to Hero as a direct influence, using the phrase "tribute to Andrej Karpathy." Both series share the same method: implement everything from scratch in code, explain the mathematics and architecture decisions as the code grows.

The key differences are scope and language. Karpathy's series covers the path to GPT-2 and tokenisation in English. LLMs-Zero-to-Hero extends the curriculum into MoE, DeepSeek MLA, and GRPO, which are architectures and training methods that emerged after Karpathy's series was published. The explanations in LLMs-Zero-to-Hero are in Chinese, targeting the Chinese-speaking practitioner audience. Karpathy's materials are in English.

An engineer who wants to understand the internals of more recent production architectures, including the specific MLA attention design used in DeepSeek, will find coverage here that Karpathy's original series does not address. An engineer who works primarily in English and wants a stable, finished curriculum should start with Karpathy's series first.

Limitations and What the Repository Does Not Cover

The curriculum is incomplete. The README explicitly marks pre-training from scratch, SFT, DPO, RLHF, Code-LLM, and deployment as sections that are planned but not yet published. A reader who wants a finished, end-to-end LLM training course should check the current state of the repository before committing to it as a primary resource.

The repository has no formal releases and no versioning. There is no pip-installable package. Code written in one notebook may not be reusable across chapters without modification. This is expected for a tutorial series, but it means the codebase is not a library that can be dropped into a production pipeline.

All explanatory text, README files for each chapter, and companion blog posts are in Chinese. Practitioners who do not read Chinese will be able to run the notebooks but will not be able to follow the written explanations without translation. The Bilibili and YouTube videos are in Chinese audio without confirmed English subtitles in the materials provided.

Maintenance and Licensing

The last push to the repository was on 2026-08-16. The repository has no GitHub releases. The README carries an author note written in October 2025 announcing that a companion book is being written, with followers encouraged to track updates through the author's WeChat public account or blog.

The code is Apache-2.0 licensed. The Apache-2.0 licence permits use, modification, and redistribution with attribution. The notebooks themselves are the primary artifact; they do not depend on any packages that would impose additional licence constraints beyond what JAX, PyTorch, or similar training libraries carry.

The author maintains additional projects referenced in the README: ApeRouter (an LLM API router for routing to Claude, GPT, and Kimi from within China) and ApeCode.ai (an AI coding and learning platform). These are separate products and not part of the tutorial content.

Editorial conclusion

LLMs-Zero-to-Hero is a good fit for Chinese-speaking practitioners who want a code-first path through modern LLM internals, from a basic transformer up to MoE, DeepSeek MLA, and GRPO, with paired video explanations. The minimum hardware the README states is an RTX 3090 or 4090. Anyone who prefers English-language instruction should look at Karpathy's Neural Networks: Zero to Hero, which the README directly acknowledges as its inspiration. The repository has no formal releases, and the last commit was on 2026-08-16.

Frequently asked questions

Does LLMs-Zero-to-Hero require a GPU to run?

The README states the minimum hardware is an RTX 3090 or 4090. Training notebooks require a GPU. The README links to a GPU rental platform for practitioners who do not own compatible hardware.

Is LLMs-Zero-to-Hero available in English?

The primary README, chapter explanations, and companion blog posts are in Chinese. The Bilibili and YouTube videos are in Chinese audio. The top-level repository entries do not include an English README.

Where are the LLMs-Zero-to-Hero notebooks located in the repository?

Notebooks are in the src/ directory, organized into chapter subdirectories (src/chapter01/, src/chapter02/, and so on) and a src/video/ subdirectory with notebooks recorded for the video series.

Official sources

  1. bbruceyuan/LLMs-Zero-to-Hero on GitHub
  2. Issues
  3. License: Apache-2.0
  4. Project website
  5. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/bbruceyuan-llms-zero-to-hero.svg)](https://hysenlabs.com/projects/bbruceyuan-llms-zero-to-hero)