Model or dataset
fancyboi999/ai-engineering-from-scratch-zh avatar
fancyboi999/ai-engineering-from-scratch-zh

ai-engineering-from-scratch-zh Review: A 523-Lesson Chinese AI Engineering Course

Agent工程师最全学习路径 · 从零精通 AI 工程 · 20 阶段 503 课 · 中文全量翻译 + 配套站点 + 动画讲解视频 · 如何成为 AI Agent 工程师的修成指南

1,149 stars182 forksPythonMIT

At a glance

What is it?
The Simplified Chinese derivative of AI Engineering from Scratch ships 523 lessons across 20 phases, a companion site at aieng-zh.cn, and CI checks on lesson counts. Here is what it is good for and where it stops.
Who is it for?
Adopt this if you read Chinese and want the math-to-agent path rather than API snippets, starting from the clone command in the README. Skip it if you need English prose, an installable package, or a pinned release: there are no retrieved releases and the license is MIT, so redistribution is permitted but the translated text is a derivative that must keep the upstream MIT notice.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 3 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What ai-engineering-from-scratch-zh actually is

This repository is a Simplified Chinese derivative of AI Engineering from Scratch, credited in the README to Rohit Ghumare under the MIT license. The headline numbers in the README are 523 lessons, 20 phases, and roughly 348 hours, with code in Python, TypeScript, Rust, and Julia. The description on the repository page says 503 lessons and 20 phases, so the two counts disagree and the README is the one with the badge and the CI check behind it.

The README frames the audience with a statistic: 84% of students already use AI tools, but only 18% feel they can use them well in professional settings. Whether or not you accept that framing, the implied reader is someone who can already program, wants to understand what happens inside a transformer or an agent loop, and is willing to write the algorithm before importing the library. The prerequisites section states two things: you can code in any language (Python preferred), and you want to know how AI works rather than only calling APIs.

It is not a package. There is no pip install for the course itself. The repository holds lesson folders, a site build, agent skills, and a requirements.txt for the lesson code.

The build-it-then-use-it loop behind 523 lessons

Every lesson follows the same six beats, shown in the README as a mermaid diagram: core idea, problem context, core concepts, build it, use it, then take it away. The load-bearing split is build it versus use it. You implement the algorithm from raw math first, then run the same task with PyTorch or scikit-learn. The stated payoff is that when PyTorch appears, you already know what it is doing underneath.

The phase graph is a dependency chain, not a menu. Math foundations feed machine learning, which feeds deep learning core, which branches into computer vision, NLP, speech, and reinforcement learning. Transformer sits above NLP, and from transformer you get generative AI and LLM-from-scratch, then LLM engineering, then tools and protocols, then agent engineering, then autonomous systems, then multi-agent, with infrastructure and ethics feeding a capstone. The README is blunt about skipping: jump ahead if you already know a lower layer, but do not jump and then wonder why the upper layer collapses.

The lesson folder layout is uniform: code/ for runnable implementations, docs/zh.md for the Chinese lesson text, and outputs/ for the artifact that lesson produces. Those artifacts are prompts, skills, agents, or MCP servers. That is the concrete difference from a video course: the end state of a lesson is a file you can drop into a workflow, not a certificate.

Installing and running your first lesson

The README gives three entry paths. Reading needs no setup: open aieng-zh.cn or expand a phase in the table of contents. Cloning is path B, and it is the one that matters if you want to run the code. The README gives this exact sequence, and the second command runs a vectors script from phase 1.

bash
git clone https://github.com/fancyboi999/ai-engineering-from-scratch-zh.git
cd ai-engineering-from-scratch-zh
python phases/01-math-foundations/01-linear-algebra-intuition/code/vectors.py

If that script imports numpy or matplotlib, install the dependencies first. The repository ships a requirements.txt at the top level with version floors rather than pins.

bash
pip install -r requirements.txt

That file pulls in numpy, matplotlib, jupyter, torch, torchvision, torchaudio, transformers, datasets, tokenizers, accelerate, scikit-learn, pandas, pillow, librosa, soundfile, tiktoken, anthropic, and openai, each with a minimum version. Installing all of it is a heavy download because torch and torchaudio are in the same list as the phase 1 math code. For an early phase you can install numpy and matplotlib alone and leave the rest until you reach the deep learning phases.

The third path is the one the README recommends: a level test. Inside an agent that supports the bundled skills, such as Claude, Cursor, Codex, OpenClaw, or Hermes, you run a slash command.

bash
/find-your-level

It asks ten questions, maps your answers to a starting phase, and produces a path with hour estimates. After finishing a phase you can quiz yourself.

bash
/check-understanding 3

The README shows that this is followed by listing the outputs folder for phase 3, lesson 5, which contains prompt-loss-function-selector.md and prompt-loss-debugger.md. That is the shape of the promise: a quiz, then two prompt files you keep.

What the Chinese edition adds beyond translation

The README is explicit that this is not a machine-translated mirror. On top of the translation it localizes 523 lesson bodies, an 83-entry glossary, quiz questions, mermaid diagrams, and interactive chart labels into Simplified Chinese, while keeping technical terms such as agent, token, and transformer in English by convention. That last decision is the right one. Translating token or transformer into Chinese coinages would make the text harder to line up with papers, library documentation, and error messages.

The companion site at aieng-zh.cn is a separate deliverable, not a GitHub Pages dump. The README lists a searchable course catalog, learning paths, progress tracking, draggable interactive charts, a command palette on Cmd/Ctrl+K, and dark mode. There is also a set of animated explainer videos in the style of 3Blue1Brown with Chinese narration, embedded in lesson pages. Phase 1, the 22 math lessons, is live, and the README says the remaining phases are in production. Read that carefully: the video layer covers one phase out of twenty today, and the README itself calls the videos a supplement to the hands-on derivation rather than a shortcut.

Two build-level details are worth noting. The site build generates sitemap.xml, llms.txt, and structured data at build time, which is a deliberate bet on being cited by search engines and AI assistants. And a CI guard runs node site/build.js --check to compare the lesson count in the list against what is actually on disk. Lesson-count drift is a real failure mode in a 523-lesson course, and putting a check in CI is a cheaper fix than a manual audit.

Where this course is the wrong tool

The language boundary is the first limit. The lesson text is docs/zh.md. If you cannot read Chinese comfortably, the repository is a directory tree of code you can run with explanations you cannot read, and the upstream English course is the correct destination instead.

Second, the translation trails the source. The README states that the course structure and code stay consistent with upstream and that translations keep following upstream updates. Following means a lag, not a guarantee. A lesson that was rewritten upstream last week may still show the previous Chinese text, and the repository does not document a per-lesson sync status or a rollback path for a bad translation. If you need the newest version of a lesson, check upstream before trusting the local zh.md.

Third, the dependency list is a floor, not a lock. Every entry uses >=, so a fresh install months from now can resolve to newer major versions of torch, transformers, or tokenizers than the lesson was written against. There is no lockfile, so reproducing an exact environment is on you.

Fourth, the video coverage is one phase. Anyone arriving for animated explanations of agents or MCP servers will not find them yet, because the README says Phase 1 is live and the rest are being produced. And there is no release: the repository page returns no releases, so there is no versioned snapshot to pin your study to. The last push was on 2026-09-08, which is recent enough that the repository is not dormant, but a commit stream is not the same thing as a stable edition.

How it differs from fast.ai and the original English course

The closest comparison is the upstream AI Engineering from Scratch repository by Rohit Ghumare, also MIT. The difference is entirely in the reading experience. Upstream gives you the English lesson text; this fork gives you Chinese lesson text plus a Chinese site with progress tracking and a command palette, plus the localization of diagrams and quizzes. The code and structure are meant to match upstream, so if you read English, the fork adds nothing except the site features and the video layer.

Against fast.ai, the split is pedagogical order. fast.ai is top-down: you get a working model in the first lesson and learn the internals later. This course is bottom-up by construction. The phase graph starts at math foundations and does not reach transformer until phase 7, with LLM-from-scratch at phase 10 and agent engineering at phase 14. That ordering is the entire argument of the project, and it is also its cost: you cannot get to an agent in a weekend, because the path to phase 14 runs through linear algebra, backpropagation, and attention written by hand.

A third contrast is with API-first tutorials. Those teach you to call a model and stop. Here the README's promise is that you build the smaller version first so the framework stops being opaque. That is a defensible position, but it is slower, and it assumes you want the internals rather than a working prototype.

Licence, maintenance, and the cost of keeping up

The licence is MIT, stated in the README badge and present as a LICENSE file at the top level. MIT is permissive: you can reuse, modify, and redistribute, including commercially, provided the copyright notice and permission notice travel with the copies. For a derivative translation, that means the upstream MIT notice has to be preserved alongside the translation, which the README already does by crediting the original author and linking the upstream repository. This is a description of the licence terms, not legal advice; if you plan to republish the lessons inside a paid product, read the LICENSE file and the upstream one yourself.

Upgrade cost is the interesting part. There is no release channel and no lockfile, so there is no upgrade in the usual sense. You pull, and you get whatever the last commit contains. The requirements.txt floors mean your environment drifts independently of the repository. If you are teaching from this material, the practical move is to freeze your own environment after the first successful install rather than reinstalling from requirements.txt each term.

Maintenance signals: the last push was on 2026-09-08, and the repository is not archived. The CI lesson-count check and the CHANGELOG.md file suggest a maintained translation process, but the README does not document how a lesson is marked as synced with upstream, so you cannot tell from the repository alone which lessons are current.

Editorial conclusion

Adopt this if you read Chinese and want the math-to-agent path rather than API snippets, starting from the clone command in the README. Skip it if you need English prose, an installable package, or a pinned release: there are no retrieved releases and the license is MIT, so redistribution is permitted but the translated text is a derivative that must keep the upstream MIT notice. Verify first that the lesson you want has a docs/zh.md file, and check the upstream repository for the same lesson, because the README states the translation follows upstream updates rather than leading them.

Frequently asked questions

Can I learn AI engineering from scratch with ai-engineering-from-scratch-zh?

Yes, that is the stated purpose. The README describes 523 lessons across 20 phases, starting at math foundations and ending at a capstone, with each lesson requiring you to implement the algorithm before using a production library.

What exactly is AI engineering in this course's terms?

The README treats it as the full stack from linear algebra to autonomous agent clusters, with each algorithm first written out from raw math before the equivalent production library is used.

Does ai-engineering-from-scratch-zh include videos?

It includes animated explainer videos with Chinese narration embedded in lesson pages, but the README says Phase 1, the 22 math lessons, is live and the remaining phases are still being produced.

Official sources

  1. fancyboi999/ai-engineering-from-scratch-zh on GitHub
  2. Issues
  3. License: MIT
  4. Project website
  5. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/fancyboi999-ai-engineering-from-scratch-zh.svg)](https://hysenlabs.com/projects/fancyboi999-ai-engineering-from-scratch-zh)