All-in-RAG: A Structured Chinese Tutorial for Building Production RAG Systems
🔍大模型应用开发实战一:RAG 技术全栈指南,在线阅读地址:https://datawhalechina.github.io/all-in-rag/
At a glance
- What is it?
- All-in-RAG is an open-source Python tutorial repository from Datawhale that covers the full RAG stack from data loading to system evaluation, organized into ten chapters with code examples, and licensed under CC BY-NC-SA 4.0.
- Who is it for?
- All-in-RAG suits Python developers with basic language skills who want a structured path through the full RAG stack, from document loading to Graph RAG, rather than assembling fragments from blog posts. It is the wrong choice for teams that need production-ready code they can drop into a system unchanged; the repository is a tutorial with illustrative examples, not a reusable library.
- Can I use it commercially?
- Not without permission. GitHub finds no licence file in the repository, and without a licence all rights are reserved by default: you may read the code but not reuse it. Check the README, or ask the authors, before using it.
- Is it still maintained?
- Yes. The repository last received commits 27 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 17, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What All-in-RAG Is and What Problem It Addresses
All-in-RAG is a tutorial repository for developers who want to learn Retrieval-Augmented Generation from the ground up. RAG is a pattern where a language model retrieves relevant documents from an external store before generating a response, which allows the model to answer questions about data that was not in its training set. The README's stated motivation is that existing RAG tutorials are scattered and unsystematic, leaving beginners without a complete technical picture. This repository attempts to fill that gap with a ten-chapter curriculum covering theory, code, and multi-modal extensions. The target audience, as the README specifies, is Python developers with basic language skills who have some familiarity with Docker and basic Linux command-line usage. Researchers and AI engineers building intelligent question-answering systems are also named as the intended readers.
How the Repository Is Organized
The top-level directory layout separates content by function:
all-in-rag/
├── docs/ # 教程文档
├── code/ # 代码示例
├── data/ # 示例数据
├── models/ # 预训练模型
├── Extra-chapter/ # 扩展章节与社区实践内容
└── README.md # 项目说明The docs/ directory holds the tutorial text, organized by chapter number. The code/ directory holds the Python code examples that accompany each chapter. The data/ directory holds sample data for the exercises. The models/ directory holds pretrained models used in examples. The Extra-chapter/ directory contains community-contributed content, including topics like a Neo4j application and a multimodal embedding practice using Jina v5-omni, which are separate from the main curriculum and held to a different contribution standard. Each chapter in docs/ is a Markdown file with a corresponding code directory.
The Ten-Chapter Curriculum Path
The ten chapters follow a progression from fundamentals to advanced architecture. Part one (chapters one and two) covers the RAG concept, environment setup, and data preparation including document loading and text chunking. Part two (chapters three) covers index construction: vector embeddings, multimodal embeddings, vector databases, Milvus in particular, and index optimization. Part three (chapter four) covers retrieval techniques: hybrid search combining dense and sparse retrieval, query construction, Text2SQL, query rewriting, and advanced retrieval algorithms. Part four (chapters five and six) covers generation formatting and system evaluation with common evaluation tools. Part five (chapters seven through nine) covers advanced topics: knowledge graph RAG, and two complete project walkthroughs including one optimized with Graph RAG. Chapter ten is listed as planned but not yet published. The curriculum is primarily in Chinese, though the repository includes an English README.
Using the Repository: Cloning and Running Examples
The README does not reproduce a specific git clone command, but the repository is hosted at github.com/datawhalechina/all-in-rag. Environment setup instructions are covered in docs/chapter1/02_preparation.md, with a supplementary Python virtual environment guide at docs/chapter1/virtualenv.md contributed by a community member. Docker knowledge is listed as a prerequisite in the README, which suggests some examples use containerized dependencies such as vector databases. The README notes that the docs/ directory holds tutorial text and the code/ directory holds examples, so navigating the repository means reading a chapter in docs/ and running the corresponding code in code/. The Extra-chapter/ directory has its own contribution guide at Extra-chapter/README.md.
Milvus as the Primary Vector Database Example
Chapter three includes a dedicated section on Milvus at docs/chapter3/09_milvus.md, covering multimodal retrieval that combines text and image vectors in a single index. Chapter nine, the Graph RAG project optimization, also uses Milvus for index construction at docs/chapter9/03_index_construction.md. This means Milvus appears as the primary vector database example throughout the more advanced parts of the curriculum, with the basic vector database chapter at docs/chapter3/08_vector_db.md covering the concept more generally before the Milvus-specific exercises. Developers who want to use a different vector database such as Qdrant, Weaviate, or Chroma will need to adapt the Milvus-specific code, since the repository does not abstract the database layer. The README does not document interchangeable database support; Milvus is named as the concrete implementation in the project examples, which is a practical constraint for anyone who has already standardized on a different vector store.
Limitations as a Learning Resource
All-in-RAG is a tutorial, not a library or a framework. The code examples illustrate concepts rather than providing production-ready modules you can import into an application. Chapter ten is listed as planned but not yet released, which means the second project walkthrough in the curriculum is incomplete at the time of the last repository push. Community contributions to the Extra-chapter/ directory are accepted at a different quality standard than the main chapters, and the README notes they are evaluated based on completeness, practical depth, and reference value before being merged. The main tutorial chapters accept corrections and documentation improvements but not general content pull requests according to the README; new chapters go through the Extra-chapter contribution process instead. Text2SQL coverage is in a single chapter (docs/chapter4/13_text2sql.md), and the README does not describe which SQL dialects or databases are used in that example. Similarly, the hybrid retrieval chapter combines dense and sparse retrieval but the README does not name specific sparse retrieval implementations.
License and Maintenance
All-in-RAG is licensed under Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International (CC BY-NC-SA 4.0). The NonCommercial clause prohibits using the content for commercial purposes, which includes incorporating it into paid training courses or commercial documentation products. The ShareAlike clause requires derivative works to use the same license, meaning any adapted tutorial must also be released under CC BY-NC-SA 4.0 and may not be used commercially. The last push to the repository was on 2026-09-04. The project is maintained by the Datawhale community, with the primary contributor listed as the project lead at github.com/FutureUnreal. The Extra-chapter/ directory accepts community contributions on an ongoing basis, and the README lists specific evaluation criteria that determine whether a submitted chapter is merged into the repository.
Editorial conclusion
All-in-RAG suits Python developers with basic language skills who want a structured path through the full RAG stack, from document loading to Graph RAG, rather than assembling fragments from blog posts. It is the wrong choice for teams that need production-ready code they can drop into a system unchanged; the repository is a tutorial with illustrative examples, not a reusable library. Developers who use a different vector database than Milvus will need to adapt the project examples, since the repository does not abstract that layer. The CC BY-NC-SA 4.0 license prohibits commercial use of the content, so companies building internal training materials from it need to confirm that restriction covers their use case before distributing anything derived from it. The last repository push was on 2026-09-04.
Frequently asked questions
What does RAG actually mean in the context of this project?
RAG stands for Retrieval-Augmented Generation. All-in-RAG is a tutorial covering the pattern where a language model retrieves relevant documents from an external store before generating a response, which allows it to answer questions about data outside its training set.
What programming language and prerequisites does All-in-RAG require?
The README specifies Python as the required language, along with basic familiarity with Docker and Linux command-line operations. Some knowledge of large language model concepts is recommended but listed as non-mandatory.
Can All-in-RAG content be used commercially?
No. The repository is licensed under CC BY-NC-SA 4.0, which prohibits commercial use of the content. The ShareAlike term also requires any derivative works to be released under the same license.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/datawhalechina-all-in-rag)