Model or dataset
huangjia2019/rag-in-action avatar
huangjia2019/rag-in-action

rag-in-action: A Modular Jupyter Course for End-to-End RAG System Development

End-to-end RAG system design, evaluation, and optimization. 极客时间RAG训练营,RAG 10大组件全面拆解,4个实操项目吃透 RAG 全流程。RAG的落地,往往是面向业务做RAG,而不是反过来面向RAG做业务。这就是为什么我们需要针对不同场景、不同问题做针对性的调整、优化和定制化。魔鬼全在细节中,我们深入进去探究。

836 stars324 forksJupyter NotebookLicense varies

At a glance

What is it?
rag-in-action is a Geekbang course code repository that walks through building a retrieval-augmented generation system in 11 modules, from a minimal SimpleRAG to GraphRAG and multi-agent patterns. It supports both LangChain and LlamaIndex, with separate environment configuration files for each, and runs on CPU or GPU depending on what you have available.
Who is it for?
Engineers who want a structured, code-first walkthrough of the full RAG pipeline, from document loading and chunking through retrieval, re-ranking, and evaluation, will find this repository directly usable. The course material is designed around a paid Geekbang class, so the repository works best when combined with the course videos.
Can I use it commercially?
Not without permission. GitHub finds no licence file in the repository, and without a licence all rights are reserved by default: you may read the code but not reuse it. Check the README, or ask the authors, before using it.
Is it still maintained?
Yes. The repository last received commits 36 days ago.
What is it written in?
Mainly Jupyter Notebook, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 28, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What rag-in-action is and who it targets

RAG systems fail in practice not because the retrieval or generation components are individually broken, but because the configuration connecting them to a specific business problem is wrong. The repository's description makes this explicit: the approach is to build RAG around the actual business need, rather than adapting a business need to fit a generic RAG template.

rag-in-action is the companion code repository for a Geekbang training course on RAG system development. The course homepage is at https://u.geekbang.org/subject/airag/1009927. A companion book titled RAG In Action was published by 人民邮电出版社 and is available separately.

The repository is aimed at developers with Python experience who want to understand every layer of a production RAG pipeline. It is not a library or a deployable application; it is a learning resource where each module corresponds to a phase of the pipeline and includes working Jupyter notebooks with examples.

Eleven modules from data loading to advanced RAG patterns

The repository is organized into 11 numbered directories plus supporting folders for environment configuration, documentation, images, and English-language content:

Module 00 (SimpleRAG) gives a minimal working RAG system as a starting point. Module 01 (DataLoading) covers data ingestion with pandas and PyPDF2. Module 02 (DocChunking) works through document splitting strategies using LangChain Splitters. Module 03 (Embedding) handles text vectorization with HuggingFace models and BGE. Module 04 (VectorDB) covers operations on Milvus and Chroma vector databases. Module 05 (PreRetrieval) addresses query expansion and other retrieval optimization techniques. Module 06 (Indexing) works with hierarchical and keyword indexes. Module 07 (PostRetrieval) covers re-ranking and filtering of retrieval results. Module 08 (Generation) integrates LLMs for answer generation. Module 09 (Evaluation) applies RAGAS and TruLens for measuring system performance. Module 10 (AdvanceRAG) covers GraphRAG and multi-agent patterns.

Two additional directories handle hybrid retrieval (Module 12, dual-path recall) and knowledge sources beyond vector databases (Module 11, three-path knowledge base). The 90-docs and 91-environment directories hold sample data and requirements files.

Setting up the Python environment for LangChain or LlamaIndex

The repository provides separate requirements files for LangChain and LlamaIndex, each in GPU and CPU variants. For a LangChain environment on Ubuntu with GPU:

bash
python -m venv venv-rag-langchain
source venv-rag-langchain/bin/activate
pip install -r 91-环境-Environment/requirements_langchain_20250413_Ubuntu-with-GPU.txt

For a CPU-only LangChain environment on macOS or Windows:

bash
python -m venv venv-rag-langchain
source venv-rag-langchain/bin/activate
pip install -r 91-环境-Environment/requirements_langchain_无GPU版_Mac-Win.txt

The LlamaIndex path uses a separate virtual environment and a different requirements file:

bash
python -m venv venv-rag-llamaindex
source venv-rag-llamaindex/bin/activate
pip install -r 91-环境-Environment/requirements_llamaindex_Ubuntu-with-CPU.txt

PDF processing requires additional system packages. On Ubuntu, the README lists ghostscript and python3-tk. On macOS, ghostscript and tcl-tk are available through Homebrew. Windows requires a manual Ghostscript installation. Python 3.10 is specified as the minimum version.

GPU versus CPU configuration and hardware trade-offs

The README gives two hardware profiles. The GPU configuration specifies an NVIDIA GPU with at least 8GB of VRAM, CUDA 11.8 or higher, and cuDNN 8.0 or higher. Ubuntu 22.04 LTS is the recommended operating system for the GPU path.

The CPU configuration requires at least 16GB of RAM and a processor with at least four cores. macOS (both Intel and Apple Silicon) and Windows 10 or 11 through WSL2 are supported in the CPU path. Apple Silicon can use MPS acceleration for PyTorch operations, which the README notes is available but does not describe in further detail.

The difference matters most in Module 03 (Embedding) and Module 08 (Generation). Embedding models like BGE run significantly faster on GPU when processing large document sets. The LLM integration in Module 08 is routed through the DeepSeek API by default, so local GPU resources are not required for generation unless you configure a locally hosted model.

Evaluation with RAGAS and TruLens in Module 09

Module 09 (Evaluation) is the section that most applied teams skip and later regret. The repository includes examples using RAGAS, which measures answer faithfulness, answer relevance, and context precision against a question-answer dataset. TruLens is included as an alternative framework that adds observability and logging to the evaluation pipeline.

Running evaluation well requires a ground-truth dataset, which the repository provides through the 90-docs folder with sample data. The evaluation workflows in Module 09 assume a working retrieval and generation setup from earlier modules, so the recommended approach in the README is to work through modules in order before running evaluation.

A properly configured evaluation loop catches the most common failure mode in RAG: retrieval returning context that is correct but not useful for the actual query. The chunking and retrieval configuration decisions in Modules 02, 05, 06, and 07 directly affect evaluation scores, and the repository structure makes it possible to modify and re-evaluate without rebuilding the entire pipeline.

Limitations: Geekbang paywall, Chinese-first content, and lack of automated tests

The repository is designed as supplementary material for a paid course. The code works independently, but the motivation behind specific design decisions in each module is explained in course videos that require a Geekbang subscription to access. Engineers who encounter confusing choices in the notebooks may not find explanations within the repository itself.

The README and most notebook comments are written in Chinese. The repository includes a 99-EN directory for English content, but the main documentation is not available in English within the repository. Technical terms, file paths, and API names remain in their original forms, so the code itself is usable regardless of language, but reading the comments requires Chinese.

The repository has no automated test suite. There are no CI workflows visible in the top-level structure. Changes to individual notebooks can introduce inconsistencies without a testing mechanism to catch them. The last push was on 2026-08-25, and the repository is not archived. The README states MIT as the license, though the metadata does not identify a formal license file.

How rag-in-action compares to LangChain and LlamaIndex documentation

The official documentation for LangChain and LlamaIndex covers individual components in isolation: how to create a retriever, how to call an LLM, how to build a chain. These docs are built around showcasing each framework's API surface rather than teaching how to combine components for a real use case.

rag-in-action takes a different path. Each module builds on the previous one, and the progression from SimpleRAG through Evaluation to AdvancedRAG reflects the actual sequence a team would follow when moving a RAG prototype to production. The evaluation module in particular addresses something neither LangChain nor LlamaIndex documentation covers in depth: measuring whether the complete system actually works for the target problem.

The cost is that rag-in-action is tied to a specific course rather than being an independent reference. Updates happen when the course material changes, not on a public versioning schedule.

Editorial conclusion

Engineers who want a structured, code-first walkthrough of the full RAG pipeline, from document loading and chunking through retrieval, re-ranking, and evaluation, will find this repository directly usable. The course material is designed around a paid Geekbang class, so the repository works best when combined with the course videos. Readers who want a completely self-contained English resource should be aware that the primary content and README are in Chinese, and the included notebooks assume access to a DeepSeek API key.

Frequently asked questions

What LLM frameworks does rag-in-action support?

The repository provides requirements files for both LangChain and LlamaIndex, with GPU and CPU variants for each. The virtual environments are kept separate, so both frameworks can be installed and used without conflicts.

Does rag-in-action require a GPU to run the notebooks?

No. The README describes a CPU configuration requiring at least 16GB of RAM. GPU is optional and speeds up the embedding and local model modules. Apple Silicon can use MPS acceleration on the CPU path.

What evaluation frameworks are covered in the course repository?

Module 09 includes examples using RAGAS, which measures faithfulness and answer relevance, and TruLens, which adds observability and logging. Both are listed in the requirements files for the relevant modules.

Official sources

  1. huangjia2019/rag-in-action on GitHub
  2. Issues
  3. Project website
  4. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/huangjia2019-rag-in-action.svg)](https://hysenlabs.com/projects/huangjia2019-rag-in-action)