# Agentic RAG for Dummies: a LangGraph RAG agent you can actually read

> GiovanniPasq's repository pairs a step-by-step notebook with a modular LangGraph app that adds query clarification, parallel retrieval agents and self-correction to ordinary RAG. The hard part is not the code, it is running the 8B-plus local model the documentation asks for.

**GiovanniPasq/agentic-rag-for-dummies** — A modular Agentic RAG built with LangGraph — learn Retrieval-Augmented Generation Agents in minutes.

- Repository: https://github.com/GiovanniPasq/agentic-rag-for-dummies
- Stars: 4,220 · Forks: 555
- Language: Jupyter Notebook
- License: MIT
- Published: 2026-09-09 · Updated: 2026-09-09 · Language: en
- Canonical page: https://hysenlabs.com/projects/giovannipasq-agentic-rag-for-dummies

## What Agentic RAG for Dummies actually solves

A plain RAG pipeline embeds a question, searches a vector store, and pastes the top chunks into a prompt. That works until the question is ambiguous, or contains two questions at once, or the first search returns nothing useful. At that point a fixed pipeline has no move left: it either answers from bad context or refuses.

The repository addresses that gap. Its README frames the motivation directly: most RAG tutorials "show basic concepts but lack guidance on building modular, agent-driven systems." The project supplies both halves of that sentence, a notebook that walks through the concepts and a modular app under project/ whose components can be swapped independently.

The intended reader is an engineer or student who already knows what embeddings and a vector store are, and now wants to see how a graph of agents handles the failure cases. It is not an introduction to retrieval itself, and it is not a hosted service.

## Hierarchical indexing and the four-stage query workflow

Two mechanisms carry the design. The first is hierarchical indexing during document preparation. Documents are split twice: parent chunks are bounded large sections cut on Markdown headers (H1, H2, H3), and child chunks are small fixed-size pieces derived from those parents. Search runs against the child chunks for precision, then the parent chunk is fetched to give the generator real context. That is a standard parent-document retrieval pattern, and the README is honest that the split is header-dependent, which means non-Markdown source material needs conversion first.

The second mechanism is the query workflow, which the README lays out as a chain: user query, conversation summary, query rewriting, query clarification, parallel agent reasoning, aggregation, final response. Stage 1 keeps a rolling summary plus recent history so context does not grow without bound. Stage 2 rewrites references ("How do I update it?" becomes "How do I update SQL?"), splits multi-part questions, and pauses for human input when the query is genuinely unclear. Stage 3 spawns one agent subgraph per sub-query; the README's example is a two-part question about JavaScript and Python producing two parallel agents. Each agent searches child chunks, fetches parents, self-corrects when results look insufficient, compresses context to avoid redundant fetches, and falls back when its search budget runs out. Stage 4 merges the agent responses into one answer.

The graph is the orchestration layer, and LangGraph supplies the state and the interruption point for the clarification step. Observability is delegated to Langfuse and evaluation to RAGAS, both listed in requirements.txt.

## Installing it and running a first query

The repository is a Python project. requirements.txt pins Python 3.11 or newer alongside langgraph==1.2.11, langchain-ollama==1.1.0, langchain-qdrant==1.1.0, langfuse==4.15.1, ragas==0.4.3, pymupdf4llm==1.28.2 and gradio==6.26.0. Install them into a virtual environment:

```bash
python -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt
```

The runnable app is Ollama-first, so the local model comes before anything else. The README gives this pair of commands, and the model name is the one it uses:

```bash
# Install Ollama from https://ollama.com
ollama pull granite4.1:8b
```

The chat model is then constructed with LangChain's Ollama integration. Temperature 0 and a fixed seed are what the README shows, which makes repeated runs comparable:

```python
from langchain_ollama import ChatOllama

llm = ChatOllama(model="granite4.1:8b", temperature=0, seed=42)
```

If you would rather not install anything, the README links a Colab notebook at notebooks/agentic_rag.ipynb, and the badge points at the main branch. That is the fastest way to see the graph run before committing to a local environment. For the full app, the README points to the Modular Architecture and Installation & Usage sections for the remaining steps; it does not spell out a single start command in the excerpt available here, so read those sections in the repository before assuming a one-liner exists.

## Where the design runs into trouble

The README itself flags the sharpest limitation, in a warning next to the Ollama example: for reliable tool calling and instruction following, prefer models 8B+ because smaller models "may ignore retrieval instructions or hallucinate." An agentic RAG graph depends on the model deciding when to search, when to rewrite and when to stop. A model that cannot follow those instructions turns the agent loop into an expensive way to produce a confident wrong answer. Anyone planning to run this on a 3B model on a laptop should treat that warning as the project's own answer.

The second constraint is architectural. Parallel sub-query agents multiply the number of model calls per user question, and each agent may self-correct and re-query. On a local Ollama instance that cost lands on your own hardware, and the README does not publish latency or throughput figures, so there is no number to plan against.

The third is operational. The README does not document rollback, and it does not describe a migration path for the vector store when the chunking strategy changes. Re-chunking documents means rebuilding the index; nothing in the repository suggests an incremental path. The README also warns that model names change frequently and tells you to check official documentation for current identifiers, which is a maintenance burden the project pushes onto you rather than solving.

## How it differs from a framework-agnostic RAG stack

The obvious alternative is a plain RAG chain built directly on LangChain retrievers, or a pipeline assembled from an orchestration-free stack such as a Qdrant client plus a prompt template. The difference is not the components, since both use an embedding model and a vector store. The difference is control flow.

A retriever chain is a straight line: query in, documents out, prompt, answer. There is no state machine deciding that the first search was weak, no branch that asks the user a question, and no fan-out into parallel sub-queries. Agentic RAG for Dummies puts LangGraph in the middle precisely so those branches exist as graph edges rather than as if-statements buried in a function. If your questions are single-topic and your corpus is clean, the straight line is cheaper and easier to debug, and the agent graph is the wrong tool.

A second alternative is a heavier agent framework where the retrieval loop is configured rather than written. Those give you less visibility into the state transitions. This repository's value is that the graph is small enough to read, which is the stated point of the project.

## Licence and the cost of staying current

The licence is MIT, stated in the README badge and in the LICENSE file at the repository root. MIT permits commercial use, modification and redistribution provided the copyright notice and permission notice are kept. That is the general shape of the licence; whether it fits your product is a question for your own legal review, not something this article can settle.

Upgrade cost is the more practical concern. requirements.txt pins exact versions, including langgraph==1.2.11 and langchain-qdrant==1.1.0. LangGraph and the LangChain integration packages move quickly, and the README concedes that model names change frequently enough that you must check official documentation before deploying. Every bump of LangGraph is a possible change to graph state or interruption semantics, and the repository publishes no compatibility matrix. The releases page shows v2.3 on 2026-06-21, v2.2 on 2026-06-10 and v2.1 on 2026-04-01. The last push to main was on 2026-08-30, which is recent, but the README does not document a deprecation policy or a supported-version window, so pinning is your only real protection.

## Conclusion

Adopt it if you already understand basic RAG and want a working reference for LangGraph state, parallel sub-query agents and human-in-the-loop clarification, or if you are teaching those ideas and need a notebook students can run in Colab. Do not adopt it as production retrieval infrastructure: the runnable app is Ollama-first, the README does not document rollback, and the repository states that small local models ignore retrieval instructions. Before committing, verify three things on your own machine: that your chosen Ollama model handles tool calling without hallucinating, that the Qdrant instance in the project is reachable from wherever you deploy, and that the pinned LangGraph 1.2.11 and langchain-qdrant 1.1.0 versions satisfy your existing dependency set. The last push to main was on 2026-08-30, so check the commit history rather than assuming the README describes the current code.

## FAQ

### Can you explain agentic RAG in a simple way?

In this project, agentic RAG means the retrieval step is decided by a model inside a graph rather than by a fixed pipeline. The README describes four stages: conversation summary, query rewriting and clarification, parallel agent reasoning over sub-queries, then aggregation into one answer. Each agent can self-correct and re-query when its first search is insufficient.

### Can you explain RAG to a beginner?

The repository's document preparation step splits files twice: parent chunks cut on Markdown headers, and small child chunks derived from them. Search runs on the child chunks for precision, and the parent chunk is retrieved to supply context for the answer. The README calls this hierarchical indexing.

### What does "agentic" mean in simple terms for Agentic RAG for Dummies?

The README uses the term for a LangGraph workflow where the system rewrites the query, pauses to ask the user when the question is unclear, spawns one parallel agent per sub-query, and re-queries when results are insufficient. Those decisions are graph steps, not fixed code paths.

### Is ChatGPT an agentic AI?

The repository does not discuss ChatGPT or compare it with this project, so this question cannot be answered from what it documents. What the README does describe is a local, Ollama-first app whose chat model can be swapped for any provider supported by LangChain.

## Sources

- [GiovanniPasq/agentic-rag-for-dummies on GitHub](https://github.com/GiovanniPasq/agentic-rag-for-dummies)
- [Issues](https://github.com/GiovanniPasq/agentic-rag-for-dummies/issues)
- [License: MIT](https://github.com/GiovanniPasq/agentic-rag-for-dummies/blob/main/LICENSE)
- [README](https://github.com/GiovanniPasq/agentic-rag-for-dummies/blob/main/README.md)
- [Releases](https://github.com/GiovanniPasq/agentic-rag-for-dummies/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/giovannipasq-agentic-rag-for-dummies
