Model or dataset
NirDiamant/Controllable-RAG-Agent avatar
NirDiamant/Controllable-RAG-Agent

Controllable-RAG-Agent: Graph-Driven Answers for Complex Questions Over Private Data

This repository provides an advanced Retrieval-Augmented Generation (RAG) solution for complex question answering. It uses sophisticated graph based algorithm to handle the tasks.

1,625 stars268 forksJupyter NotebookApache-2.0

At a glance

What is it?
Controllable-RAG-Agent is an open-source Python project that replaces simple vector similarity retrieval with a deterministic graph orchestrator, allowing it to answer multi-hop questions from private document collections where semantic search alone returns the wrong chunks.
Who is it for?
Controllable-RAG-Agent suits teams with a PDF corpus and multi-hop questions that simple vector retrieval answers incorrectly. It is not suitable for production deployments without significant refactoring: the Jupyter notebook structure, the pinned LangGraph 0.0.49 dependency, and the absence of a stable API surface mean it works best as a research prototype or a starting point for a more hardened system.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 8 days ago.
What is it written in?
Mainly Jupyter Notebook, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 25, 2026, and from our analysis. They are not legal advice.

Editorial analysis

Why Semantic Search Alone Fails Complex Questions

Standard RAG pipelines split a document into chunks, embed each chunk, store the embeddings, and then retrieve the top-k nearest neighbors for any incoming question. That approach works well for questions with a single, localized answer. It breaks down when the answer requires combining information from multiple parts of the document: identifying which character appears in chapter three and then verifying a claim made about that character in chapter twelve, for example.

Controllable-RAG-Agent addresses this by using a deterministic graph as the reasoning layer. The graph decides which retrieval step to take next, keeps track of intermediate answers, and adjusts its plan based on what the previous step returned. The project description names the graph as the "brain" of a highly controllable autonomous agent. The controllability refers to the fact that the execution path through the graph is determined by explicit logic, not by a second LLM call that might behave unpredictably. The target audience is developers and researchers who need reliable, auditable answers from their own document collections and who have found that simpler RAG pipelines miss compound questions.

Pipeline Architecture: From PDF Loading to Graph Orchestration

The pipeline has five named phases. PDF documents are loaded and split into chapters. Text preprocessing then cleans and normalizes the content for better summarization and embedding. In the summarization phase, large language models generate extensive summaries of each chapter and store them in a dedicated chapter summaries vector store. A separate book quotes database is built for questions that require access to specific verbatim passages. The final phase builds the vector store over document chunks used for general retrieval.

The repository reflects this structure directly in its top-level directories: `chapter_summaries_vector_store/`, `chunks_vector_store/`, and `book_quotes_vectorstore/` each hold a separate index. The primary notebook, `sophisticated_rag_agent_harry_potter.ipynb`, uses a Harry Potter text as the example corpus, which gives a concrete demonstration of multi-hop reasoning over fiction.

The graph itself is built with LangGraph 0.0.49. LangGraph at that version models the agent as a state machine: each node in the graph is a function that reads the current state, does its work, and returns an updated state. The graph can route to different nodes depending on what a retrieval step returned, which allows it to formulate follow-up queries without returning to the user.

Hallucination prevention is enforced at the retrieval level: the graph is constrained to answer only from the retrieved content. It does not fall back to the LLM's pretrained knowledge when no relevant chunk is found. The README lists `Ragas` metrics for evaluating answer quality after the fact.

Running the Agent: Docker Setup and Required API Keys

The project ships with a Docker Compose file that wraps the application in a Streamlit container exposed on port 8501.

First, copy the environment template and fill in your API keys:

bash
cp .env.example .env

The `.env.example` file defines two variables that must be set:

bash
OPENAI_API_KEY=
GROQ_API_KEY=

With the keys in place, start the application:

bash
docker compose up

This builds the image defined in `dockerfile`, installs the dependencies from `requirements.txt`, and starts the Streamlit server on port 8501. For local development outside Docker, install the dependencies directly:

bash
pip install -r requirements.txt

The requirements list pins `langgraph==0.0.49`. This is version 0.0.49 of LangGraph, released before the 0.1 API restructuring. If you have a LangChain or LangGraph installation at a later version in the same environment, the pinned dependency may conflict. The `simulate_agent.py` script provides a command-line entry point for running the agent without the Streamlit UI.

Hallucination Prevention and Ragas Evaluation

The project's primary defense against hallucinated answers is architectural. Because the graph retrieves content from the private document corpus before generating an answer, and because it is structured to use only that retrieved content when responding, the LLM cannot easily interpolate from pretrained knowledge.

This constraint is harder to enforce than it sounds. LLMs will often blend retrieved content with pretrained knowledge in ways that are difficult to detect from the output alone. The README lists Ragas metrics as the evaluation mechanism: Ragas is a library that measures retrieval-augmented generation quality by scoring faithfulness, answer relevancy, and context precision and recall.

The `functions_for_pipeline.py` and `helper_functions.py` files at the top level of the repository contain the supporting functions for the graph nodes. The `graphs/` directory holds the actual graph definitions. The `full_graph_visualization.ipynb` notebook allows you to render the graph structure, which is useful for understanding the execution flow before running the agent on real data.

Limitations: Notebook-First Design and Version Pins

The primary deliverable is a Jupyter notebook. The code is organized to demonstrate the approach rather than to serve as a library with a stable API. Calling the agent programmatically from another Python module requires reading through the notebook and extracting the relevant functions, since there is no installed package with a public API surface.

The langgraph==0.0.49 pin is a significant constraint. LangGraph 0.1 introduced breaking changes to the graph construction API. Code written for 0.0.49 will not run against a current LangGraph installation without modification. For anyone maintaining a LangChain-based stack at a recent version, this means creating an isolated virtual environment for this project.

The project also depends on FAISS for vector storage. FAISS does not persist to disk automatically in all configurations, and the repository includes pre-built vector stores (`chapter_summaries_vector_store/`, `chunks_vector_store/`, `book_quotes_vectorstore/`) that appear to be built from the Harry Potter example corpus. Using the agent on a different document corpus requires rebuilding all three stores, which is time-consuming for long documents.

Controllable-RAG vs. Simple Vector-Store Pipelines

A standard LangChain retrieval chain does one thing: embed a question, retrieve the top-k chunks by cosine similarity, and pass those chunks to an LLM. This is fast, easy to set up, and works well for single-hop questions. It has no concept of planning, sub-task decomposition, or iterative query refinement.

Controllable-RAG-Agent introduces a planning layer on top of retrieval. The graph can issue multiple retrieval steps, evaluate intermediate results, and adjust subsequent queries based on what was found. This makes it meaningfully more capable for compound questions, but also meaningfully more complex to debug when the graph takes an unexpected path.

The relevant comparison for simpler use cases is a plain LangChain conversational retrieval chain or a retrieval chain built with LlamaIndex. Both are easier to set up and deploy. Neither supports the multi-step graph-based planning that Controllable-RAG-Agent provides. The trade-off is clear: simpler pipelines for single-hop questions, graph-based orchestration for compound questions over large corpora.

License, Companion Resources, and Activity Status

The project is licensed under Apache 2.0. The README links to two companion repositories by the same author: RAG_Techniques, which documents additional retrieval strategies, and GenAI_Agents, which covers general agent patterns. These are separate repositories and are not dependencies of Controllable-RAG-Agent itself.

The README also promotes a paid book, "RAG Made Simple," and a course, "Prompt to Production." These are commercial products by the author. They are not required to use the repository, but the README dedicates significant space to them.

The last push to the repository was on 2026-09-21. The project does not use GitHub releases; contributions and updates appear directly on the main branch.

Editorial conclusion

Controllable-RAG-Agent suits teams with a PDF corpus and multi-hop questions that simple vector retrieval answers incorrectly. It is not suitable for production deployments without significant refactoring: the Jupyter notebook structure, the pinned LangGraph 0.0.49 dependency, and the absence of a stable API surface mean it works best as a research prototype or a starting point for a more hardened system. Before adopting it, check that langgraph==0.0.49 is compatible with your existing LangChain installation, since that version predates the API changes introduced in LangGraph 0.1 and later.

Frequently asked questions

What is an agentic RAG agent?

An agentic RAG agent combines retrieval-augmented generation with a planning or orchestration layer that can issue multiple retrieval steps, evaluate intermediate results, and adjust its query strategy before producing a final answer. Controllable-RAG-Agent implements this using a deterministic LangGraph state machine as the orchestrator.

What is the difference between a RAG and an agent?

A standard RAG pipeline retrieves relevant content and passes it to an LLM in a single step. An agent adds a planning or decision layer that can take multiple actions, such as issuing follow-up queries, before producing an answer. Controllable-RAG-Agent uses a LangGraph graph as the agent layer to handle questions that require combining information from multiple parts of a document.

What LLM providers does Controllable-RAG-Agent support?

The .env.example file defines OPENAI_API_KEY and GROQ_API_KEY, and the requirements.txt includes both the openai and groq packages. The README does not document support for other providers.

Official sources

  1. Issues
  2. License: Apache-2.0
  3. NirDiamant/Controllable-RAG-Agent on GitHub
  4. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/nirdiamant-controllable-rag-agent.svg)](https://hysenlabs.com/projects/nirdiamant-controllable-rag-agent)