Model or dataset
umbertogriffo/rag-chatbot avatar
umbertogriffo/rag-chatbot

umbertogriffo/rag-chatbot: a local RAG chatbot over your Markdown files

RAG (Retrieval-augmented generation) ChatBot that provides answers based on contextual information extracted from a collection of Markdown files.

445 stars112 forksPythonApache-2.0

At a glance

What is it?
A Python RAG chatbot that indexes a docs folder into Chroma, runs Llama 3.1 through a local llama.cpp server, and answers with context from your own pages. It is a self-hosted project with a CUDA or Metal requirement and a real rebuild cost when you change embedding models.
Who is it for?
Adopt it if you keep documentation in Markdown, already have an NVIDIA GPU or an Apple Silicon machine, and want the retrieval and generation loop on your own hardware rather than behind an API. Do not adopt it if you need a hosted endpoint, a managed vector store, or a corpus that is not Markdown: the upload allowlist is [".md"], and the README does not document rollback or a migration path for switching embedding models.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 110 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What problem this solves, and who it is actually for

Most RAG demos assume you will call a hosted model and pay per token. This project takes the opposite position. It loads a GGUF model into a llama.cpp server that runs on your machine, embeds your documents locally with Sentence Transformers, and stores the vectors in Chroma. Nothing in the described pipeline requires an external inference API.

The intended input is narrow and deliberate: a collection of Markdown files in the docs folder. The README states that the Memory Builder loads Markdown pages from docs, splits them, embeds the sections, and saves them in Chroma. The upload allowlist in .env.example is ALLOWED_UPLOAD_EXTENSIONS=[".md"], so this is not a general document ingestion tool. If your knowledge lives in PDFs, tickets, or a wiki API, you are converting it to Markdown first.

The audience is an engineer with a workstation GPU who wants a working, inspectable RAG loop they can read end to end. The disclaimer is explicit about what was tested: Ubuntu 22.04.2 LTS on a Lenovo Legion 5 Pro with an i7-12700H and an RTX 3060, plus macOS Sonoma 14.3.1 on an M1 MacBook Pro. Other hardware is not covered, and the README points readers to the upstream llama.cpp issue tracker when models will not load. That is a fair warning, not a hedge: the hardware dependency is the first thing that will break for a new user.

The retrieval loop: question rewriting, then reading

The pipeline has a step that many small RAG projects skip. The README says that because the original question cannot always retrieve well, the system first prompts an LLM to rewrite the question, and only then performs retrieval-augmented reading. So the flow is: user question, LLM rewrite, vector search over Chroma, context assembly, final answer. That extra inference call costs latency on a local GPU, and it is the kind of trade-off worth measuring on your own hardware rather than assuming.

Chat history is persisted and fed back in. The config exposes CHAT_HISTORY_LENGTH=2, so the conversation memory is short by default. Two turns is enough to resolve pronouns in a follow-up question; it is not enough to hold a long design discussion, and the README does not describe a summarization step for older turns.

Context overflow is handled by two named strategies. Create And Refine the Context synthesizes a response sequentially through all retrieved contents. Hierarchical Summarization of Context generates an answer for each relevant section independently and then combines them hierarchically. The .env.example sets SYNTHESIS_STRATEGY="tree-summarization", and NUM_RETRIEVALS=2. Two retrievals is a low number, which keeps the context small and the answer fast, but it also means a question whose answer is spread across three sections may miss one. Both the strategy and the retrieval count are configuration, not fixed behavior.

The chunking is worth noting for anyone who has fought LangChain's splitter. The README says the authors took the RecursiveCharacterTextSplitter class from LangChain and refactored it, specifically to avoid adding LangChain as a dependency. Defaults are CHUNK_SIZE=1000 and CHUNK_OVERLAP=50.

Incremental indexing, and the one change that forces a full rebuild

The Memory Builder does not rebuild the whole index on every run, and the mechanism is described in detail. Each chunk is tagged with a source document ID and a version hash. When a document changes, the pipeline regenerates chunks for that document only, deletes the old chunks by metadata filter, and inserts the new ones. A diff step compares source documents against what is already indexed using those hashes, so only changed or new documents are processed.

Deletions get their own structure. A separate mapping table from doc_id to chunk_ids lives in a SQLite database, so removing a document means looking up its chunk IDs rather than scanning the vector store. That is a sensible design for a corpus that grows, and the README's own framing is that it keeps compute costs reasonable as the corpus grows.

The constraint that undoes all of this is stated in a callout: if you swap embedding models, you must rebuild the vector store from scratch, because the vector spaces are not compatible. The incremental machinery is keyed on document versions, not on which model produced the vectors. There is no migration path in the README for re-embedding an existing store under a new model. Plan the embedding model before you index anything, because the cost of changing your mind later is a full re-embed of the corpus.

Installing it and running a first query

Prerequisites are listed plainly: Python 3.12+, a CUDA 12.4+ GPU or an Apple Silicon M-series chip, Poetry 2.3.0+, Docker 24.0.6+ with Compose 5.0.2+, and for the UI Node 22.12+ with Yarn 1.22+. The NVIDIA Container Toolkit is optional but required for CUDA support.

The Makefile is the intended entry point. Run check first to confirm your shell resolves to the Python you expect, then run the setup target for your hardware. The README says to run Setup as your init command, or after Clean.

bash
make check
make setup_cuda

The setup_cuda target chains four steps: install_dependencies, install_pre_commit, migrate_db, and start_llama_server_cuda. That last step runs docker compose up -d, which starts the llama.cpp server defined in docker-compose.yml. The server is published on port 8080, and the README says it will be available at http://0.0.0.0:8080 and will show the llama-ui. The healthcheck curls http://localhost:8080/health with a 60 second start period, so give the container time before assuming it failed.

The model and retrieval settings live in .env.example. Copy it, pick your model, and set the embedding model. The default LLM is Meta-Llama-3.1-8B-Instruct-Q4_K_M, fetched from a Hugging Face GGUF repository, and the default embedding model is jinaai/jina-embeddings-v5-text-small-retrieval.

bash
cp .env.example .env
# MODEL="Meta-Llama-3.1-8B-Instruct-Q4_K_M"
# EMBEDDING_MODEL="jinaai/jina-embeddings-v5-text-small-retrieval"
# SYNTHESIS_STRATEGY="tree-summarization"

With the server up, build the index over your Markdown files in docs, then start the application. The README separates these: Build the memory index, then Run the Chatbot, which is what make start does. The Makefile comment says start launches both backend and frontend, ensuring the backend is running and ready before the frontend comes up. The backend reads HOST="0.0.0.0" and PORT=8000 from the environment.

bash
make start

To stop the inference server without touching your Python environment, the Makefile provides a dedicated target that runs docker compose down.

bash
make stop_llama_server

One caveat on the compose file: the llama-server image is pinned to ghcr.io/ggml-org/llama.cpp:server-cuda-b9501, and a TODO comment in the file notes that the tag should eventually come from an environment variable. Until that changes, upgrading llama.cpp means editing docker-compose.yml by hand.

Where this project is the wrong tool

The hallucination warning is not boilerplate. The README states that the large language model sometimes generates hallucinations or false information. Every RAG system has this property, but it matters more here because the answers are meant to be grounded in your own documentation. A confident wrong answer about your own product is worse than no answer, and there is no described citation mechanism that forces the model to point at the chunk it used.

The hardware requirement is the harder limit. CUDA 12.4+ or Apple Silicon is a prerequisite, not a recommendation. On a machine without either, the documented setup path does not apply, and the README directs you to llama.cpp's issue tracker rather than offering a CPU fallback. If your team runs Windows workstations or CPU-only cloud instances, this is the wrong starting point.

The Markdown-only ingestion is the third boundary. If your source of truth is a Confluence space, a PDF archive, or a database, you are building a conversion pipeline before you can evaluate the retrieval quality, and the project gives no help with that. It also means the chatbot's knowledge is only as current as the last time someone exported and indexed the docs.

Finally, the project is a working codebase rather than a product. There are no retrieved releases, and the README links to notes/todo.md for planned improvements. Treat version 0.7.0 as something you read and modify, not something you install and forget.

How it differs from a LangChain plus hosted-API stack

The obvious alternative is a LangChain pipeline pointed at a hosted model such as an OpenAI endpoint. The difference is not cosmetic. In that stack, retrieval is a chain of library abstractions, embeddings come from an API, and generation is billed per token. Here, the README explicitly says LangChain is not a dependency: the authors refactored the recursive splitter to avoid pulling it in. Embeddings come from Sentence Transformers running locally, and generation comes from a GGUF model served by llama.cpp.

That changes the failure modes. A hosted stack fails when the network or the API key fails, and it scales by paying more. This stack fails when the GPU driver or the pinned container image fails, and it scales by buying hardware. The .env.example even carries openai as a backend dependency and an LLAMA_SERVER_BASE_URL pointing at localhost, which suggests the backend talks to the llama.cpp server through an OpenAI-compatible interface. The README does not document switching that base URL to a remote provider, so do not assume it is a supported configuration.

The other difference is the incremental index. A naive LangChain tutorial re-embeds everything on each run. This project tracks document version hashes and a doc_id to chunk_ids mapping in SQLite so that only changed documents are reprocessed. If your corpus is large and changes often, that is the part worth studying in the source.

Maintenance, upgrades and licence

The repository is not archived, and the last push was on 2026-06-12. That is roughly three months before this writing, so the project has recent activity, but there are no retrieved releases, which means upgrades are tracked by commits and by the version field in pyproject.toml rather than by changelogs.

Upgrade cost concentrates in three places. The llama.cpp image tag is hardcoded in docker-compose.yml, so a llama.cpp upgrade is a manual edit. The dependency set is pinned tightly, including grpcio==1.78.0 with a comment that version 1.78.1 used by chromadb is yanked, which tells you the authors have already hit a resolution conflict in this tree. And the embedding model change is the expensive one: a new embedding model means re-embedding the entire corpus, as the README warns.

The licence is Apache-2.0, which is permissive and includes an explicit patent grant. That covers this repository's code. It does not cover the model weights you download: Meta Llama 3.1 has its own licence terms, and the default MODEL_URL points at a third-party GGUF repository on Hugging Face. Check the terms attached to whichever weights you actually load. This is a description of the files, not legal advice.

Editorial conclusion

Adopt it if you keep documentation in Markdown, already have an NVIDIA GPU or an Apple Silicon machine, and want the retrieval and generation loop on your own hardware rather than behind an API. Do not adopt it if you need a hosted endpoint, a managed vector store, or a corpus that is not Markdown: the upload allowlist is [".md"], and the README does not document rollback or a migration path for switching embedding models. Before committing, verify that the pinned llama.cpp image tag server-cuda-b9501 runs on your driver, that jinaai/jina-embeddings-v5-text-small-retrieval downloads at the size you expect, and that you can afford a full index rebuild if that embedding model ever changes.

Frequently asked questions

What is a RAG chatbot, in the context of umbertogriffo/rag-chatbot?

It is a chatbot that answers from a retrieved context rather than from the model's training data alone. In this project, the context comes from a collection of Markdown files that are chunked, embedded, and stored in Chroma, then retrieved when a question is asked.

Does umbertogriffo/rag-chatbot use RAG?

Yes. The README describes it as a RAG ChatBot that takes a collection of Markdown files as input and, when asked a question, provides the corresponding answer based on the context provided by those files.

How much does a RAG chatbot cost to run?

The project does not publish cost figures. Because inference runs through a local llama.cpp server and embeddings run locally with Sentence Transformers, the running cost is your own hardware rather than per-token API billing. The README's stated concern is compute cost as the corpus grows, which the incremental indexing pipeline is designed to limit.

What are examples of RAG AI in this project?

The README gives two context-overflow strategies as concrete examples: Create And Refine the Context, which synthesizes responses sequentially through all retrieved contents, and Hierarchical Summarization of Context, which answers each relevant section independently and then combines them.

Is a RAG chatbot an AI agent?

The README does not describe agent behaviour. It describes a fixed pipeline: rewrite the question with an LLM, retrieve relevant sections from Chroma, and generate an answer from that context, with chat history carried forward for CHAT_HISTORY_LENGTH turns.

Official sources

  1. Issues
  2. License: Apache-2.0
  3. README
  4. umbertogriffo/rag-chatbot on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/umbertogriffo-rag-chatbot.svg)](https://hysenlabs.com/projects/umbertogriffo-rag-chatbot)