All concepts
Concept

What is Retrieval-augmented generation?

Retrieval-augmented generation (RAG) is a pattern in which a language model's answer is built from documents fetched at question time instead of from the model's weights alone. The model writes the answer, but a retrieval step decides what it gets to read.

Published September 28, 2026

How RAG works

The mechanism has two phases, and they are usually built and debugged separately.

Indexing happens ahead of time. Source documents are split into chunks, each chunk is converted into a vector by an embedding model, and the vectors are written to a store that supports nearest-neighbour search. Chunk size, overlap and any metadata attached to each chunk are decided here. Everything downstream inherits those choices.

At query time the user's question is embedded with the same model, the store returns the top-k chunks by similarity, and those chunks are inserted into the prompt as context. The model then produces an answer conditioned on the question plus the retrieved text. A citation layer can map sentences in the answer back to chunk identifiers, but that is an add-on, not part of the core loop.

Three variants are common. Vector-only retrieval is the simplest and fails on exact identifiers. Hybrid retrieval combines vector search with keyword search such as BM25, which helps when the query contains a part number or a function name. Graph retrieval walks a knowledge graph instead of a flat vector index, which suits questions about relationships between entities.

The retrieval step is where most quality is won or lost. If the right chunk is not in the top-k, no prompt instruction recovers it. If the context window is filled with near-duplicate chunks, the model has less room for the question and the answer degrades quietly.

When you need RAG and when you do not

RAG is the right tool when the answer depends on facts the model was never trained on, when those facts change, or when you must show where an answer came from. Internal policy documents, product catalogues, support tickets and code repositories all fit. The retrieval step also acts as an access boundary: if a user cannot retrieve a document, the model cannot quote it.

It is the wrong tool when the task is style, format or reasoning over material already in the prompt. Fine-tuning changes how a model writes; RAG changes what it knows at answer time. If your failure mode is tone or output schema, retrieval will not fix it. If the corpus is small enough to fit in the context window on every request, a plain long-context prompt is simpler and has fewer moving parts.

Cost and latency are the trade. Every query pays for an embedding call, a vector search and a longer prompt. The longer prompt is usually the dominant cost, because retrieved chunks are billed as input tokens on every turn. Caching helps for repeated questions and does nothing for the long tail.

A middle path exists for stable, narrow domains: put a small curated set of documents directly in the system prompt and skip the index entirely. That approach does not scale past a few hundred pages, but below that threshold it removes an entire service from your stack.

Pitfalls and limits

Chunking is the first failure point. Split too small and a chunk loses the antecedent it needs; split too large and the embedding averages several topics into a vector that matches nothing well. Tables, code and PDFs with multi-column layouts break naive splitters, which is why several projects in this space ship dedicated document parsers.

Retrieval quality is hard to measure without a labelled set. Teams often tune top-k by eyeballing a handful of queries, which overfits to those queries. A small evaluation set of question, expected chunk and expected answer pairs catches regressions that a demo will not.

Stale indexes are a quiet failure. If the source changes and the index is not rebuilt, the model answers from an old version with full confidence. Incremental re-indexing is a real engineering cost, and it is the part most prototypes omit.

Security has a retrieval-shaped hole. If the index is shared across tenants, a prompt injection in one document can surface in another tenant's answer. Filtering by tenant at query time, not at prompt time, is the usual mitigation.

Finally, retrieval does not eliminate hallucination. It reduces the space in which the model can invent, but a model can still misread a retrieved chunk, merge two chunks incorrectly, or answer from its weights when the retrieved context is thin. Grounding is a probability shift, not a guarantee.

How RAG shows up in open-source projects

The projects below sit at different layers of the same stack, and their descriptions and analyses mark where each one stops.

langgenius/dify bundles a visual workflow builder, a RAG pipeline, agent tools and model management behind one Docker Compose stack. It fits teams that want to ship LLM apps without writing orchestration code, and it costs a multi-service deployment to run. The RAG pipeline is one node type inside a larger product, not a standalone library.

infiniflow/ragflow is described as an Apache-2.0 RAG engine that pairs deep document understanding with a visual chunking UI and agent components. It installs through Docker Compose, and the stated resource floor of 16 GB RAM and 50 GB disk is the part most teams underestimate. That floor is the price of the document parsing layer.

pathwaycom/pathway is a Python ETL framework for stream processing, real-time analytics, LLM pipelines and RAG. It targets engineers who want one pipeline to work for batch and streaming, and it pays for that with a Rust engine, a BUSL-1.1 core and a licensing split between free and enterprise editions. The licence is the constraint to check before building on it.

pathwaycom/llm-app ships Docker-friendly AI pipeline templates that stay in sync with Google Drive, Sharepoint, S3, Kafka and PostgreSQL. Its interesting property is subtraction: no separate vector database, no cache and no API framework. That is a strong opinion about how much infrastructure a RAG pipeline needs.

zylon-ai/private-gpt is not a model runner. It is a Claude-shaped API layer in front of any OpenAI-compatible inference server, supplying retrieval, tools, MCP and database access. It assumes you already have an inference server and stops there.

headroomlabs-ai/headroom is an Apache-2.0 Python project that compresses tool outputs, logs, files and RAG chunks before they reach an LLM, available as a library, a local proxy and an MCP server. The savings on JSON are large, but the reversible cache and the wrapped-agent setup are the parts to check before adopting. This is a cost lever on the retrieval payload, not a retrieval system.

abhigyanpatwari/GitNexus indexes a codebase into a graph of dependencies, call chains and execution flows, then exposes it to editors through MCP, with a built-in Graph RAG agent. The npm install path is the practical one, and the licence is not open source in the usual sense. Treat it as a code-search tool with a graph index.

Shubhamsaboo/awesome-llm-apps packages over 100 open-source AI agents, agent skills and RAG apps as ready-to-run templates. It is a copy-and-run repository, not a framework, with real constraints around maintenance and vendor lock-in. dair-ai/Prompt-Engineering-Guide is a reading resource first and a runnable app second, covering prompt engineering, context engineering, RAG and agents. ruvnet/ruflo adds swarms, persistent memory and hooks around Claude Code and Codex, with two install paths that do not expose the same tools.

Choosing a starting point

Pick by the layer you are missing, not by the longest feature list. If you need a working pipeline today and can run Docker Compose, langgenius/dify and infiniflow/ragflow both deliver one, at the cost of a multi-service deployment and, for RAGFlow, a documented resource floor. If you want a library to embed in existing code, headroomlabs-ai/headroom addresses the token cost of retrieved context rather than retrieval itself.

If your data is already in a warehouse or a stream, pathwaycom/pathway and pathwaycom/llm-app are built around that assumption, with the BUSL-1.1 core as the licensing constraint. If you already run an OpenAI-compatible inference server, zylon-ai/private-gpt adds the API and retrieval layer without replacing it. For codebases specifically, abhigyanpatwari/GitNexus takes a graph approach.

Before adopting any of them, check two things: when the repository was last pushed, and what the licence permits for your use. Neither is visible from a feature list.

In practice

RAG is a retrieval problem wearing a generation costume. The model is the easy part to swap; the index, the chunking and the freshness of the data decide whether answers are right. Start by measuring retrieval on a small labelled set, then pick a project from the layer you are actually missing. Read the licence and the last push date before you commit.

langgenius/difyDify is an open-source LLM app platform combining agentic workflows, RAG pipelines and model management, deployable on cloud, VPC, or self-hosted infrastructure.157,512 stars · TypeScriptShubhamsaboo/awesome-llm-appsGitHub describes it as 100+ AI Agents, Agent Skills and RAG Apps - Free and Open Source.. The repository metadata lists Python as its primary language. The metadata lists the Apache-2.0 license. This article stays within the project description and details documented in the GitHub repository README.140,206 stars · Pythoninfiniflow/ragflowRAGFlow is a leading open-source Retrieval-Augmented Generation (RAG) engine that fuses reliable RAG with Agent capabilities to create a superior context layer for LLMs.91,300 stars · Godair-ai/Prompt-Engineering-GuideGitHub describes it as 🐙 Guides, papers, lessons, notebooks and resources for prompt engineering, context engineering, RAG, and AI Agents.. The repository metadata lists MDX as its primary language. The metadata lists the MIT license. This article stays within the project description and details documented in the GitHub repository README.78,721 stars · MDXheadroomlabs-ai/headroomCompress tool outputs, logs, files, and RAG chunks before they reach the LLM. 20% fewer tokens for coding agents, 60-95% fewer tokens for JSON, same answers. Library, proxy, MCP server.74,088 stars · Pythonruvnet/rufloRuflo is an agent meta-harness for Claude Code and Codex, adding 100+ specialized agents, coordinated swarms, self-learning memory, and federation across machines.73,502 stars · TypeScriptpathwaycom/pathwayPathway is a Python ETL framework for stream processing, real-time analytics, LLM pipelines and RAG, driven by a scalable Rust engine using Differential Dataflow.62,215 stars · Pythonpathwaycom/llm-appReady-to-run cloud templates for RAG, AI pipelines, and enterprise search with live data. Docker-friendly.Always in sync with Sharepoint, Google Drive, S3, Kafka, PostgreSQL, real-time data APIs, and more.58,897 stars · Jupyter Notebookzylon-ai/private-gptComplete API layer for private AI applications on local models: RAG, skills, tools, MCP, text-to-sql, and more. Works with any OpenAI-compatible inference server.57,537 stars · Pythonabhigyanpatwari/GitNexusGitNexus: The Zero-Server Code Intelligence Engine - GitNexus is a client-side knowledge graph creator that runs entirely in your browser. Drop in a git repository (Github, Gitlab, Azure, Local) or ZIP file, and get an interactive knowledge graph with a built in Graph RAG Agent. Perfect for code exploration47,648 stars · TypeScripthpcaitech/ColossalAIMaking large AI models cheaper, faster and more accessible. Build your AI agents, chatbots, and RAG applications with HPC-AI Model APIs!41,443 stars · PythonHKUDS/LightRAG[EMNLP2025] "LightRAG: Simple and Fast Retrieval-Augmented Generation.39,860 stars · Python

Sources

  1. langgenius/dify repository
  2. Shubhamsaboo/awesome-llm-apps repository
  3. infiniflow/ragflow repository
  4. dair-ai/Prompt-Engineering-Guide repository
  5. headroomlabs-ai/headroom repository