# NVIDIA RAG Blueprint: a reference pipeline for NIM-backed retrieval

> The NVIDIA RAG Blueprint is a Python reference stack that wires NeMo Retriever models, a Milvus vector store and a LangChain orchestrator into a question-answering service. It is a starting point for teams already on NVIDIA hardware, not a drop-in product.

**NVIDIA-AI-Blueprints/rag** — This NVIDIA RAG blueprint serves as a reference solution for a foundational Retrieval Augmented Generation (RAG) pipeline.

- Repository: https://github.com/NVIDIA-AI-Blueprints/rag
- Website: https://build.nvidia.com/nvidia/build-an-enterprise-rag-pipeline
- Stars: 779 · Forks: 332
- Language: Python
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/nvidia-ai-blueprints-rag

## The gap the NVIDIA RAG Blueprint fills

Most RAG projects stall between a notebook demo and something a platform team will run. The hard parts are not the prompt. They are document extraction across tables and charts, hybrid retrieval, reranking, multi-turn state, and an API surface other services can call. The NVIDIA RAG Blueprint is a reference solution for that middle layer, published under Apache-2.0 with a sample UI and multiple deployment paths. It is written for engineers building question answering over their own enterprise documents, especially where governance, latency and scale requirements rule out a hosted black box.

The blueprint does not pretend to be vendor-neutral. Its retriever and extraction models are NVIDIA NIM microservices: llama-nemotron-embed-1b-v2 for embeddings, llama-nemotron-rerank-1b-v2 for reranking, and separate NIMs for page elements, table structure, graphic elements and OCR. Response generation defaults to a Nemotron model. If your organisation has no path to those endpoints, local or hosted, the blueprint's value drops sharply. That is the central trade-off, and the README is upfront about it rather than hiding it behind abstraction.

## How the orchestrator, vector store and NIMs fit together

The architecture has three visible layers. The bottom layer is NIM microservices: inference for generation, retrieval and reranking models, plus extractors for text, tables, charts and graphics. The middle layer is the RAG Orchestrator Server, which the README describes as LangChain-based and responsible for coordinating the user, retrievers, vector database and inference models, including multi-turn and context-aware handling. The top layer is the reference UI and the OpenAI-compatible API.

Storage is Milvus, and the dependency list confirms it: langchain-milvus and pymilvus with the milvus_lite and bulk_writer extras. The README describes the vector database as accelerated with NVIDIA cuVS, and lists GPU-accelerated index creation and search among the retrieval features. Hybrid search combines dense and sparse retrieval, with reranking applied afterwards. Query processing adds decomposition and dynamic filter expression creation, so a single user question can become several retrieval calls with generated filters.

The optional agentic path is the most interesting design choice. Instead of a fixed retrieve-then-generate chain, it adds a LangGraph plan-and-execute pipeline with scope discovery, parallel sub-tasks, synthesis and optional verification. It can be switched on for one request with agentic: true on /v1/generate, or for the whole deployment with ENABLE_AGENTIC_RAG, and the reference UI exposes it as Pipeline then Agentic. Stage events and reasoning traces stream back to the UI and API. That is a meaningful difference from most reference RAG repos, which stop at a single retrieval hop.

## Installing the NVIDIA RAG Blueprint and running a first query

The repository does not ship a single pip install line for the whole stack, because the stack is partly containers and partly NIM endpoints. What the repository does give you is a Python package definition and a deployment directory. The package metadata sets the interpreter range and the core dependencies, so the first thing to check is your Python version.

```toml
[project]
name = "nvidia_rag"
version = "2.6.0.rc1"
requires-python = ">=3.11,<3.14"
```

If your environment runs 3.10 or 3.14, the install will be refused before any model question comes up. The runtime dependencies pull in FastAPI and uvicorn for the service, langchain and langgraph for orchestration, langchain-milvus and pymilvus for the store, and langchain-nvidia-ai-endpoints for the NIM client. Redis appears in the dependency list, which is consistent with the multi-session support the README advertises.

The deployment story is directory-driven. The README names local docker, with and without NVIDIA Hosted endpoints, and Kubernetes as the supported options, with NIM Operator support listed under deployment and operations. The repository carries a deploy/ directory and a variables.env file at the top level for configuration. Read those before editing anything, because the blueprint expects NIM endpoint URLs and model names to be supplied rather than hardcoded.

Once the service is up, the API is OpenAI-compatible, and the agentic mode is a per-request flag rather than a separate endpoint. A request that turns it on looks like this:

```json
{
  "agentic": true,
  "messages": [
    { "role": "user", "content": "Compare the warranty terms across the three supplier contracts." }
  ]
}
```

That flag is documented as part of the /v1/generate surface. With it set, the response comes back through the plan-and-execute path and the UI shows stage events; without it, you get the standard retrieve-then-generate chain. The same switch exists at deployment level as ENABLE_AGENTIC_RAG, which is the simpler option if every caller wants the agentic path.

## Where the blueprint stops being the right tool

The dependency on NIM microservices is the first real limitation. The README lists a specific set of models for generation, embedding, reranking and extraction, and while the vector database is described as pluggable, the retrieval models are not presented as interchangeable. A team that has standardised on a different embedding model family is looking at replacing components, not configuring them.

The second limitation is hardware. GPU-accelerated index creation and search, cuVS-backed Milvus, and NIM inference all assume NVIDIA GPUs somewhere in the path. The blueprint does support NVIDIA Hosted endpoints for local docker deployments, which removes the GPU requirement from the application host but not from the request path. Latency and cost then depend on a network hop you do not control.

The third is scope. This is a reference solution, and the README calls it a foundational starting point that developers adapt and extend. There is no documented rollback procedure for an ingestion that goes wrong, no documented migration path between vector store backends despite the pluggable claim, and no stated compatibility guarantees across the 2.6.x releases. The changelog and release notes are the only place to check what moved between v2.6.0, v2.6.1 and v2.6.2. Treat upgrades as code changes, not as patch installs.

## How the NVIDIA RAG Blueprint differs from a plain LangChain pipeline

The obvious alternative is assembling the same chain yourself with LangChain or LlamaIndex and a vector store of your choice. The difference is not the retrieval algorithm, which is conventional hybrid search plus reranking. The difference is what comes bundled: extraction NIMs for tables and charts, programmable guardrails for content safety, RAGAS-based evaluation scripts, OpenTelemetry and Prometheus instrumentation, and a reference UI with multi-session support.

A hand-rolled pipeline lets you use any embedding model and any store, and it runs on CPU if your corpus is small. It also means you own the document extraction quality, the reranking model choice, the observability wiring and the UI. The blueprint trades that freedom for a known-good combination of parts. If your documents are mostly plain text and your query volume is low, the trade is not worth it. If your corpus is full of scanned tables and infographics, the extraction NIMs are the part you would struggle to reproduce quickly.

## Licence, upgrades and the cost of staying current

The project is Apache-2.0, and pyproject.toml declares license = "Apache-2.0" with a LICENSE file at the root. That covers the blueprint code. It does not cover the NIM microservices, which are separate NVIDIA products with their own terms, and the repository carries a LICENSE-3rd-party.txt for bundled dependencies. Anyone planning to redistribute a modified blueprint should read both files rather than assuming the Apache grant extends to the models. This is a description of what the repository states, not legal advice.

The upgrade surface is wider than a typical library. Releases arrive as version tags, with v2.6.2 on 2026-08-20, v2.6.1 on 2026-08-06 and v2.6.0 on 2026-06-04, and the last push to main was on 2026-09-03. Between releases you are tracking LangChain, LangGraph, pymilvus and FastAPI version bounds that shift independently. The pyproject file pins ranges rather than exact versions, so a fresh install today and a fresh install in three months will not resolve to the same dependency set. Pin your own lockfile if reproducibility matters. The repository ships a uv.lock, which is the intended mechanism for that.

## Conclusion

Adopt the NVIDIA RAG Blueprint if you already run NVIDIA GPUs or NIM endpoints and want a working retrieve-then-generate service you can fork, with an optional LangGraph agentic pipeline behind the same API. Do not adopt it if you need a CPU-only stack, a managed service with an SLA, or a pipeline that treats the vector store as anything other than Milvus-shaped. Before committing, verify three things in your own environment: that your NIM endpoints for llama-nemotron-embed-1b-v2 and llama-nemotron-rerank-1b-v2 respond, that the Milvus connection in your deploy configuration resolves, and whether you need ENABLE_AGENTIC_RAG set at deployment level or per request through the agentic flag on /v1/generate.

## FAQ

### What are NVIDIA AI blueprints?

They are reference solutions published by NVIDIA. This one is a foundational Retrieval Augmented Generation pipeline that developers adapt and extend for their own data and requirements.

### What is an AI RAG tool?

Retrieval-Augmented Generation combines an LLM with retrieval from trusted data sources, so answers are grounded in enterprise knowledge rather than model memory alone. The NVIDIA RAG Blueprint is a reference implementation of that pattern.

### Does generative AI use RAG?

The NVIDIA RAG Blueprint is built on the premise that it should, for enterprise question answering: it grounds LLM responses in retrieved documents from your own data sources to reduce hallucinations and keep answers current.

### How does the RAG pipeline work?

The blueprint runs a LangChain-based orchestrator that coordinates the user, retrievers, a Milvus vector database and inference models. Documents are extracted and embedded, queries go through hybrid dense and sparse search with reranking, and generation is handled by a NIM inference model. An optional agentic mode replaces the single retrieval hop with a LangGraph plan-and-execute pipeline.

## Sources

- [License: Apache-2.0](https://github.com/NVIDIA-AI-Blueprints/rag/blob/main/LICENSE)
- [NVIDIA-AI-Blueprints/rag on GitHub](https://github.com/NVIDIA-AI-Blueprints/rag)
- [Project website](https://build.nvidia.com/nvidia/build-an-enterprise-rag-pipeline)
- [README](https://github.com/NVIDIA-AI-Blueprints/rag/blob/main/README.md)
- [Releases](https://github.com/NVIDIA-AI-Blueprints/rag/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/nvidia-ai-blueprints-rag
