NVIDIA GenerativeAIExamples: reference RAG and agent workflows for NVIDIA's inference stack
Generative AI reference workflows optimized for accelerated infrastructure and microservice architecture.
At a glance
- What is it?
- The repository is a collection of Jupyter notebooks and Docker Compose reference applications for RAG, agentic workflows, NeMo microservices and Vision NIMs. It is a fast way to see NVIDIA's inference stack working end to end, and a poor fit if you need a supported product rather than sample code.
- Who is it for?
- Adopt it if you are building on NVIDIA NIMs or NeMo microservices and want a runnable reference for RAG or agentic pipelines before writing your own service layer. Skip it if you need a supported, versioned product with an upgrade path, or if you are not already committed to NVIDIA inference infrastructure.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 20 days ago.
- What is it written in?
- Mainly Jupyter Notebook, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
DEEP OPEN-SOURCE ANALYSIS
What NVIDIA GenerativeAIExamples is for
The repository is a starting point for developers who want to build generative AI systems on NVIDIA's software ecosystem. That is the framing the README gives, and it matches the directory layout: RAG/, nemo/, nemotron/, finetuning/, vision_workflows/, industries/, community/ and oss_tutorials/ sit side by side at the top level.
The audience is narrow and specific. If you are evaluating whether to use NVIDIA NIM microservices for retrieval, or NeMo microservices for fine-tuning and guardrails, this repository shows the pieces wired together. If you are choosing a vector database, an orchestration framework, or a model provider, the examples assume answers to those questions already.
The unit of delivery is the example, not the library. There is no package to install and no API surface to depend on. You clone the repository, pick a directory, and run what is inside it.
How the RAG and NIM examples are put together
Two shapes of example appear throughout. The first is a Jupyter notebook that calls hosted endpoints, such as the agentic RAG pipeline built with Llama 3.1 and NeMo Retriever NIM microservices. The second is a Docker Compose application, such as the basic RAG pipeline under RAG/examples/basic_rag/langchain/, which builds containers for the chain server, the retrieval service and a front end.
The Compose examples are the more informative ones for architecture, because the service boundaries are visible in the compose file rather than implied by notebook cells. A chain server handles generation, a retrieval component handles embedding and search, and the playground UI sits in front of both. Swapping a component means changing a container or an endpoint, not rewriting the application.
GPU acceleration enters through the model serving layer. TensorRT and Triton Inference Server appear in the repository topics, and the vision examples lean on NIM containers directly. The notebooks themselves are mostly orchestration code; the heavy work happens in the NIM or NeMo service the notebook calls.
The repository also carries a vision_workflows tree pulled in as a git submodule, which is why the documented clone command uses --recurse-submodules. Without that flag the directory is empty and the Vision NIM links resolve to nothing.
Installing and running the basic RAG pipeline
The README's Try it Now section is the shortest path to something running. It starts with an API key from the NVIDIA API Catalog, which is exported as an environment variable. The value has the nvapi- prefix.
export NVIDIA_API_KEY=nvapi-...Clone the repository. If you also want the Vision NIM workflows, add the recursive flag the README documents for that case.
git clone https://github.com/nvidia/GenerativeAIExamples.gitMove into the basic LangChain RAG example and bring the stack up. The build step compiles the containers locally, so the first run takes noticeably longer than later ones.
cd GenerativeAIExamples/RAG/examples/basic_rag/langchain/
docker compose up -d --buildWhen the containers report healthy, the sample RAG Playground is served on port 8090. Submit a query there and you should get an answer grounded in the example's document set. Tear everything down with docker compose down from the same directory.
docker compose downTwo constraints are worth noting before you start. The NVIDIA_API_KEY export has to be present in the shell that runs Compose, because the containers read it at startup. And the port is fixed at 8090 in the documented example, so a local service already bound there will collide.
Where the repository stops being the right tool
The examples are reference implementations, and the README presents them that way. Nothing in the repository promises a stable interface between releases. The release history supports that reading: v0.6.0, v0.7.0 and v0.8.0 arrived between May and August 2024, and the most recent push is dated 2026-08-20, so the tree has moved since the last tagged release. If you fork an example and build a product on it, you are maintaining that fork yourself.
The dependency surface is the other problem. A Compose example pulls container images, model endpoints and a Python stack. Any of those can move independently of the notebook you copied. There is no lockfile discipline described for the examples as a whole.
Some material also points outside the repository. The Llama 3.1 agentic RAG entry links to a v0.7.0 path for the LangChain local NIM notebook and to a developer blog post, and the Morpheus event-driven RAG example links to a v0.7.0 tree. Those links pin to old tags, so following them gives you code that does not match the current main branch.
Finally, the examples assume GPU-capable infrastructure and, in several cases, a running NVIDIA API Catalog account. If your constraint is CPU-only deployment or a provider-neutral stack, this repository is the wrong starting point.
NeMo microservices and the Data Flywheel tutorials
The Data Flywheel material is the most involved part of the repository. It is not a single example but a set of tutorials that chain NeMo Datastore, NeMo Entity Store, NeMo Customizer, NeMo Evaluator and NeMo Guardrails microservices together with NIMs.
The tool-calling tutorial is the concrete one. It customizes Llama-3.2-1B-Instruct on the xLAM function-calling dataset, evaluates the result, then applies guardrails. The README describes the NeMo microservices platform as deployable across Kubernetes clusters in cloud or on-premises environments, which is the deployment story: these are not laptop examples.
That is the trade-off. The flywheel tutorials show a full customization loop, but they assume a Kubernetes cluster with the NeMo microservices deployed. The Compose RAG example runs on a single machine with an API key. The gap between those two entry points is large, and the README does not bridge it.
The Safer Agentic AI tutorials sit in the same tree. One audits LLMs with NeMo Auditor to surface vulnerabilities to unsafe prompts; another runs inference with multiple rails in parallel to reduce latency. Both are notebooks rather than deployable services.
How it compares to LangChain templates and LlamaIndex examples
LangChain maintains its own template and example collection, and LlamaIndex ships starter projects for retrieval. The difference is what the examples are optimized around.
LangChain templates are organized around the framework's abstractions. You get a chain or an agent expressed in LangChain primitives, and the model and vector store are configuration choices. Swapping the LLM is expected.
NVIDIA's examples are organized around NVIDIA's serving stack. The model endpoint is a NIM, the retrieval components come from NeMo Retriever, and the acceleration story runs through TensorRT and Triton. LangChain appears as the orchestration layer in several RAG examples, but it is the part you could replace.
So the choice is not which collection has better code. It is which layer you want held fixed. If you want to compare model providers freely, the LangChain templates keep that door open. If you have already committed to NIMs and want to see how retrieval, guardrails and fine-tuning fit around them, this repository is the more direct answer, and it is the one whose examples will not fight your serving layer.
Licence, maintenance and upgrade cost
The repository is Apache-2.0, and the top level also carries a separate LICENSE.DATA file. That split matters in practice: code and data can carry different terms, and the data file is the one to read before you reuse any dataset or document collection an example ships with. This is not legal advice; read both files and the terms attached to any model or dataset an example pulls.
Maintenance is best judged by the push date rather than the release tags. The last push was on 2026-08-20 and the repository is not archived, so the tree is being touched. The most recent tagged release, v0.8.0, is dated 2024-08-21, which means the release cadence and the commit cadence have diverged.
Upgrade cost is the practical consequence. Because examples are copied rather than depended on, upgrading means re-reading the example you forked and reconciling it with the current tree. Notebooks that call hosted endpoints are cheap to keep current. Compose applications with pinned container images are not, and the NeMo microservices tutorials carry the highest cost because a Kubernetes deployment sits underneath them.
Editorial conclusion
Adopt it if you are building on NVIDIA NIMs or NeMo microservices and want a runnable reference for RAG or agentic pipelines before writing your own service layer. Skip it if you need a supported, versioned product with an upgrade path, or if you are not already committed to NVIDIA inference infrastructure. Before adopting, verify the container images and model endpoints your target example actually calls, because the repository mixes notebooks that run against hosted NVIDIA API Catalog endpoints with examples that expect locally deployed NIM containers.
Frequently asked questions
Is ChatGPT a generative AI?
The repository does not discuss ChatGPT. Its own examples cover generative systems built on NVIDIA software: RAG pipelines, agentic workflows, fine-tuning with NeMo microservices, and Vision NIM workflows such as video stream monitoring and natural language image search.
What is the best example of generative AI?
The README does not rank examples. It groups them by use case: a basic RAG pipeline under RAG/examples/basic_rag/langchain/, an agentic RAG pipeline with Llama 3.1 and NeMo Retriever NIM microservices, Data Flywheel tutorials, and Vision NIM workflows for video alerts, NV-CLIP image search and text extraction.
What are the top 5 generative AI tools?
The repository does not list tools by popularity. It names the NVIDIA components its examples use: NIM microservices, NeMo Datastore, NeMo Entity Store, NeMo Customizer, NeMo Evaluator, NeMo Guardrails, NeMo Auditor, TensorRT and Triton Inference Server.
What is the difference between AI and generative AI?
The repository covers generative systems specifically: RAG pipelines, agentic workflows, fine-tuning, guardrails and vision-language models. Traditional CV components appear only where they feed a generative pipeline, as in the vision text extraction example that combines VLMs, LLMs and CV models.
What are some generative AI examples?
The repository's own examples include a basic RAG pipeline run through Docker Compose, an agentic RAG pipeline with Llama 3.1, tool-calling fine-tuning with NeMo microservices, knowledge graph RAG with RAPIDS, and Vision NIM workflows for video monitoring, NV-CLIP search and few-shot classification with NVDINOv2 and Milvus.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/nvidia-generativeaiexamples)
Community notes