# OPEA GenAIExamples: Docker and Kubernetes Samples for RAG, Summarization and Code Generation

> GenAIExamples is OPEA's collection of microservice-based GenAI samples, from ChatQnA to DocSum, deployable with Python, Docker Compose or Kubernetes. It is a reference architecture catalogue, not a framework you build a product on.

**opea-project/GenAIExamples** — Generative AI Examples is a collection of GenAI examples such as ChatQnA, Copilot, which illustrate the pipeline capabilities of the Open Platform for Enterprise AI (OPEA) project.

- Repository: https://github.com/opea-project/GenAIExamples
- Website: https://opea.dev
- Stars: 741 · Forks: 339
- Language: Shell
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/opea-project-genaiexamples

## What GenAIExamples actually solves, and for whom

Getting a retrieval-augmented generation stack running is mostly assembly work: a vector store, an embedding service, a reranker, an LLM server, a gateway, and the glue between them. GenAIExamples ships that assembly as numbered directories. ChatQnA is the RAG chatbot, DocSum summarizes documents, CodeGen generates code, DocIndexRetriever handles retrieval without generation, and the repository also carries VisualQnA, Text2Image, Translation, AudioQnA, GraphRAG, HybridRAG, AgentQnA and others.

The audience is narrower than the name suggests. This is not a library you add to an application. It is a catalogue of deployable topologies, aimed at engineers who want to stand up a pipeline, measure it, and then decide what to build. The README describes the examples as microservice-based samples that simplify deploying, testing and scaling GenAI applications, and states they are compatible with both Docker and Kubernetes across Gaudi, Xeon, AMD EPYC CPUs, AMD Instinct GPUs and NVIDIA GPUs.

That framing matters. If you need a RAG component inside an existing Python service, the useful artifact here is the architecture, not the code. If you need to know what a four-service RAG pipeline looks like when deployed and how it behaves on a specific accelerator, this repository is the shortest path to that answer.

## The GenAIComps, GenAIInfra and GenAIEval split

GenAIExamples is one of four repositories, and the boundaries explain most of its design. GenAIComps holds the microservice components: llm, embedding, reranking and others. GenAIExamples composes those components into applications such as ChatQnA and DocSum. GenAIInfra is the containerization and cloud-native layer that deploys the examples. GenAIEval measures throughput, latency and accuracy so you can compare hardware configurations.

So the data flow in a ChatQnA deployment is a chain of separately addressable services, each backed by a component from GenAIComps, wired together by a compose file or a Kubernetes manifest from GenAIInfra, with GenAIEval attaching to the running services to produce metrics. Nothing in that chain is monolithic, which is the point: you can swap the LLM backend or the reranker without rewriting the application.

The cost of that split is that debugging spans repositories. A failure in a ChatQnA deployment could originate in a GenAIComps service image, in the GenAIInfra orchestration, or in the example's own configuration. The repository layout reflects this: each use case is a top-level directory, and the actual wiring lives in deployment files inside it rather than in shared Python modules. Plan for that when you file an issue or read a stack trace.

## Installing GenAIExamples and running a first example

The README lists three deployment methods: Python startup, Docker Compose and Kubernetes. Docker Compose is the least demanding of the three. The prerequisite section states you should have docker compose installed, and that deployment is based on released Docker images by default, with a docker image list in docker_images_list.md for details.

Start by confirming Compose is present:

```bash
docker compose version
```

You should see a Compose v2 version string. If the command is not found, the README points to the Docker Compose install page rather than bundling an installer.

The repository also carries a requirements.txt at the top level, which pulls in opea-eval, kubernetes, locust, prometheus_client and related packages. That file supports the benchmarking and evaluation scripts (benchmark.py, deploy.py, deploy_and_benchmark.py) rather than the examples themselves:

```bash
pip install -r requirements.txt
```

After that, the workflow is per example: enter a use case directory such as ChatQnA and use its deployment files. The README does not inline those commands, it points at the documentation site and at a supported_examples.md file that maps use cases to supported deployment types. Check supported_examples.md before assuming your chosen example supports Compose on your accelerator.

For Kubernetes, the prerequisites are stricter. You need a cluster, optionally Helm version 3.15 or later for Helm chart deployment, and optionally GMC installed in the cluster if you want to deploy through the microservices connector. The README links out to installation guides for each rather than providing them.

## Hardware sizing is the real gate

The README's reference configuration table is the most concrete constraint in the repository, and it is easy to skim past. For the Xeon row with Intel/neural-chat-7b-v3-3, the reference is 64 vCPUs, 365 GB disk, 100 GB RAM and Ubuntu 24.04. The Gaudi row for the same model calls for 1 or 2 Gaudi cards with 16 vCPUs, 365 GB disk and 100 GB RAM. An AWS Xeon configuration is listed as c7i.16xlarge with 64 vCPUs, 100 GB disk, 64 GB RAM and Ubuntu 24.04.

Those numbers describe a 7B model. A laptop or a small VM will not run this stack, and the failure will look like an out-of-memory kill partway through model loading rather than a clear error. If your target is a smaller footprint, the examples are the wrong starting point; the component repositories are a better place to look.

The table is also the clearest statement of intent. Two of the three reference rows point at Intel Tiber Developer Cloud, and the hardware list leads with Gaudi and Xeon. This is an Intel-aligned catalogue. The README does state support for AMD EPYC, AMD Instinct and NVIDIA GPUs, so the examples are not exclusive, but the sizing guidance and the validated_configurations.md file are where the tested combinations live. Treat any configuration not in those files as unverified.

## Where GenAIExamples is the wrong tool

The repository is a set of deployment samples, and that shapes what it cannot do. There is no packaging story for embedding an example inside your own service. The top-level pyproject.toml configures isort, black, codespell and ruff with a 120-character line length and a Python 3.10 target, which is lint configuration for the repository's own scripts, not a library build. If you need a RAG pipeline as an importable dependency, you are looking at the wrong layer.

Second, the deployment surface is not fully documented in the README. It defers to the documentation site for architecture and deployment guides, and to supported_examples.md for the use case and deployment-type matrix. The README does not document rollback, and it does not describe what happens when a component image version drifts from the example's expected interface. Since the components come from a separate repository with its own release cadence, that drift is a plausible failure mode rather than a theoretical one.

Third, the model choice is fixed in the reference configurations. If you need a model that is not in the validated set, you are outside what the project has tested, and the sizing table will not tell you what to provision. That is a normal boundary for a sample repository, but it should be a conscious decision rather than a surprise after the first deployment.

## GenAIExamples compared with LangChain or LlamaIndex templates

The obvious alternative for a RAG starting point is a Python framework such as LangChain or LlamaIndex, and the difference is architectural rather than qualitative. Those frameworks give you an in-process library: you import retrievers and chains, and the application is a single Python program. GenAIExamples gives you a deployed topology: separate services for embedding, retrieval, reranking and generation, wired by Compose or Kubernetes manifests, with GenAIEval available to measure the result.

The trade-off is direct. A framework is faster to prototype and easier to embed, but you own the serving, scaling and hardware placement yourself. GenAIExamples hands you a running multi-service deployment with hardware-specific configurations already written, but every change crosses a service boundary and the code you edit is YAML and shell rather than Python.

A second alternative is building directly on GenAIComps without the examples. That gets you the same components with fewer opinions about how they are combined, which suits a team that already knows its target architecture. The examples are most valuable when you do not yet know what that architecture should look like and want a working reference to measure against.

## Maintenance, releases and the Apache-2.0 licence

The repository is not archived, and the last push was on 2026-09-09, so it is being worked on. Recent releases are v1.5 on 2025-12-22, v1.4 on 2025-08-25 and v1.3 on 2025-05-14, which gives a rough cadence of one release per quarter. The repository also carries a pre-commit configuration and a prettier ignore file, so contributions are linted on the way in.

Upgrade cost depends on which layer you pin. If you deploy from released Docker images as the README recommends, an upgrade means pulling new component images and reconciling them with the example's deployment files. Because the components live in GenAIComps and the orchestration in GenAIInfra, a version bump can require changes in three places. The release notes for v1.5 are the place to check what moved between v1.4 and v1.5 before you pull.

The licence is Apache-2.0, which permits commercial use and modification and includes an explicit patent grant. It also requires that you preserve copyright and licence notices and state significant changes. The repository carries a LEGAL_INFORMATION.md file alongside LICENSE. Note that the licence covers the repository's own code; the models the examples download carry their own terms, and the README's reference configurations name Intel/neural-chat-7b-v3-3 specifically. Check the model licence separately. This is a description of the licence text, not legal advice.

## Conclusion

Adopt GenAIExamples if you need a working RAG or summarization topology to compare hardware and serving stacks on, and you already have Docker Compose or a Kubernetes cluster plus the RAM the reference configurations list. Do not adopt it as a library to import into an existing application: these are deployment samples built from GenAIComps services, and the repository's own files are the integration surface. Before committing, check the example directory's compose file against the hardware you actually have, confirm the model size against the README's reference configuration table (the Xeon row lists 64 vCPUs and 100 GB RAM), and read the v1.5 release notes for what changed since v1.4.

## FAQ

### What is GenAIExamples from the OPEA project?

It is a collection of microservice-based generative AI samples, including ChatQnA, DocSum, CodeGen and DocIndexRetriever, that illustrate the pipeline capabilities of the Open Platform for Enterprise AI. The README states the examples are compatible with both Docker and Kubernetes.

### Is ChatGPT an example of GenAI?

The repository does not discuss ChatGPT. Its own examples are chat and retrieval applications such as ChatQnA and VisualQnA, built from GenAIComps services and deployed with Docker Compose or Kubernetes.

### What are the top GenAI tools covered by GenAIExamples?

The repository does not rank tools. Its README groups use cases by scenario: question answering (ChatQnA, VisualQnA), image generation (Text2Image), summarization (DocSum), code generation (CodeGen), retrieval (DocIndexRetriever) and fine-tuning (InstructionTuning).

### What are some real life examples of general AI in GenAIExamples?

The README's use case table maps scenarios to concrete examples: question answering to ChatQnA and VisualQnA, image generation to Text2Image, content summarization to DocSum, code generation to CodeGen, information retrieval to DocIndexRetriever, and fine-tuning to InstructionTuning.

### What types of GenAI are there in GenAIExamples?

The repository covers several types through its examples: retrieval-augmented chat and visual question answering, text-to-image generation, document summarization, code generation, document retrieval, translation and instruction tuning. The README notes the full list and supported deployment types are in the documentation.

## Sources

- [License: Apache-2.0](https://github.com/opea-project/GenAIExamples/blob/main/LICENSE)
- [opea-project/GenAIExamples on GitHub](https://github.com/opea-project/GenAIExamples)
- [Project website](https://opea.dev)
- [README](https://github.com/opea-project/GenAIExamples/blob/main/README.md)
- [Releases](https://github.com/opea-project/GenAIExamples/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/opea-project-genaiexamples
