Model or dataset
Azure-Samples/azure-search-openai-demo avatar
Azure-Samples/azure-search-openai-demo

azure-search-openai-demo: a RAG reference app you deploy with azd, not a library you import

A sample app for the Retrieval-Augmented Generation pattern running in Azure, using Azure AI Search for retrieval and Azure OpenAI large language models to power ChatGPT-style and Q&A experiences.

7,773 stars5,470 forksPythonMIT

At a glance

What is it?
The Python sample wires Azure AI Search retrieval to Azure OpenAI chat behind a web frontend, and its main value is a working end-to-end pipeline plus a productionizing guide that warns you not to ship it as-is.
Who is it for?
Adopt it if you want a working Azure RAG pipeline to read, fork or use as an architecture reference before building your own service. Do not adopt it as a production application: the README states the code was built to showcase Azure services and advises against putting it into production without additional security work.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 5 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 28, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What azure-search-openai-demo actually is, and what it is not

This is a sample application, not a package. You do not pip install it into an existing service; you clone the repository, provision Azure resources, and run the web app it contains. The README describes it as a solution that creates a ChatGPT-like frontend over your own documents using Retrieval Augmented Generation, with Azure OpenAI Service for GPT models and Azure AI Search for indexing and retrieval. The backend is Python, and the repository also points to JavaScript, .NET and Java samples based on the same design.

The audience is engineers who need to see the whole RAG loop assembled: document ingestion, chunking, embedding, index creation, query, prompt construction, answer generation with citations, and a UI that exposes the retrieval settings. The repository ships sample data about a fictitious company called Zava, so the pipeline runs end to end without you supplying documents first. That matters more than it sounds. Most RAG tutorials stop at the retrieval call and leave the ingestion and index schema as an exercise.

The framing in the README is unusually blunt for a Microsoft sample. A security notice states that the template and its configuration were built to showcase Azure services, and advises customers not to make the code part of production environments without implementing or enabling additional security features. Treat that as the project's own scope statement. If you want a reference implementation to read and adapt, it fits. If you want a hosted product, it does not.

The retrieval pipeline: ingestion, index, then grounded answers

The mechanism is a two-phase pipeline. Offline, a preparation step reads your documents, splits them into chunks, generates embeddings, and writes them into an Azure AI Search index. Online, each user question is sent to the search index, the top matching chunks come back, and those chunks are inserted into the prompt sent to the Azure OpenAI model. The answer returned to the browser carries citations pointing back at the source chunks, and the UI renders the thought process alongside the response.

The repository layout reflects that split. There is an app directory for the running service, a data directory for the sample documents, and a scripts directory for the ingestion tooling. The pyproject.toml adds app/backend, scripts, app/functions and evals to pythonpath, which is the clearest signal that these are separate executables sharing code rather than one importable package. The same file shows the project targets Python 3.10 and runs ruff, black and ty for linting, formatting and type checking, with tests under tests and a separate evals tree for evaluation code.

Retrieval behaviour is tunable rather than fixed. The README lists a settings panel in the UI for tweaking behaviour and experimenting with options, so the retrieval parameters are visible to whoever is testing the app instead of buried in configuration. The repository also supports many document formats and cloud data ingestion, and there are optional paths for multimodal models over image-heavy documents, speech input and output, and Microsoft Entra login with data access control. Each of those is opt-in, which is why the docs directory has separate pages for them.

Installing azure-search-openai-demo and running a first question

There are three documented starting points: GitHub Codespaces, VS Code Dev Containers, and a local environment. All of them assume an Azure subscription and the Azure Developer CLI. The README lists the required account permissions explicitly: your account needs Microsoft.Authorization/roleAssignments/write, held by roles such as Owner or User Access Administrator, and Microsoft.Resources/deployments/write at the subscription level. If you lack subscription-level rights, the docs describe deploying into an existing resource group where you have been granted RBAC.

For the local path, the README points at the local environment section for the full prerequisite list, which includes installing azd. The deployment itself is a single command from the repository root:

bash
azd auth login
azd up

The azd up command provisions the infrastructure defined under infra and deploys the application. The README notes that Azure Container Apps is the default host as of 2024-10-28, with App Service provisioned only if you follow the App Service deployment guide. Expect provisioning to take several minutes, and expect it to fail early if the role assignment permissions above are missing.

Once deployment finishes, the README's workflow moves to the development server for local iteration. The repository includes a locustfile.py at the top level, which indicates load testing is part of the sample's tooling, though the README does not walk through running it. When the app is up, the Zava sample data is already indexed, so you can ask about benefits, internal policies, or job descriptions and see citations come back with the answer. To remove everything the deployment created, the README provides a clean up section, which is worth reading before you start rather than after.

Where this sample stops being the right tool

The security notice is the first limitation and the project states it plainly. Code written to demonstrate a service is not the same as code hardened for untrusted input, and the README directs readers to a productionizing guide and to the Azure OpenAI Landing Zone reference architecture rather than claiming the sample is sufficient. If your requirement is a compliant deployment, this repository is a starting point for a design review, not the artifact that passes it.

The second constraint is cost and account shape. The README is explicit that pricing varies by region and usage and that exact costs cannot be predicted. The listed resources include Azure Container Apps on a consumption plan with a minimum of zero replicas, Azure Container Registry on the Basic tier, Azure OpenAI billed per 1K tokens with at least 1K tokens used per question, and Azure App Service only if you take that path. You can model this with the Azure pricing calculator link the README provides, but you cannot get a fixed number from the documentation.

Third, the project is Azure-bound by construction. Retrieval runs on Azure AI Search and generation on Azure OpenAI, so the value of the sample is the Azure wiring. If your stack is already on another cloud, or if you need to run retrieval locally with no managed service, the architecture here will not transfer without replacing both halves. The README does not present a provider-agnostic mode, and nothing in the repository layout suggests one.

Finally, the sample is opinionated about the answer format. Citations and the thought process are rendered for every answer, which is what you want when evaluating retrieval quality and noise when you want a minimal chat surface. Stripping those is your change to make, not a configuration toggle the README documents.

How it compares with a framework such as LangChain

The real alternative is not another Azure sample but a RAG framework, most commonly LangChain or LlamaIndex. The difference is in what you are handed. A framework gives you abstractions: document loaders, text splitters, vector store interfaces, retrievers, and chain composition, with the managed services plugged in behind those interfaces. This repository gives you a working application with the Azure services wired directly, plus the infrastructure templates to provision them.

That difference decides the choice. If you need to swap vector stores or model providers repeatedly, or you are prototyping across several backends, the framework's abstraction layer earns its keep and this sample will feel rigid, because the retrieval target is Azure AI Search and the generation target is Azure OpenAI. If you have already chosen Azure and your problem is that you cannot see how the pieces fit together, the sample answers a question the framework documentation does not: what does a complete, deployable Azure RAG app look like, including the Bicep templates and the role assignments?

There is a second axis, the client language. The README links JavaScript, .NET and Java samples based on this one, so a .NET team is not forced into the Python codebase to reuse the design. That is a different kind of alternative from a framework swap: same architecture, different runtime, which is often the more useful comparison when the constraint is your team's language rather than your cloud.

Maintenance, upgrade cost and the MIT licence

The repository is not archived, and the last push was on 2026-09-02. The release history is close together: 2026-07-09b upgraded the msal JS packages to 5.x, 2026-07-10 migrated to a Foundry project and upgraded the evaluation model judge to gpt-5.4, and 2026-07-17 upgraded the agentic knowledge base to GPT-5.4 and defaulted retrieval effort to minimal. That cadence tells you the sample tracks Azure service changes rather than sitting still.

It also tells you what upgrading costs. Model and platform migrations land as commits you have to reconcile with your fork, and if you have modified the ingestion code, the prompt construction, or the infrastructure templates, those are exactly the files that move. The repository carries a pre-commit configuration and a markdownlint configuration, so contributing back follows a defined path, but there is no documented upgrade procedure for downstream forks. The README does not describe how to pull upstream changes into a modified copy.

On licensing, the repository is MIT, which permits commercial use, modification and redistribution provided the copyright notice and permission notice are retained. That is the usual reading, and it is a permissive licence rather than a copyleft one. It says nothing about the Azure services the app calls, which are billed and governed by their own terms, and nothing about the sample data's suitability for your use. This is a description of the licence text, not legal advice; if the deployment touches regulated data, get the licence and the service terms reviewed by someone qualified.

Editorial conclusion

Adopt it if you want a working Azure RAG pipeline to read, fork or use as an architecture reference before building your own service. Do not adopt it as a production application: the README states the code was built to showcase Azure services and advises against putting it into production without additional security work. Before you commit, verify that your account holds Microsoft.Authorization/roleAssignments/write and Microsoft.Resources/deployments/write, since the deployment creates role assignments and will fail without them.

Frequently asked questions

Can I use Azure OpenAI for free?

The README does not claim a free tier for Azure OpenAI. It points new Azure users at a free account with credits and a guide to deploying with the free trial, and separately links the Azure pricing calculator for the resources involved.

Is Azure AI Search free to use?

The README does not state that Azure AI Search is free. It lists the resources the sample provisions with links to their pricing pages and notes that pricing varies by region and usage, so exact costs cannot be predicted from the documentation.

How do I use Azure AI Search?

In this sample you do not call Azure AI Search directly. The ingestion step writes document chunks and embeddings into an index, and each question retrieves matching chunks from that index before they are inserted into the prompt sent to the Azure OpenAI model.

Is Azure Cognitive Search the same as Azure AI Search?

The repository topics list both azure-ai-search and azurecognitivesearch, which indicates the older name still appears in the project's metadata. The README itself uses Azure AI Search throughout.

Official sources

  1. Azure-Samples/azure-search-openai-demo on GitHub
  2. License: MIT
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/azure-samples-azure-search-openai-demo.svg)](https://hysenlabs.com/projects/azure-samples-azure-search-openai-demo)