Model or dataset
Azure/GPT-RAG avatar
Azure/GPT-RAG

GPT-RAG: a Microsoft Foundry accelerator for agentic RAG on Azure

Enterprise-grade accelerator for agentic RAG on Azure. Built on Microsoft Foundry with Foundry IQ as the default retrieval backend, Microsoft Agent Framework orchestration, Zero-Trust architecture and IaC.

1,178 stars323 forksPythonMIT

At a glance

What is it?
Azure/GPT-RAG is a deployment accelerator, not a library: it ships Bicep infrastructure, a Microsoft Agent Framework orchestration layer, and Foundry IQ as the default retrieval backend. The judgement below is about when that trade is worth it.
Who is it for?
Adopt GPT-RAG if your data already lives in Azure and you want the Zero-Trust topology, the Bicep templates and the Agent Framework orchestration assembled for you rather than designed from scratch. Do not adopt it if you need a vendor-neutral retrieval layer, if you have no Azure subscription, or if you want a Python package you can import into an existing service.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 1 day ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What GPT-RAG actually is, and who it is built for

The repository describes itself as a solution accelerator that provides architecture templates and deployment assets, which is a precise description of what you get. There is no single package to install. The top-level layout is telling: azure.yaml, infra, main.parameters.json, config, scripts, contracts, hosted-agent, tests and util sit alongside the Python code. That is the shape of a deployment project, not a library.

The target reader is an Azure platform team or an enterprise engineering group that has already decided to build retrieval-augmented generation on Microsoft Foundry and does not want to design the network topology, the identity model and the retrieval plumbing from zero. The README frames the scope as customer support through to decision automation, and names NL2SQL query generation, multi-source grounding and MCP tool integration as supported scenarios.

If you are a solo developer who wants to add retrieval to a FastAPI service in an afternoon, this is the wrong shape of project. The value here is in the infrastructure and the security model, and both assume an Azure subscription with enough governance to care about them.

Foundry IQ, Agent Framework, and the rollback path that matters

Retrieval is the part worth reading carefully. The README states that Foundry IQ is the default retrieval backend and that it blends Blob, Azure AI Search, Work IQ, Fabric, SharePoint, OneLake, Web and MCP sources with permission trimming. Orchestration runs on the Microsoft Agent Framework on Microsoft Foundry.

The design decision embedded in that sentence is that grounding is federated rather than centralised. Instead of ingesting every document into one index you control, the accelerator points at the systems where the data already lives and relies on permission trimming to keep results scoped to the caller. That is attractive when SharePoint and Fabric already hold the authoritative copies and duplicating them into a separate index would create a staleness and access-control problem.

It is also the source of the main operational risk. When retrieval spans eight source types, a wrong answer can originate in any of them, and the debugging surface is wider than a single index. The README acknowledges this by noting that Azure AI Search direct remains fully supported as a rollback path. The phrase rollback path is doing real work: it implies the maintainers expect some deployments to try Foundry IQ and move back. The README does not document what triggers that rollback, how to migrate an existing index, or what changes in the configuration when you switch.

Hosted conversation continuity and its fail-closed contract

The most specific technical detail in the README concerns hosted continuity. It uses Responses protocol 2.0.0 with delegated x-ms-user-identity ownership from the trusted UI BFF. Activation fails closed until the UI identity holds the exact Foundry Agent Consumer role and the GPT-RAG user-identity impersonation role directly at the individual agent scope.

Two consequences follow. First, role assignment must be direct at the agent scope, not inherited from a broader scope, which is a stricter requirement than many Azure deployments assume. Second, the hosted container receives neither role, no ownership key, and no Cosmos DB in the no-panel topology. That is a deliberate blast-radius reduction: the container that runs the agent cannot impersonate a user on its own.

The README points to contracts/README.md for role IDs, validation rules and the disabled capability fallback. That file is where the real specification lives, and anyone evaluating this accelerator should read it before estimating effort. The README itself gives the shape of the contract but not the identifiers.

Deploying it: what the documentation actually tells you to run

The README does not contain install commands. It directs readers to the documentation site, and specifically to the Deployment Guide at azure.github.io/GPT-RAG/deploy/, which it says covers Basic, Zero Trust and network-isolated deployments, preflight checks, a jumpbox workflow and container image builds. Because the README gives no command sequence, the honest starting point is the preflight step described there rather than a guessed azd invocation.

The repository does expose the usual Azure Developer CLI surface through azure.yaml and main.parameters.json at the top level, and azd-templates appears in the topic list, so the deployment model is azd-driven Bicep. Confirm the current command set against the Deployment Guide before running anything, since the README does not pin it.

What you should expect to see after a successful deployment, based on the repository layout, is the infrastructure defined in infra, the parameter set from main.parameters.json, and the agent code under hosted-agent. The README also ships a UI, shown in media/gpt-rag-homepage.png, which is where a first conversation would be exercised. Treat the deployment guide as the source of truth for flags and ordering; nothing in the README substitutes for it.

Where this accelerator is the wrong tool

The clearest limitation is portability. Foundry IQ is the default retrieval backend and the Microsoft Agent Framework handles orchestration, so the retrieval and agent layers are tied to Microsoft Foundry. If your organisation runs a multi-cloud estate or has a policy against a single-vendor AI control plane, the accelerator's main value proposition is exactly the thing you cannot use.

A second limit is the identity requirement. Fail-closed activation means that a misconfigured role assignment does not degrade gracefully, it disables the capability. Teams used to permissive defaults will find the failure mode abrupt, and the README does not describe a diagnostic path beyond the contract document.

A third is documentation depth in the README itself. It is a signpost document: it links to the documentation site for grounding sources, deployment and release notes. Anyone who reads only the README will not know how to deploy, what the role IDs are, or how to switch retrieval backends. That is a deliberate split, but it means the repository alone is not sufficient for an evaluation.

Finally, the README does not document rollback mechanics, cost, or upgrade procedures. If those are decision criteria for you, they are unanswered by the documentation shipped in the README.

How it differs from a plain Azure AI Search RAG stack

The obvious alternative is building the same thing directly on Azure AI Search with your own orchestration code, which the accelerator itself supports as the Azure AI Search direct path. The difference in approach is where grounding happens. A conventional stack ingests documents into an index you own, chunks and embeds them, and queries that index. The accelerator's default instead federates across Blob, SharePoint, Fabric, OneLake, Work IQ, Web and MCP sources through Foundry IQ and trims results by permission.

That changes the engineering work rather than removing it. With a single index you control chunking strategy, embedding model and refresh schedule. With federation you inherit the source systems' refresh behaviour and permission models, and you gain a much larger set of places for a retrieval failure to hide. The trade is duplication and staleness on one side against debugging surface and dependency on Foundry IQ on the other.

The second alternative is a framework such as LangChain or LlamaIndex with your own infrastructure. Those give you a Python library and leave the network topology, identity model and IaC to you. GPT-RAG gives you the topology and the IaC and asks you to accept its platform choices. Neither is strictly better; they optimise for different teams.

Maintenance, licence and what the release history implies

The repository is not archived and the last push was on 2026-09-09. Three releases landed on 2026-09-03: v3.8.1, v3.8.2 and v3.8.3, two of them within about half an hour of each other. That pattern suggests active iteration with quick patch follow-ups rather than a slow release train, and it is worth knowing because it means the surface you deploy can move within a week.

The licence is MIT, which is permissive and permits modification and redistribution. Two caveats that are not legal advice: the README states that the project may contain trademarks or logos and that authorised use must follow Microsoft's Trademark and Brand Guidelines, with modified versions required not to imply sponsorship or cause confusion. If you fork and rebrand, that clause is the one to read. The README also notes that third-party trademarks are subject to their own policies.

Upgrade cost is the open question. The README does not describe an upgrade procedure, and the CHANGELOG.md at the top level is where release-level change detail would live. For a deployment that provisions Azure infrastructure, an upgrade is not a pip install; it is a re-run of the deployment with changed parameters, and the README does not describe how that is validated.

Editorial conclusion

Adopt GPT-RAG if your data already lives in Azure and you want the Zero-Trust topology, the Bicep templates and the Agent Framework orchestration assembled for you rather than designed from scratch. Do not adopt it if you need a vendor-neutral retrieval layer, if you have no Azure subscription, or if you want a Python package you can import into an existing service. Before committing, read contracts/README.md for the exact Foundry Agent Consumer role IDs and the fail-closed behaviour of hosted conversation continuity, and confirm which grounding source your data actually sits in, because Foundry IQ is the default and Azure AI Search direct is the documented rollback path.

Frequently asked questions

What is GPT-RAG?

It is a solution accelerator from Azure that provides architecture templates and deployment assets for building agentic RAG on Microsoft Foundry, with Foundry IQ as the default retrieval backend and the Microsoft Agent Framework handling orchestration. It is a deployment project rather than an installable library.

Is GPT-RAG a Python package I can install?

No. The repository is structured around azure.yaml, infra, main.parameters.json and scripts, and the README points to a deployment guide rather than an install command. It deploys Azure infrastructure and containers, so the entry point is the Deployment Guide on the documentation site.

Does GPT-RAG require Azure AI Search?

Not as the default. The README states that Foundry IQ is the default retrieval backend and that Azure AI Search direct remains fully supported as a rollback path, so an Azure AI Search deployment is possible but is positioned as the fallback rather than the primary path.

What security model does GPT-RAG use?

It is built on a Zero-Trust architecture, with network access tightly governed and least-privilege communication between services. Hosted conversation continuity fails closed until the UI identity holds the exact Foundry Agent Consumer role and the GPT-RAG user-identity impersonation role directly at the individual agent scope.

Which grounding sources does GPT-RAG support?

The README lists Blob, Azure AI Search, Work IQ, Fabric, SharePoint, OneLake, Web and MCP sources, blended through Foundry IQ with permission trimming. The documentation site has a grounding sources overview that covers how to choose between them.

Official sources

  1. Azure/GPT-RAG on GitHub
  2. License: MIT
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/azure-gpt-rag.svg)](https://hysenlabs.com/projects/azure-gpt-rag)