Azure-Samples AI-Gateway: Hands-On Labs for Azure API Management as an AI Gateway
Labs to explore AI Models, MCP servers, and Agents with the AI Gateway powered by Azure API Management and Microsoft Foundry 🚀
At a glance
- What is it?
- AI-Gateway is a collection of more than 30 Jupyter notebook labs that demonstrate how to use Azure API Management as an enterprise AI gateway for controlling access to large language models, MCP tool servers, and agentic applications. Each lab includes Bicep infrastructure templates and ready-to-deploy APIM policies.
- Who is it for?
- AI-Gateway is the right resource for Azure developers who need to understand how to put API Management in front of language model endpoints and MCP servers, and who want working notebook-to-deployment code rather than conceptual documentation. It is not useful to teams outside Azure or to anyone who wants a gateway that runs on-premises without a cloud subscription.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 14 days ago.
- What is it written in?
- Mainly Jupyter Notebook, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What the repository is and why AI gateways matter
An AI gateway sits between your application and the language model APIs it calls. Without one, each application team manages its own API key, rate limiting, cost controls, and logging independently. A shared gateway centralizes those concerns: one place where token consumption is metered, where rate limits are enforced, where requests are logged for compliance, and where failover between model backends is configured.
The AI-Gateway repository translates that concept into runnable Azure code. Each lab is a Jupyter notebook that walks through a specific gateway pattern, provisions the required Azure infrastructure with a Bicep template, and produces an APIM policy file that can be deployed to a real subscription. The result is a working example, not a slide deck.
The three categories of labs reflect three layers of an AI application: managing access to language models (rate limiting, caching, routing), enabling tools via the MCP protocol, and orchestrating multi-model agentic systems. A team building any one of these layers can go directly to the relevant lab.
What is Azure AI gateway in this context
The AI Gateway capability referenced throughout the repository is the set of features in Azure API Management designed specifically for LLM traffic. These include token-based rate limiting (counting tokens rather than HTTP requests), semantic caching (using vector similarity to return a cached response for semantically equivalent queries), and load balancing across multiple model backend endpoints.
The repository documents five labs for LLM management: backend pool load balancing, token rate limiting, semantic caching, model routing by version and model name, and a FinOps framework for budget controls and automated quota management.
These are APIM policies expressed in APIM's XML policy language, deployed via Bicep. The notebooks step through each policy's effect, test it against a deployed endpoint, and show how to verify it is working. This is the difference between AI-Gateway and reading the Azure documentation: the notebooks run the test for you and show expected output.
Setting up the environment and opening a lab
The prerequisites are Python 3.12 or higher, the uv package manager, VS Code with the Jupyter extension, an Azure subscription with Contributor and RBAC Administrator roles, and the Azure CLI authenticated to that subscription.
Setup from a local machine:
git clone https://github.com/Azure-Samples/AI-Gateway.git
cd AI-Gateway
uv sync
uv pip install -r pyproject.toml
code .When opening a notebook in VS Code, select the .venv interpreter created by uv sync as the Jupyter kernel. Each lab also has its own pyproject.toml for labs with extra or pinned dependencies; run uv sync inside the lab directory for those.
The repository also supports GitHub Codespaces, accessible from the README badge. Codespaces provisions a pre-configured cloud development environment with all dependencies installed. This is the fastest path to running the first notebook if the Azure CLI authentication and Python setup are already handled.
Installing uv on Linux or macOS:
curl -LsSf https://astral.sh/uv/install.sh | shMCP and agent labs: what they demonstrate
Four labs address MCP integration. The Model Context Protocol lab shows how to use APIM as a gateway for MCP servers, with OAuth credential management handled at the gateway layer. The MCP Client Authorization lab implements the MCP client authorization flow specifically. The Realtime Audio plus MCP lab combines the OpenAI real-time voice API with MCP tools.
Four agent labs build on top of the model and tool layers. The AI Agent Service lab uses the Azure AI Foundry Agent Service. The OpenAI Agents SDK lab runs OpenAI Agents against Azure OpenAI with APIM-managed tools. The Gemini MCP Agents lab integrates Google Gemini models with MCP tools through APIM. The A2A Enabled Agents lab demonstrates agent-to-agent communication using the A2A protocol alongside MCP tools.
This progression reflects a real architecture decision path: start with model access control, add tool capabilities via MCP, then build agentic flows on top. The labs are independent, so a team that only needs the MCP authorization pattern can go directly to that lab without working through the model management labs first.
Developer tooling: Copilot Agent Skills for building new labs
The repository includes a .github/skills/ directory with Copilot Agent Skills for VS Code. These are not user-facing instructions; they are skills that GitHub Copilot can invoke to help developers create new labs using the repository's conventions.
Six skills are documented: lab-creator for scaffolding notebooks with Bicep and policies, apim-bicep for generating Bicep templates, apim-terraform for Terraform equivalents, apim-policies for APIM XML policies, apim-kql for KQL queries targeting APIM logs, and mcp-builder for MCP server integration.
The README shows a sample Copilot prompt that creates a new lab called multi-model-failover covering circuit breaker patterns and retry policies. This is an unusual feature for a sample repository: the repository is designed to generate more of itself using AI. The prompt asks for a backend pool with priority-based routing, a retry policy with exponential backoff, a circuit breaker pattern for unhealthy backends, and built-in LLM logging across all backends. The result would be a new notebook, Bicep template, and APIM policy file generated by the Copilot agent using the existing labs as templates.
The tools/ folder provides five standalone utilities for testing and development: a tracing notebook for AI Foundry API calls with tracing enabled, a streaming test notebook for verifying streaming responses, a rate limit tester for validating rate limiting configurations, an OpenAI API mock server for local development that avoids hitting live endpoints during testing, and an OAuth client notebook for testing authentication flows. These tools support iterative lab development without requiring full Azure deployments for every change.
Scope, limitations, and what the labs require
Every lab requires an active Azure subscription. The Bicep templates create real Azure resources that incur cost. The labs are not emulated locally; they provision APIM instances, Azure OpenAI deployments, Azure AI Foundry services, or other Azure services depending on the lab's subject. The Contributor and RBAC Administrator roles are required because the templates create role assignments.
The repository pins azure-cli>=2.66 in pyproject.toml. Older Azure CLI versions may produce errors when running the Bicep deployments. The mcp package is pinned at mcp==1.21.2 in the root environment to avoid a conflict with the autogen packages, which pin different MCP version ranges.
Azure API Management is the only gateway the labs cover. Teams using a different gateway product such as Kong, Cloudflare Workers AI Gateway, AWS API Gateway, or Apigee will find the policy concepts transferable in principle, but will need to rewrite the APIM XML policies in their own gateway's configuration format. The labs assume APIM as the control plane throughout, and the Bicep templates are specific to Azure Resource Manager.
The AGENTS.md file in the repository root documents the coding agent conventions used during development, including tool call patterns and worktree management. It is included because the repository itself was partially built using AI coding agents, consistent with its own subject matter.
The last push to the repository was on 2026-09-16. The project is licensed MIT. All 30-plus labs are linked from the landing page at aka.ms/ai-gateway/labs. Each lab is independent, with its own Bicep template and notebook, so there is no required order for working through them. The .python-version file pins the Python version for the repository, and the pyproject.toml uses a uv workspace configuration that explicitly excludes each lab directory so that individual labs can manage their own dependency pins without conflicting with the root environment.
Editorial conclusion
AI-Gateway is the right resource for Azure developers who need to understand how to put API Management in front of language model endpoints and MCP servers, and who want working notebook-to-deployment code rather than conceptual documentation. It is not useful to teams outside Azure or to anyone who wants a gateway that runs on-premises without a cloud subscription. Before starting, confirm you have Contributor and RBAC Administrator roles on your Azure subscription, since the Bicep templates create role assignments. The GitHub Codespaces option skips local setup entirely but requires a GitHub account with Codespaces enabled.
Frequently asked questions
What is an AI gateway in Azure?
In the context of this repository, an AI gateway is Azure API Management configured to manage LLM traffic with token-based rate limiting, semantic caching, load balancing across model backends, and MCP protocol support. The Azure documentation for these capabilities is at learn.microsoft.com/azure/api-management/genai-gateway-capabilities.
What are AI gateways used for?
An AI gateway centralizes rate limiting, cost controls, logging, and request routing for LLM API calls across an organization. The labs in this repository demonstrate token rate limiting, semantic caching, multi-model load balancing, and MCP tool authorization as specific gateway functions.
Does running the AI-Gateway labs cost money?
The README does not document costs, but the labs provision real Azure resources including APIM instances and Azure OpenAI deployments. An Azure subscription with Contributor and RBAC Administrator roles is a prerequisite. Azure resources incur charges based on usage.
Can the AI-Gateway labs run without an Azure subscription?
No. The labs deploy Bicep templates to a real Azure subscription and require the Azure CLI authenticated to that subscription. The tools/ directory includes a mock OpenAI server for local testing, but the full lab flow requires Azure.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/azure-samples-ai-gateway)