microsoft/call-center-ai: an API call that places an AI phone call
Send a phone call from AI agent, in an API call. Or, directly call the bot from the configured phone number!
At a glance
- What is it?
- Microsoft's call-center-ai is a Python service that turns a POST request into an outbound phone call handled by a GPT agent, with the transcript, structured claim and follow-up reminders stored afterwards. The catch is that the whole thing runs on Azure, so the install is a subscription deployment, not a pip install.
- Who is it for?
- Adopt it if you already run Azure and want an outbound or inbound voice agent that writes structured claim data, not just a transcript. Skip it if you need a self-hosted stack on your own hardware, if you cannot create Azure Communication Services and OpenAI resources, or if you need production support with an SLA, since the repository carries no support statement.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 91 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem: outbound calls that end in a structured record, not a transcript
Most voice bot demos stop at conversation. A human still has to listen back and type the outcome into a ticketing or claims system. call-center-ai is built around the opposite assumption: the call exists to fill a schema. You declare the fields you want collected in the request body, and the agent's job is to gather them while talking.
The README frames the audience as insurance, IT support and customer service teams, and says the bot "can be customized in few hours (really) to fit your needs." That is the pitch: an engineering team with an Azure subscription wires up a phone number and a prompt, and gets a working intake line without building telephony, speech-to-text, LLM orchestration and storage themselves. It is not a product for a call center manager with no cloud access. The unit of work here is a deployment.
The scope is deliberately narrow. The feature list mentions handling "low to medium complexity calls," with human agent fallback for the rest. That is an honest boundary. An agent that fills a claim schema is not the same thing as one that negotiates a refund.
How a POST to /call becomes a conversation, a claim and a todo list
The README's first example is the clearest description of the data flow. You send a JSON body to a /call endpoint with bot_company, bot_name, phone_number, task, agent_phone_number and a claim array. The claim array is the schema: each entry has a name and a type, and the README's example uses types text and datetime. The task string is the system-level instruction for the call, and in the example it explicitly says the assistant works for the IT support department and should gather information into the claim.
From there the pipeline is assembled from named Azure services, which the pyproject.toml dependency list makes explicit rather than hiding. azure-communication-callautomation handles the call itself, azure-cognitiveservices-speech does speech, azure-ai-inference talks to the LLM, azure-ai-translation-text provides translation, azure-search-documents backs the retrieval-augmented generation path, azure-cosmos is the store, and azure-storage-queue plus azure-eventgrid move events around. FastAPI serves the HTTP surface and granian is the server. Redis caching is mentioned in the feature list, and azure-appconfiguration is there so configuration can change without a redeploy.
The stored output shape is shown in the README's demo extract. A finished call produces claim (the filled fields), messages (an ordered list with created_at, action, content, persona, style and tool_calls), next (an action such as case_closed plus a justification), reminders (title, description, due_date_time, owner) and synthesis (long, short, satisfaction and improvement_suggestions). The synthesis field is the part worth noting: the model writes its own summary of the call and rates satisfaction, which means the summary is generated text, not a template. Treat it as a draft.
Two mechanisms in the feature list are easy to miss. Conversations are streamed in real time, which the README says is to avoid delays, and they can be resumed after disconnections. Resumption matters because phone calls drop. A stateless bot would lose the context; keeping the message list in Cosmos is what makes the resume possible.
Installing call-center-ai: an Azure deployment, not a pip install
There is no pip install path in the README. The repository layout points at deployment instead: a Makefile, a cicd/ directory, config-local-example.yaml, config-remote-example.yaml and a .devcontainer/ for GitHub Codespaces. The README offers a Codespaces quickstart badge, which is the fastest way to see the code without provisioning anything.
The Makefile defines the container image and the Azure locations used for deployment. It pins cognitive_communication_location to westeurope, default_location and openai_location to swedencentral, and search_location to francecentral, with a comment warning that some regions do not support all services or capabilities, including Cognitive Services TTS voices. Those variables are the ones to change first if you deploy outside Europe.
The Makefile also shows how the app URL is recovered after deployment, by reading a Bicep output from the subscription deployment:
az deployment sub show --name $(name_sanitized) | yq '.properties.outputs["appUrl"].value'That tells you the deployment is a subscription-scoped Bicep template, and app_url, blob_storage_public_name and container_app_name all come out of it as outputs. The image referenced is ghcr.io/clemlesne/call-center-ai with image_version set to main, so the default deploy pulls the main tag rather than a release tag. If you want the v17.4.2 release, that is a change you make yourself.
Optional credentials go in a .env file. The example lists a service principal and Application Insights:
AZURE_CLIENT_ID=xxx
AZURE_CLIENT_SECRET=xxx
AZURE_TENANT_ID=xxx
APPLICATIONINSIGHTS_CONNECTION_STRING=xxx
OTEL_TRACES_SAMPLER_ARG=0.5The README marks both blocks optional, so a deployment that relies on managed identity and skips tracing does not need them. The sampler argument of 0.5 means half the traces are kept, which is a cost control rather than a correctness one.
Once the app is running, the first real use is the call request. This is the README's own example, trimmed to the essential fields:
curl \
--header 'Content-Type: application/json' \
--request POST \
--url https://xxx/call \
--data '{"bot_company":"Contoso","bot_name":"Amélie","phone_number":"+11234567890","task":"Help the customer with their digital workplace.","agent_phone_number":"+33612345678","claim":[{"name":"hardware_info","type":"text"}]}'Replace the placeholder host with the appUrl output from the deployment. What you should get back is a call placed to phone_number, with the agent identifying itself as Amélie from Contoso and collecting hardware_info. The README's demo section shows that the conversation, claim and todo list are written to the database during and after the call.
Where the Azure dependency becomes a real constraint
The honest limitation is that this is not portable. Every major capability maps to a specific Azure service, and the dependency list in pyproject.toml is the proof. If you want to run speech-to-text locally with Whisper, or store conversations in Postgres, you are not configuring the project, you are rewriting it. Teams that need data residency outside an Azure region, or that have a policy against sending audio to a cloud speech API, should stop here.
The second constraint is the model choice. The README names gpt-4.1 and gpt-4.1-nano and states plainly that these carry a "10-15x cost premium." The nano model is presumably for cheaper steps, but the README does not document which calls use which model. That is a gap you would have to close by reading the code before you can estimate cost per call.
The third is operational. The README says the architecture is serverless and scales elastically, and it lists Application Insights for monitoring. It does not document rollback, version pinning, or what happens to in-flight calls during a deployment. The Makefile default of image_version := main means a redeploy can pull newer code than you tested. For a phone line that customers dial, that is a decision worth making deliberately rather than inheriting.
Finally, the README's roadmap mentions automated callbacks and IVR-like workflows as future items. If your requirement is a menu tree with keypad input, the current state of the project is not that.
call-center-ai versus a self-hosted voice agent stack
The obvious alternative is assembling the pieces yourself: a media stream from a SIP provider, a local or hosted speech recognizer, an LLM API, and your own state store. Projects in that space let you point the agent at any OpenAI-compatible endpoint and keep the audio path under your control.
The difference is not quality, it is who owns the integration. A self-hosted stack gives you provider choice and lets you run on your own metal, but you write the call state machine, the reconnection handling and the storage schema. call-center-ai ships those as Azure resources and a Bicep template, which is why the README can promise customization in hours. You are trading portability for a working pipeline.
The second alternative is a commercial AI call center product. Those come with support contracts and an operator console, and they do not require you to hold an Azure subscription. What they do not give you is the source, the prompt files, or the ability to change how a claim is extracted. If your intake schema is unusual, that matters more than the support line.
Licence, maintenance and what an upgrade actually costs
The project is Apache-2.0, and pyproject.toml references the LICENSE file rather than declaring a classifier. Apache-2.0 permits commercial use and modification and includes an explicit patent grant. It also requires that you keep the licence and notice files when you redistribute. The repository carries no separate support statement, so the licence is the whole of the legal picture; for anything involving recorded customer calls, the compliance question is about your own obligations under wiretapping and consent rules, not about the licence.
On maintenance, the last push was on 2026-07-01, which is recent, and the repository is not archived. The most recent tagged release is v17.4.2 from 2025-05-27, with v17.4.1 and v17.4.0 a few days earlier. That gap between the last release and the last push means the main branch is ahead of the last tag, which is consistent with the Makefile defaulting to image_version := main.
Upgrade cost is dominated by the Azure side, not the Python side. Because the deployment is a Bicep template, moving to a newer version means re-running the deployment and accepting whatever the template changes, including resource configuration. The versioning is handled by a script under cicd/version/ that the Makefile wraps as make version and make version-full, so the project tracks its own version rather than relying on a tag alone. The practical risk is a model or API version change in one of the azure-* packages, several of which are pinned to alpha or preview releases in pyproject.toml.
Editorial conclusion
Adopt it if you already run Azure and want an outbound or inbound voice agent that writes structured claim data, not just a transcript. Skip it if you need a self-hosted stack on your own hardware, if you cannot create Azure Communication Services and OpenAI resources, or if you need production support with an SLA, since the repository carries no support statement. Before committing, verify three things: that your target region supports the models and speech voices you need, that your Twilio or Azure phone number can be provisioned in that region, and what the per-minute cost of Azure Communication Services plus gpt-4.1 inference looks like against your expected call volume.
Frequently asked questions
How are call centers using AI, and where does microsoft/call-center-ai fit?
The README positions the project for insurance, IT support and customer service, handling low to medium complexity calls with human agent fallback for the rest. It answers inbound and outbound calls on a dedicated phone number, collects structured claim fields during the conversation, and stores the transcript, a synthesis and follow-up reminders.
Is AI taking over call centers in microsoft/call-center-ai?
No. The feature list includes human agent fallback, and the README scopes the bot to low and medium complexity calls rather than all of them. The stored output includes a next action such as case_closed with a justification, so the model is deciding whether a case can close, not replacing the whole function.
Which AI call center agent is the best, and how does microsoft/call-center-ai compare?
The README does not rank agents, so no comparison can be made from it. What can be said is that call-center-ai is Apache-2.0, written in Python, and requires an Azure subscription because its telephony, speech, search, storage and inference all map to named Azure services.
What is the 80/20 rule in call centers, and does microsoft/call-center-ai address it?
The README does not mention the 80/20 rule or any call-volume distribution. The closest statement is that the bot handles low to medium complexity calls, which is a scope claim rather than a routing rule.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/microsoft-call-center-ai)