Skip to content

Repository files navigation

Voice Agent in Foundry Agent Service - Community Hub

Voice Agent in Foundry Agent Service offers below key values to customers:

Build and launch an enterprise-ready voice agent in under two minutes. Choose speech-to-speech or cascaded pipelines powered by OpenAI, Microsoft AI, and Azure real-time models. Extend your agent with Foundry tools, monitor and measure its performance, and connect inbound and outbound calls through Teams and Twilio. Create engaging conversations with voices optimized for call centers and lifelike avatars.

This repository is the official community hub for Voice Agent in Foundry Agent Service. Here you'll find:

🐛 Report Issues — File bugs, feature requests, and feedback via GitHub Issues

📚 Resources — Curated links to docs, videos, blogs, and community content for Voice Agent

🧪 Samples — Hands-on samples and extended solutions

What's New New

Voice Agent is now available in public preview, making voice a first-class modality in Foundry Agent Service. We've received great feedback from customers and are working with them to move their voice agents into production as we prepare for general availability (GA).

Voice Agent and Voice Live

Voice Live is the foundation for real-time voice interaction, and Voice Agent builds on that foundation to deliver complete voice-first agents. Voice Live handles the real-time voice experience, while Voice Agent adds the intelligence, tools, knowledge, and orchestration required to build end-to-end agentic applications.

For most customers building voice-first agents, we recommend starting with Voice Agent. Voice Agent provides the more complete, integrated experience for building and operating an agent, bringing together voice, reasoning, instructions, knowledge, tools, and orchestration. Voice Live is a better fit when customers already have their own agent stack and primarily need real-time voice capabilities with greater control over the voice application architecture.

Quick Links

Resource Description
Product home page Explore Foundry Agent Service, including voice capabilities.
New Foundry portal — Create & manage agents Create and manage voice agents in the new Foundry portal.
Documentation & quickstart Create a voice-based prompt agent, customize its instructions, and test spoken conversations.
Configure a voice agent Configure your voice agent's behavior and voice settings.
Use a hosted agent as the conversation engine Use a hosted agent for conversation logic and tools, while Voice Live handles speech, turn-taking, and interruptions.
Use a subagent in a voice-based agent Delegate specialized requests to prompt or hosted subagents in the same Foundry project.
Voice Agent tracing, monitoring & evaluation Monitor voice sessions and evaluate conversation transcripts using datasets, traces, or simulations.
Pricing & billing Understand Voice Agent pricing and billing.
Foundry Voice Agent samples Python · Java · JavaScript · TypeScript · C# / .NET · Browser voice console
Extended Voice Agent samples Explore runnable samples and extended solutions.
Voice Agent integration with Teams Teams meeting delegate sample connecting a Foundry Voice Agent and avatar to Teams through Azure Communication Services.
Regions & quotas Check the shared Foundry Agent Service region and quota documentation.
Contact the Voice Agent team Email voiceagent@microsoft.com with questions and feedback.
Python SDK azure-ai-projects for agent management and realtime voice sessions.
Java SDK com.azure:azure-ai-agents for voice sessions, conversations, and telephony.
JavaScript / TypeScript SDK @azure/ai-projects for agent management and realtime voice sessions.
C# / .NET SDK Azure.AI.Projects.Agents for voice agent definitions and management.

In the Foundry portal, open your project, go to Build > Agents, select Build an agent, and choose Voice as the interaction mode.

Voice Agent is now available in public preview. Customers can start building voice agents and preparing for production.

Explore Voice Agent in Foundry Agent Service for voice agent guides, portal links, and runnable samples.

Voice Agent Videos

Short demo (720p, 1:38)

voice-agent-short.mp4

Full walkthrough (1080p, 15:11)

voice-agent-foundry-demo.mp4

Voice Agent Blogs

Voice Agent Customer Stories

Voice Agent Quality And Scalability

Latency

Speech-to-speech models can achieve latency as low as 500 ms on service side. Cascaded pipelines (STT → LLM → TTS) can also deliver low latency < 1s with the right LLM and streaming speech configuration. Actual latency depends on the model, region, network conditions, and tool calls.

Speech and model quality

Voice Agent brings together advanced speech and language models, including GPT Realtime, GPT Live, Azure speech-to-text and HD voices, and MAI Transcribe and MAI Voice. Choose the combination that best fits your languages, domain, and conversational experience. For more resources and samples for Azure Speech HD / MAI voices, visit the Azure Text-to-Speech repository.

Recent speech benchmark references:

Customization

  1. Speech-to-speech — Azure Realtime: Custom voices are available upon request. Contact voiceagent@microsoft.com to discuss access. See the Azure Realtime voice configuration reference for the publicly documented native voice settings.
  2. Cascaded pipelines — speech recognition and synthesis: Adapt recognition to your vocabulary and audio conditions with Azure Custom Speech. Create a distinctive voice with Azure Custom Voice, or use MAI Voice customization to create a voice from a short reference recording through the documented access and consent workflow.

Scalability

Voice Agent can scale to meet your business needs. Customers are already using it to run more than 3,000 concurrent sessions and are continuing to scale. For large-scale deployments with high concurrency requirements, contact us to discuss your capacity and scalability needs.


Voice Agent Extended Samples

Start here

This directory is the entry point for the Voice Agent portal, runnable samples, the Finance reference workflow, and its shared MCP service. Detailed setup and operation instructions live with the component that owns them; this file routes users and coding agents to the correct entry point.

See Platform support for workflow-specific operating-system requirements.

For an unqualified request such as "run the UI" or "start the portal", use portal/. It is the general Voice Agent UI and runs on http://127.0.0.1:9527 by default. Do not start the Finance MCP stack unless the request mentions Finance, Templates, or the shared MCP.

Portal + local MCP quickstart

Use this path when the portal must publish or run the checked-in Templates. It is the recommended end-to-end workflow:

cd /path/to/voice-agent
az login
./scripts/setup-local-examples.sh \
   --project-endpoint "https://<account>.services.ai.azure.com/api/projects/<project>"
./scripts/login-devtunnel.sh
./scripts/setup-local-examples.sh --check
./scripts/manage-local-mcp-and-ui.sh restart

login-devtunnel.sh uses GitHub device-code authentication. Follow the printed https://github.com/login/device prompt. Azure operations continue to use the separate identity selected by az login.

Success requires local_mcp_and_portal=ready. Open http://localhost:18098, then verify with:

./scripts/manage-local-mcp-and-ui.sh status
curl -fsS http://127.0.0.1:18003/healthz
curl -fsS http://127.0.0.1:18098/healthz

Port 9527 is only for a portal-only session without the managed local MCP lifecycle. Do not run a second manual portal after this quickstart.

Choose an entry point

Goal Start here What it owns
Run the general Voice Agent UI portal/README.md Agent editor, YAML version editing, Templates, voice playground, and standalone WebRTC page
Run Python or .NET samples samples/README.md Common Python setup, microphone samples, REST lifecycle, IQ, Toolbox, local functions, downloads, and the C# sample
Run realtime speech transcription samples/realtime_stt/README.md MAI Transcribe 2 or Azure Speech with VAD, microphone or WAV input, disabled LLM responses, and reported session usage
Run Data Zone and regional realtime voice agents samples/realtime_datazone_voice_agent.md East US 2 gpt-realtime-2.1-datazone with US inference, Japan East azure-realtime with confirmed GPU routing, and eligible Central India gpt-realtime; model-specific voices and microphone setup
Run the GPT Live terminal sample samples/gpt_live/README.md Create or reuse a GPT Live agent, stream microphone audio with the OpenAI SDK, and view independently scrollable GPT Live, delegation, and user transcripts
Run the complete Finance workflow docs/README.md Ordered subscription, MCP, sample, portal, and debugging guides
Work on or deploy the Finance MCP shared_mcp/README.md Shared MCP image, Finance routes, local Dev Tunnel hosting, and Azure Container Apps deployment
Create an IQ + voice + avatar Agent samples/create_agent_with_iq_avatar_voice/README.md Portal-first Andrew Dragon HD, Harry Business, Knowledge IQ, and optional Python creation
Inspect SDK package information dist/README.md Public Python and .NET SDK dependencies and historical preview build records
Use the coding-agent workflows skills/ Voice Agent creation, IQ/Toolbox provisioning, and local-session debugging

Instructions for coding agents

When the user asks to run or debug something from this directory:

  1. Verify the environment follows Platform support for the selected workflow before running commands.
  2. Select the component from the table above and read its README.md before running commands.
  3. Treat UI without a qualifier as the general portal/.
  4. Treat Finance UI, portal Templates with Finance, or shared MCP UI as the workflow documented in docs/03_run_samples.md.
  5. Reuse an existing component-local .env and virtual environment when they are valid. Never copy credentials or endpoints between unrelated .env files without the user's intent.
  6. Start servers as long-running processes, verify their /healthz endpoint, and report the browser URL and log location. Do not report success merely because a process was spawned.
  7. Do not silently fall back to mock data or a different Azure Project when authentication, endpoint, model, or preview checks fail.

For the default portal route, follow portal/README.md to prepare portal/.env and its .venv, start portal/demo_server.py, then verify:

GET http://127.0.0.1:9527/healthz

For the complete Finance route, follow the ordered docs/README.md workflow. The normal lifecycle commands are:

./scripts/setup-local-examples.sh --project-endpoint "https://<account>.services.ai.azure.com/api/projects/<project>"
./scripts/manage-local-mcp-and-ui.sh restart
./scripts/manage-local-mcp-and-ui.sh status
./scripts/manage-local-mcp-and-ui.sh stop

The setup command installs Node.js 22 and the Dev Tunnel CLI under the ignored .local-mcp-and-ui/tools/ directory when they are missing. It also supports minimal Linux Python installations that provide venv but omit ensurepip: each component environment receives its own bootstrapped pip. No sudo or system package change is required for those three tools. Azure CLI must already be installed and authenticated; Dev Tunnel device-code authentication remains an explicit user action because setup must not authenticate as an identity the user did not choose.

Platform support

On Windows, WSL2 is required for the combined portal + local MCP automation:

  • scripts/setup-local-examples.sh: installs Linux tools and prepares the stack.
  • scripts/manage-local-mcp-and-ui.sh: starts, stops, and monitors Linux processes.

Run this workflow from a checkout in the WSL Linux filesystem, such as ~/src/, not /mnt/c/, and authenticate Azure CLI inside WSL.

Other standalone Python and .NET samples can run natively on Windows. Follow each sample's prerequisites and commands. Bash-only wrappers still need Bash/Unix tooling or equivalent PowerShell commands.

Directory map

Path Purpose
portal/ General local Voice Agent portal and WebRTC UI
samples/ Python, .NET, and Finance samples
docs/ Finance architecture, setup, operation, and debugging guides
shared_mcp/ Shared Finance MCP source, native runtime, optional container, IaC, and deployment scripts
scripts/ Finance local setup and process lifecycle entry points
dist/ Public SDK package information and historical preview build records
skills/ Reusable coding-agent workflows
tests/ Offline contracts for the common Projects SDK samples

Security and environment boundaries

  • Configure only Azure resources and identities the customer is authorized to use.
  • Never place access tokens, API keys, connection secrets, or customer data in source files.
  • Use Foundry Project connections or environment-based credentials.
  • The portal and samples use Azure Foundry; running the local portal does not run the Azure voice service locally.
  • Review each component's persistence and recording behavior before using sensitive prompts, audio, transcripts, or tool output.

Voice Live Related Links

Resource Description
Voice Live samples Referrence only. Voice Live Quickstarts and examples for Python, C#, Java, and JavaScript/TypeScript, including Foundry agent integration, MCP tools, and avatars. We recomend starting with Voice Agent now instead of Voice Live
Call Center Voice Agent Accelerator Referrence only. A template for speech-to-speech call center agents using Voice Live, with browser and telephony integration and deployment to Azure Container Apps. We recomend starting with Voice Agent now instead of Voice Live. Now Voice Agent has supported Twillio and Teams Phone and we will add more support soon.
Voice Live overview Referrence only. Learn about real-time voice interactions with speech recognition, language models, and speech synthesis.
Hosted agents overview Referrence only. Learn how to deploy custom agent applications on managed infrastructure.
Voice Live with hosted agents Referrence only. Add voice interaction to hosted agents using the Responses or HTTP Invocations protocol; includes Python client and agent examples.
Voice Live Bridge sample Deploy a hosted conversation engine and a voice wrapper, with Voice Live handling speech and interruptions.
Build a voice agent with invocations_ws Referrence only. Build hosted voice agents over WebSocket with Voice Live, Pipecat, or LiveKit.
invocations_ws voice agent samples Referrence only. Explore hosted voice agent code using WebSocket audio streaming or signaling for separate media transport.