Microsoft Foundry joins native real-time voice agent to challenge Google Gemini Live

📅 2026-09-25

Abstract:

Microsoft is further strengthening the real-time voice capabilities of the Foundry platform, providing developers with native voice agents that can directly perform voice input and voice output, making it easier for enterprises to build real-time conversational AI experiences similar to Google Gemini Live.

1790309180_microsoft_foundry_voice_agents.webp

The core change of Microsoft's update is that Foundry Agent Service begins to support native real-time voice capabilities. Different from the process of traditional AI agents that first convert speech into text, then hand it over to the text model for processing, and finally convert the text into speech again, the native speech architecture can directly allow the real-time speech model to process the user's voice input and generate a voice reply, thus reducing intermediate links and resulting delays.

For users, this means that AI agents can engage in continuous conversations more naturally. Users do not need to wait for speech recognition to complete or for the model to generate a complete text answer before starting speech synthesis. The system can be closer to the real-time communication method between people, and can handle key aspects of voice interaction such as interruption and turn judgment.

Microsoft's capabilities are built on the Voice Live API. This interface integrates functions such as speech recognition, generative AI, text-to-speech, turn detection, and interruption processing into a unified real-time speech interface, allowing developers to add real-time speech interaction capabilities to their AI agents without having to build multiple speech components separately.

For enterprises that have used Foundry Agent Service to build AI agents, this means that the original text-based agents can further obtain voice interaction capabilities, while continuing to use the original tool invocation, knowledge base, memory, content security protection, and enterprise-level data and service integration capabilities.

Microsoft particularly emphasizes that the voice agent is not just a chatbot that can "read answers", but can call tools and perform tasks during real-time voice communication. For example, companies can use this technology to build customer service systems that allow users to communicate directly with AI agents over the phone; it can also be used for barrier-free services, voice-first applications, and various internal corporate systems that require natural voice interaction.

In terms of technical architecture, Foundry currently supports two main voice agent routes. One is the traditional speech pipeline, which converts user speech into text, then gives it to the text model for inference, and finally converts the generated text into speech. The other is the native Speech-to-Speech that is focused on this time, which is the speech-to-speech architecture. The real-time speech model directly processes speech input and generates speech output.

The biggest advantage of the native voice architecture is that it reduces latency while better maintaining the tone, pauses, and communication rhythm of natural conversations. For scenarios that require continuous communication, this method is closer to a real real-time conversation than the traditional multi-stage solution of "speech recognition → text model → speech synthesis".

Microsoft is also continuing to improve the real-time interaction capabilities of the Voice Live API. Recent updates have added new native real-time speech types, and further improved features such as echo cancellation, intelligent closing detection, and parallel tool invocation. Echo cancellation is particularly suitable for voice agents running on custom hardware or non-standard audio links. It allows the system to refer to the sound actually played by the client, thereby reducing the problem of misrecognition after the AI ​​hears its own voice.

In addition to voice itself, Microsoft is also strengthening Foundry's positioning as an enterprise-level AI agent development and operation platform. Foundry Agent Service can be responsible for the life cycle management of intelligent agents and provide enterprise-level security, permission control, network isolation, operation tracking and evaluation capabilities. In this way, enterprises do not need to build an infrastructure from scratch for the operation, monitoring and voice communication of AI agents.

Microsoft's move is obviously also aimed at Gemini Live, which Google has continuously strengthened in recent years. Gemini Live has become an important product for Google to demonstrate real-time voice AI capabilities to consumers, while Microsoft emphasizes providing similar real-time voice capabilities directly to enterprise developers, allowing companies to build exclusive voice intelligence based on their own data, tools and business processes.

The positioning of the two is therefore not exactly the same. Gemini Live mainly provides directly usable AI voice experience for ordinary users, while Microsoft Foundry is more oriented towards enterprises and developers. The focus is on allowing enterprises to embed real-time voice capabilities into customer service, telephone systems, internal business and other professional applications.

Microsoft has been promoting the development of AI agents from simple chatbots to systems that can perform tasks autonomously in recent years. With the addition of voice capabilities to this system, the interaction between users and agents has gradually expanded from keyboard input to more natural real-time conversations. Not only can agents answer questions, they can also invoke tools, access enterprise data, and complete specific tasks during conversations.

Currently, Microsoft is gradually integrating capabilities such as Foundry Agent Service, Voice Live, and managed agents into the same enterprise AI platform. As these technologies mature, the infrastructure that companies need to develop and maintain themselves to build real-time voice customer service, phone agents, voice office assistants, and other AI services will be further reduced.

This also means that the competition for AI voice assistants is shifting from simply comparing models “whether they can chat” to comparing who can provide lower-latency voice interaction, more complete tool calling capabilities, more reliable enterprise-level infrastructure, and a more convenient development environment. Microsoft's addition of a native real-time voice agent to Foundry is an important step in further expanding its voice agent layout in the enterprise AI market.

Related tags

Related articles

Comments

0/500
Captcha (click to refresh)
No comments yet