Abstract:
Microsoft has launched three new models for real-time voice interaction, namely the speech-to-text model MAI-Transcribe-2-Streaming, and the text-to-speech models MAI-Voice-2.1 and MAI-Voice-2.1-Flash. Microsoft hopes developers will combine them to build voice assistants and conversational agents that can process while listening and respond instantly with natural speech.

MAI-Transcribe-2-Streaming is different from the previous MAI-Transcribe-2 that can only process recording files. It can receive continuous input audio streams and start generating staged text before the speaker has finished speaking. It will continue to be revised as more context comes, and finally submits stable results. Microsoft said that the model supports 60 languages and can continuously and automatically recognize languages; the first batch of recognition is assumed to appear slightly more than 100 milliseconds after receiving audio, and is suitable for scenarios such as voice agents, customer service, conference and classroom subtitles, and voice input. Microsoft's internal testing of dictation and subtitles shows text appears about twice as fast as the nearest competitor, a comparison based on the company's own reviews.
Artificial Analysis' streaming transcription benchmark ranks the model first for final and staged text accuracy. The published test results were that the final transcribed word error rate was about 2.5%, and the final text was given about 0.13 seconds after the end of the speech was detected. The list compared 38 models at that time, tested about 8 hours of audio, and covered three types of corpus: AA-AgentTalk, VoxPopuli and Earnings22, accounting for 50%, 25% and 25% of the comprehensive weight respectively. The lower the word error rate, the better. This result illustrates the performance of the model on this set of benchmarks and does not mean that the same accuracy can be achieved in all languages, noisy environments, and real-world applications.
MAI-Transcribe-2-Streaming is currently on sale for $0.54 per hour of audio, and the offer lasts until the end of 2026; in comparison, non-streaming MAI-Transcribe-2 was previously advertised at $0.10 per hour. The real-time version is more suitable for continuous interaction, while the batch version is for recording audio transcription. The prices of the two cannot simply be regarded as a direct comparison under the same usage method.
The new speech synthesis model MAI-Voice-2.1 supports 23 languages and 26 regional variants, allowing the same synthesized voice to switch between different languages while retaining the identity of the speaker and adopting the local accent of the corresponding language. It is priced at US$22 per 1 million characters. MAI-Voice-2.1-Flash for low-latency and large-scale requests supports the same language. Microsoft says it can generate 45 seconds of audio with an end-to-end delay of about 150 milliseconds. The model inference speed is increased by 55% and the price is about 60% lower than comparable models. It is priced at $15 per 1 million characters. The latter set of speed and price advantages are among the comparisons published by Microsoft.
Both voice models support cross-language real-time voice cloning, but only with the use of authorized and consented reference voices. Microsoft documentation lists the feature as restricted access, with a reference audio recommendation of 5 to 60 seconds, and protections against abuse. MAI-Voice-2.1 is suitable for longer content, narrations and audiobooks, while Flash is more suitable for real-time scenarios such as voice agents, assistants and customer service.
All three models can be used through Microsoft Foundry, MAI Playground and Vercel; the two Voice models have also been connected to OpenRouter and can be called through Azure Voice Live. LiveKit support is expected to be launched later. Microsoft Azure documentation currently labels these models as public preview, stating that there is no service level agreement and they are not recommended for direct use in production workloads.
Learn more:
https://microsoft.ai/news/our-first-streaming-transcription-model/
Comments