Microsoft Unveils Three New Speech AI Models: Real-Time Transcription Tops Accuracy Rankings, MAI-Voice 2.1 Supports 23 Languages

Deep News
4 hours ago

Microsoft is accelerating its efforts to strengthen speech AI capabilities. On October 1 local time, Microsoft (NASDAQ: MSFT) released three new speech models, including the real-time speech-to-text model MAI-Transcribe-2-Streaming, as well as text-to-speech models MAI-Voice-2.1 and the lower-latency MAI-Voice-2.1-Flash. Among them, MAI-Transcribe-2-Streaming claimed the top spot for accuracy on the AI benchmarking platform Artificial Analysis's streaming speech recognition test.

Data published by AI benchmarking and analysis platform Artificial Analysis shows that MAI-Transcribe-2-Streaming achieved a final transcription word error rate (WER) of 2.5%, ranking first among 38 models, and takes only about 0.13 seconds from the end of speech to returning the final text. In the "first partial transcription" test, which places greater emphasis on real-time performance, its WER was also 2.5%, with a latency of about 0.12 seconds. By comparison, the previously top-ranked SpaceXAI Grok Voice Transcribe 2.0 had a final WER of 2.7% and a latency of 0.49 seconds.

Real-time transcription surges to first place as Microsoft targets "listening and acting simultaneously"

The core model in this release is MAI-Transcribe-2-Streaming. Microsoft says the model supports 60 languages and can continuously and automatically identify languages, without having to wait until a user finishes a sentence before it starts outputting text. Unlike traditional speech-to-text models that follow a "listen—process—output" pattern, streaming models first quickly generate partial transcription results and then continuously revise them as more speech information comes in, eventually forming stable text. Microsoft says the model can begin producing its first batch of results within just over 100 milliseconds after receiving audio. This means the application scenarios for speech AI can extend beyond simply "turning sound into text" toward real-time interaction. For example, in customer service scenarios, AI agents can begin identifying needs before a user has finished speaking; voice assistants can reason ahead and prepare tool calls in advance, rather than waiting until the user finishes speaking to start the entire processing flow. Microsoft says that in applications such as real-time dictation and captions, its internal tests show that text appears about twice as fast as that of its closest competitor.

Artificial Analysis's tests also demonstrate this characteristic: MAI-Transcribe-2-Streaming's final WER was 2.5% with a latency of 0.13 seconds, both better than Grok Voice Transcribe 2.0's 2.7% and 0.49 seconds; Muse Voice Transcribe had a WER of 3.1% and a latency of 0.16 seconds, while ElevenLabs Scribe v2 Realtime stood at 3.6% and 0.14 seconds. However, speed is not an absolute lead. Artificial Analysis data shows that Cartesia Ink Preview's final transcription latency was about 0.11 seconds, slightly faster than Microsoft's model, but its WER was 3.1%, lower in accuracy than Microsoft's.

At $0.54 per hour, real-time capability does not pursue "extreme low pricing"

On pricing, Microsoft is offering MAI-Transcribe-2-Streaming at a promotional rate of $0.54 per hour through the end of the year, equivalent to $9 per 1,000 minutes. Microsoft says the model is currently in a developer-focused promotion phase. From Artificial Analysis's cross-comparison, $0.54 per hour sits at a relatively high position among mainstream streaming speech model prices: comparable to Gemini 3.5 Transcribe Live's roughly $9 per 1,000 minutes, higher than ElevenLabs Scribe v2 Realtime and Deepgram Flux at about $6.50, and also notably higher than Cartesia Ink-2 at about $4 and Muse Voice Transcribe at about $3. This also means that Microsoft is not simply trying to capture the market through low prices this time, but is instead attempting to combine accuracy, latency, and enterprise-grade deployment capabilities.

Microsoft's previously launched non-streaming MAI-Transcribe-2 in September focused on capabilities such as accuracy, speaker recognition, timestamps, and complex real-world scenarios, and currently supports 60 languages. Microsoft said at the time that the model had an average WER of 5.2% on the FLEURS benchmark and was on the Pareto frontier in Artificial Analysis's accuracy and latency tests.

Fast at "hearing" on one end, fast at "speaking" on the other

In addition to the transcription model, Microsoft also released MAI-Voice-2.1 and MAI-Voice-2.1-Flash, corresponding respectively to higher expressive quality and lower latency. Among them, MAI-Voice-2.1 supports 23 languages and 26 regions/locales. Microsoft emphasizes that the same voice can remain consistent across languages while using more natural local accents for different languages. This means developers can let an AI assistant switch among English, Chinese, German, and other languages without having to configure different voices for each language. MAI-Voice-2.1 is priced at $22 per million characters, making it more suitable for scenarios with higher requirements for voice expressiveness; MAI-Voice-2.1-Flash is aimed at applications more sensitive to latency and cost, with the price lowered to $15 per million characters. Microsoft says the Flash version improves model inference speed by 55% and costs about 60% less than comparable models. Microsoft also says both models support voice cloning based on a few seconds of reference audio and include safety measures such as consent mechanisms.

From the perspective of the product portfolio, Microsoft is in fact simultaneously compressing latency at both ends of speech AI interaction: MAI-Transcribe-2-Streaming is responsible for quickly "understanding," while MAI-Voice-2.1-Flash is responsible for quickly "speaking." Microsoft says that when combined, the two can free up more time for voice agents to reason, call tools, and check answers, while maintaining an interaction rhythm close to that of natural human conversation.

Microsoft continues to fill out its self-developed AI model matrix

This release also continues Microsoft's actions this year to expand its MAI self-developed model matrix. Microsoft has previously launched MAI-Transcribe-2, MAI-Voice-2, MAI-Thinking-1, MAI-Code-1.1-Flash, and the MAI-Image series of models, and has gradually integrated these models into Microsoft Foundry as well as products and services such as Copilot, Teams, GitHub, and Dynamics 365. Among them, MAI-Transcribe-2 currently already covers scenarios such as meeting records, video captions, customer service documentation, accessibility services, and voice agents; Microsoft's addition of a real-time streaming version this time means that the competitive focus of its self-developed models is shifting from single-item recognition accuracy to the complete chain of real-time voice agents.

For Microsoft, the significance of this direction is not just about adding a few models. As AI agents move further from text interaction toward phone customer service, voice assistants, real-time translation, and multimodal workflows, the latency, accuracy, and cost of speech recognition and speech generation will directly affect user experience and enterprise deployment costs. MAI-Transcribe-2-Streaming topping the Artificial Analysis test at least shows that Microsoft is accelerating its pursuit and entering the first tier in the speech AI segment.

Disclaimer: Investing carries risk. This is not financial advice. The above content should not be regarded as an offer, recommendation, or solicitation on acquiring or disposing of any financial products, any associated discussions, comments, or posts by author or other users should not be considered as such either. It is solely for general information purpose only, which does not consider your own investment objectives, financial situations or needs. TTM assumes no responsibility or warranty for the accuracy and completeness of the information, investors should do their own research and may seek professional advice before investing.

Most Discussed

  1. 1
     
     
     
     
  2. 2
     
     
     
     
  3. 3
     
     
     
     
  4. 4
     
     
     
     
  5. 5
     
     
     
     
  6. 6
     
     
     
     
  7. 7
     
     
     
     
  8. 8
     
     
     
     
  9. 9
     
     
     
     
  10. 10