Microsoft’s New Speech Model Leads 38-Model Accuracy Ranking

Microsoft AI has released MAI-Transcribe-2-Streaming, its first streaming speech-to-text model, with a launch date of October 1, 2026. Artificial Analysis ranks it first among 38 models for both final transcript accuracy and first partial transcript accuracy.
The model is built for conversations where audio and text move at the same time. Audio streams in continuously, while text streams back before the speaker has finished talking. MAI-Transcribe-2-Streaming supports 60 languages and uses automatic, continuous language detection, so the model can track language changes without a separate manual selection.
Microsoft says the model emits its first hypotheses, called partials, just over 100 milliseconds after receiving audio. The Microsoft team also states that internal tests show words appearing 2x faster than with its closest competitor.
Accuracy and speed in the Artificial Analysis comparison
The Artificial Analysis comparison uses the AA-WER Streaming index, which covers about 8 hours of audio. The mix includes AA-AgentTalk at 50%, VoxPopuli at 25%, and Earnings22 at 25%.
MAI-Transcribe-2-Streaming records a 2.5% WER for its final transcript at 0.13 seconds after the end of speech. Its first partial transcript also records a 2.5% WER, reaching that result at 0.12 seconds after the end of speech.
The closest listed results come from Grok Voice Transcribe 2.0, which records 2.7% WER at 0.49 seconds, and Muse Voice Transcribe, which records 3.1% WER at 0.16 seconds. Cartesia Ink-2 returns final transcripts in 0.07 seconds, but its WER is 4.0%.
That comparison puts MAI-Transcribe-2-Streaming in a notable position: it combines the lowest listed WER with fast final and first partial results. Its first partial score also matches its final transcript score, based on the Artificial Analysis figures.
Pricing, access, and integration options
Microsoft prices MAI-Transcribe-2-Streaming at $0.54 per hour of audio. That is an introductory price through the end of 2026, and Artificial Analysis normalizes it to $9.00 per 1,000 minutes.
Microsoft charges more than xAI and Meta for streaming, while the price roughly matches Google’s estimated rate. For comparison, Batch MAI-Transcribe-2 costs $0.10 per hour, making the streaming model the higher-priced option between the two Microsoft offerings.
Microsoft documents two integration paths: the Realtime API and the Azure Speech SDK. The model is available in the MAI Playground, through Vercel, and through Azure Voice Live. LiveKit support is listed as coming soon.
The model also pairs with MAI-Voice-2.1-Flash for full voice loops. MAI-Voice-2.1 covers 23 languages and 26 locales and costs $22 per 1 million characters.
What the public preview includes
MAI-Transcribe-2-Streaming is available as a public preview. The preview has no SLA and no open weights, two limits that define how the release can be used and evaluated.
For now, the release combines a 60-language streaming system, automatic language detection, and text output that begins while a speaker is still talking. Its Artificial Analysis ranking also gives developers a direct way to compare its final and first partial transcript results with other models.
The main numbers are clear: 2.5% WER for both final and first partial transcripts, 0.13 seconds to the final result, 0.12 seconds for the first partial result, and a $0.54 introductory price per hour of audio. Microsoft’s first streaming speech-to-text model enters the market with accuracy, timing, and integration details already defined.
Based on




