Microsoft has released a new speech recognition model called MAI-Transcribe-2, which claims to transcribe faster, more accurately, and more cheaply than comparable models from OpenAI, Google, and ElevenLabs. One hour of audio now costs ten US cents – a decrease of about 72 percent compared to the previous model from spring. The model is now available as a public preview in Microsoft Foundry.
Benchmark measures error rate in 60 languages
On the FLEURS benchmark, which tests speech recognition in 60 languages, MAI-Transcribe-2 achieves an average word error rate of 5.2 percent, placing it at the top. The model is also ranked first in speed on the platform Artificial Analysis, which specializes in model comparisons, while it ranks second in pure error rate with 2.0 percent behind a competing model. In terms of processing speed, MAI-Transcribe-2 is said to work ten times faster than OpenAI’s GPT-Transcribe, seven times faster than ElevenLabs’ Scribe v2, and five times faster than Google’s Gemini 3.5 Transcribe – independently unverified values that come from tests by Artificial Analysis and are quoted by Microsoft in its own announcement.
For practical use, robustness against background noise and dialects matters as much as the pure recognition rate. According to the product page, additional features include speaker separation, word-accurate timestamps, and automatic language detection – characteristics that go beyond mere transcription and make the model interesting for minutes of multi-speaker meetings.
Price drops by more than two-thirds in five months
The introductory price of ten US cents per hour of audio is valid until the end of 2026 and replaces the 36 US cents that the previous model MAI-Transcribe-1 cost since its launch five months ago. Practically, this means: a complete conference with eight hours of discussions can be transcribed for about 80 US cents, and a one-hour customer call for ten cents. Microsoft leaves open how the price will develop after the transition period.
The price drop aligns with an industry-wide trend: OpenAI, Google, and ElevenLabs have been reducing costs for speech-to-text services at a similar pace for months, while accuracy is simultaneously increasing. For companies that transcribe meetings, customer calls, or interviews on a large scale, simply switching providers can significantly reduce ongoing costs. The price decline also shows individual users, who, for example, use the dictation feature in Word to capture text by voice, how quickly speech recognition is becoming cheaper in everyday life – provided that accuracy and data protection hold up in practical tests.
Access runs through Foundry, Playground, and OpenRouter
MAI-Transcribe-2 is accessible as a public preview in Microsoft Foundry as well as through the MAI Playground and the OpenRouter marketplace, with a demo available at playground.microsoft.ai. The preview status means: a service level agreement does not yet apply, and caution is recommended for production-critical applications. Microsoft does not comment on separate availability in Europe or Germany in the announcement. Since access runs through cloud services, the model should also be usable in this country, but confirmation regarding data storage in EU data centers is still pending.
In addition to speech recognition itself, the model supports keyword biasing for technical terms and selectable transcription styles – verbatim with filler words or cleaned up for minutes. The model also automatically processes code-switching between languages within a sentence, such as in bilingual meetings.
MAI-Transcribe-2 comes from the same in-house Microsoft AI team that already introduced its own MAI models in Excel and Outlook in July, replacing OpenAI and Anthropic models there. The new transcription engine continues this line and makes Microsoft more independent from external providers in speech recognition – one more building block in the growing in-house model portfolio alongside text and image generation.
It will be crucial whether Microsoft maintains or raises the introductory price after 2026, once workflows have adjusted to the low costs. It also remains to be seen how competition in transcription services will evolve when accuracy and speed are already sufficient among several providers, and price becomes the decisive criterion.


