AI in Practice

Microsoft: MAI-Transcribe-2 transcribes for ten cents per hour

3 min read

TL;DR Too Long; Didn’t read

Microsoft has released MAI-Transcribe-2, a speech recognition model for ten US cents per hour of audio. The model achieves a word error rate of 5.2 percent across over 60 languages on the FLEURS benchmark and is said to work up to ten times faster than OpenAI's GPT-Transcribe. The price is about 72 percent lower than that of the predecessor model.

A microphone with a Microsoft logo sticker on its stand emits sound waves that turn into lines of text, while a price tag reading ten cents falls beside it. Image generated with GPT Image 2

Key takeaways

  • MAI-Transcribe-2 costs ten US cents per hour of audio, about 72 percent less than its predecessor.
  • The model achieves a word error rate of 5.2 percent across over 60 languages on the FLEURS benchmark.
  • According to Artificial Analysis, it works up to ten times faster than OpenAI's GPT-Transcribe.
  • Access is available as a public preview via Microsoft Foundry, MAI Playground, and OpenRouter – still without a service level agreement.
  • The introductory price is valid only until the end of 2026, and EU availability remains officially unconfirmed.
  • The model supports speaker separation, timestamps, and code-switching between languages in a sentence.

Microsoft has released a new speech recognition model called MAI-Transcribe-2, which claims to transcribe faster, more accurately, and more cheaply than comparable models from OpenAI, Google, and ElevenLabs. One hour of audio now costs ten US cents – a decrease of about 72 percent compared to the previous model from spring. The model is now available as a public preview in Microsoft Foundry.

Benchmark measures error rate in 60 languages

On the FLEURS benchmark, which tests speech recognition in 60 languages, MAI-Transcribe-2 achieves an average word error rate of 5.2 percent, placing it at the top. The model is also ranked first in speed on the platform Artificial Analysis, which specializes in model comparisons, while it ranks second in pure error rate with 2.0 percent behind a competing model. In terms of processing speed, MAI-Transcribe-2 is said to work ten times faster than OpenAI’s GPT-Transcribe, seven times faster than ElevenLabs’ Scribe v2, and five times faster than Google’s Gemini 3.5 Transcribe – independently unverified values that come from tests by Artificial Analysis and are quoted by Microsoft in its own announcement.

For practical use, robustness against background noise and dialects matters as much as the pure recognition rate. According to the product page, additional features include speaker separation, word-accurate timestamps, and automatic language detection – characteristics that go beyond mere transcription and make the model interesting for minutes of multi-speaker meetings.

Price drops by more than two-thirds in five months

The introductory price of ten US cents per hour of audio is valid until the end of 2026 and replaces the 36 US cents that the previous model MAI-Transcribe-1 cost since its launch five months ago. Practically, this means: a complete conference with eight hours of discussions can be transcribed for about 80 US cents, and a one-hour customer call for ten cents. Microsoft leaves open how the price will develop after the transition period.

The price drop aligns with an industry-wide trend: OpenAI, Google, and ElevenLabs have been reducing costs for speech-to-text services at a similar pace for months, while accuracy is simultaneously increasing. For companies that transcribe meetings, customer calls, or interviews on a large scale, simply switching providers can significantly reduce ongoing costs. The price decline also shows individual users, who, for example, use the dictation feature in Word to capture text by voice, how quickly speech recognition is becoming cheaper in everyday life – provided that accuracy and data protection hold up in practical tests.

Access runs through Foundry, Playground, and OpenRouter

MAI-Transcribe-2 is accessible as a public preview in Microsoft Foundry as well as through the MAI Playground and the OpenRouter marketplace, with a demo available at playground.microsoft.ai. The preview status means: a service level agreement does not yet apply, and caution is recommended for production-critical applications. Microsoft does not comment on separate availability in Europe or Germany in the announcement. Since access runs through cloud services, the model should also be usable in this country, but confirmation regarding data storage in EU data centers is still pending.

In addition to speech recognition itself, the model supports keyword biasing for technical terms and selectable transcription styles – verbatim with filler words or cleaned up for minutes. The model also automatically processes code-switching between languages within a sentence, such as in bilingual meetings.

MAI-Transcribe-2 comes from the same in-house Microsoft AI team that already introduced its own MAI models in Excel and Outlook in July, replacing OpenAI and Anthropic models there. The new transcription engine continues this line and makes Microsoft more independent from external providers in speech recognition – one more building block in the growing in-house model portfolio alongside text and image generation.

It will be crucial whether Microsoft maintains or raises the introductory price after 2026, once workflows have adjusted to the low costs. It also remains to be seen how competition in transcription services will evolve when accuracy and speed are already sufficient among several providers, and price becomes the decisive criterion.

Frequently asked questions

What does MAI-Transcribe-2 cost exactly?

Until the end of 2026, Microsoft charges ten US cents per hour of transcribed audio as an introductory price. What applies thereafter has not yet been communicated by the company.

Is MAI-Transcribe-2 usable in Germany or the EU?

Microsoft does not mention any regional restrictions in the announcement; access is via internet-based cloud services. An official commitment to data storage in EU data centers is still missing.

How does MAI-Transcribe-2 compare to ChatGPT, Gemini, or Scribe?

According to tests by Artificial Analysis, it processes audio faster than OpenAI's GPT-Transcribe, ElevenLabs' Scribe v2, and Google's Gemini 3.5 Transcribe, with a similarly low error rate.

What do you need to use MAI-Transcribe-2?

Currently, access is through Microsoft Foundry, the MAI Playground, or the OpenRouter marketplace, each requiring a separate account. A demo without registration is available on the site playground.microsoft.ai.

Is the model already approved for productive use?

No, MAI-Transcribe-2 is currently running as a public preview without a service level agreement. For mission-critical applications, it is advised to test cautiously for now.

Sources (4)
  1. MAI-Transcribe-2 is the fastest, most accurate and cheapest speech recognition model in the world (Microsoft AI)
  2. MAI-Transcribe-2 (Microsoft AI model page)
  3. Speech to Text (ASR) Leaderboard (Artificial Analysis)
  4. MAI-Transcribe-2: Highest quality transcription (Azure AI Foundry Blog, Microsoft Community Hub)

Your AI update for the work week

Once a week, the most important AI news – plus one practical tip to try right away. No spam, unsubscribe anytime.

← Back to the blog