Speech AI
MAI-Transcribe-2 Adds Speaker Labels; Its Benchmark Lead Needs Replay
Microsoft released MAI-Transcribe-2 in public preview on September 3 with speaker diarization, word-level timestamps, and stated coverage across 60 languages at a promotional $0.10 per audio hour through December 31. Its FLEURS and Artificial Analysis placements are useful test leads, not proof for every language, accent, speaker mix, or production recording.
Citation-ready: Microsoft released MAI-Transcribe-2 in public preview on September 3, 2026, with speaker diarization, word-level timestamps, and stated support across 60 languages.

What happened and why it matters
No. They justify a controlled comparison, while adoption still depends on language and accent coverage, speaker attribution, domain terms, noise, timestamp accuracy, latency, privacy, regional availability, and correction cost.
Official Microsoft Foundry launch post
Primary reference: Microsoft Foundry: MAI-Transcribe-2 launch. Kaleido Field checked the event date and the article's attributed facts against this source.
| Source date | September 3, 2026 |
|---|---|
| Checked by Kaleido Field | September 4, 2026, 10:05 CST |
| Source function | current speech-model release analysis separating preview availability, language coverage, diarization, word timestamps, temporary price, first-party benchmark framing, third-party leaderboard placement, and workload reproduction |
Word error rate is not speaker error
A transcript can contain the right words and still assign them to the wrong person. That changes meeting decisions, interview quotations, contact-center analysis, and regulated records even when aggregate word error looks strong.
A production evaluation should measure word error, speaker-attributed word error, diarization error, timestamp deviation, named entities, punctuation, code-switching, overlapping speech, silence, truncation, and manual correction time.
The launch price has an expiry date
Microsoft states $0.10 per audio hour through December 31, 2026. A cost comparison should preserve that date, deployment region, storage, data movement, retries, post-processing, human review, and the share of transcripts accepted without correction.
Freeze the endpoint and model version, replay the same consented audio set across vendors, and keep human reference transcripts and scoring scripts with the result.
Evidence boundary
Official product facts: preview status, Foundry and Azure Speech access, named structural features, stated language count, and temporary launch price. Microsoft-reported or launch-curated evidence: 5.2% average FLEURS word-error rate, first place on FLEURS, second place on the Artificial Analysis WER leaderboard, and a claimed accuracy-latency Pareto lead. Not established: independent reproduction of every chart, accuracy for every language or accent, diarization error, timestamp error, live-stream behavior, data-retention terms, region and quota coverage, or production reliability.
FAQ
Is MAI-Transcribe-2 generally available?
No. Microsoft labels it a public preview.
What new structure does it return?
Speaker diarization and word-level timestamps.
How long does the launch price last?
Microsoft states $0.10 per audio hour through December 31, 2026.