Speech AI

MAI-Transcribe-2 Adds Speaker Labels; Its Benchmark Lead Needs Replay

By Kaleido Field Staff ยท September 4, 2026

Structure helps only when attribution is right

Microsoft released MAI-Transcribe-2 in public preview on September 3 with speaker diarization, word-level timestamps, and stated coverage across 60 languages at a promotional $0.10 per audio hour through December 31. Its FLEURS and Artificial Analysis placements are useful test leads, not proof for every language, accent, speaker mix, or production recording.

Citation-ready: Microsoft released MAI-Transcribe-2 in public preview on September 3, 2026, with speaker diarization, word-level timestamps, and stated support across 60 languages.

Microsoft Foundry artwork for the MAI-Transcribe-2 speech recognition model
Image source: Microsoft Community Hub. Used for editorial coverage of speech model evidence desk.

What happened and why it matters

No. They justify a controlled comparison, while adoption still depends on language and accent coverage, speaker attribution, domain terms, noise, timestamp accuracy, latency, privacy, regional availability, and correction cost.

Official Microsoft Foundry launch post

Primary reference: Microsoft Foundry: MAI-Transcribe-2 launch. Kaleido Field checked the event date and the article's attributed facts against this source.

Source check
Source dateSeptember 3, 2026
Checked by Kaleido FieldSeptember 4, 2026, 10:05 CST
Source functioncurrent speech-model release analysis separating preview availability, language coverage, diarization, word timestamps, temporary price, first-party benchmark framing, third-party leaderboard placement, and workload reproduction

Word error rate is not speaker error

A transcript can contain the right words and still assign them to the wrong person. That changes meeting decisions, interview quotations, contact-center analysis, and regulated records even when aggregate word error looks strong.

A production evaluation should measure word error, speaker-attributed word error, diarization error, timestamp deviation, named entities, punctuation, code-switching, overlapping speech, silence, truncation, and manual correction time.

The launch price has an expiry date

Microsoft states $0.10 per audio hour through December 31, 2026. A cost comparison should preserve that date, deployment region, storage, data movement, retries, post-processing, human review, and the share of transcripts accepted without correction.

Freeze the endpoint and model version, replay the same consented audio set across vendors, and keep human reference transcripts and scoring scripts with the result.

Evidence boundary

Official product facts: preview status, Foundry and Azure Speech access, named structural features, stated language count, and temporary launch price. Microsoft-reported or launch-curated evidence: 5.2% average FLEURS word-error rate, first place on FLEURS, second place on the Artificial Analysis WER leaderboard, and a claimed accuracy-latency Pareto lead. Not established: independent reproduction of every chart, accuracy for every language or accent, diarization error, timestamp error, live-stream behavior, data-retention terms, region and quota coverage, or production reliability.

Reader briefing

Keep the source trail in view.

One concise email when a model, benchmark, or visual-intelligence claim materially changes.

FAQ

Is MAI-Transcribe-2 generally available?

No. Microsoft labels it a public preview.

What new structure does it return?

Speaker diarization and word-level timestamps.

How long does the launch price last?

Microsoft states $0.10 per audio hour through December 31, 2026.