Microsoft AI ships MAI-Transcribe-1, MAI-Voice-1, and MAI-Image-2 — Mustafa Suleyman announcement
Mustafa Suleyman (CEO, Microsoft AI) announces three closed-API models shipped by the MicrosoftAI team within a few months, now all available on Microsoft Foundry: MAI-Transcribe-1 (speech-to-text, claimed most accurate on FLEURS WER across 25 languages), MAI-Voice-1 (TTS, claimed new bar for natural speech), and MAI-Image-2 (text-to-image, claimed top-3 family on the Arena.ai leaderboard). The tweet is a product-launch announcement with no benchmark numbers beyond the FLEURS-WER claim for Transcribe and the arena.ai placement claim for Image-2 — no architecture, no weights, no tech report on any of the three.
Key claims
Section titled “Key claims”- MAI-Transcribe-1 is announced as the most accurate transcription model across 25 languages on the FLEURS WER benchmark [tweet body].
- MAI-Voice-1 is announced as setting a “new standard” for natural speech, with no numbers [tweet body].
- MAI-Image-2 is announced as a top-3 model family on arena.ai (T2I leaderboard) — restating the claim already filed in Introducing MAI-Image-2: for limitless creativity [tweet body].
- All three are positioned as shipped within “a few months” by the MicrosoftAI team and now available on Microsoft Foundry [tweet body].
Method
Section titled “Method”Not disclosed. This is a CEO product announcement on X with an attached promo video; no architecture, parameter counts, training-data description, or evaluation protocol beyond the FLEURS-WER and arena.ai references.
Results
Section titled “Results”Two numerical-class claims, both leaderboard-style without raw numbers:
- FLEURS WER #1 across 25 languages for MAI-Transcribe-1 [tweet body].
- Arena.ai top-3 model family for MAI-Image-2 [tweet body] — matches the rank reported in the Microsoft AI launch post for MAI-Image-2 Introducing MAI-Image-2: for limitless creativity.
No numbers for MAI-Voice-1.
Why it’s interesting
Section titled “Why it’s interesting”Fills out the Microsoft AI release surface: the MAI-Image-2 launch was already filed via the Microsoft AI blog (Introducing MAI-Image-2: for limitless creativity) but the MAI-Transcribe-1 and MAI-Voice-1 launches were not on the wiki — this tweet is the only Luma-side pointer to those two so far. The pattern matches the closed-but-API-accessible Foundry cohort that the Open foundation-model releases page tracks as a comparison baseline; specifically, MAI-Transcribe-1’s FLEURS-WER positioning is directly comparable to the open Cohere Transcribe — open-source speech-to-text model (announcement) Cohere Transcribe and Voxtral TTS — Mistral AI frontier open-weight text-to-speech model (announcement) Voxtral TTS announcements from the same week, and MAI-Voice-1 sits in the same product slot as the open Qwen3-TTS Family is Now Open Sourced: Voice Design, Clone, and Generation TTS release and the closed PersonaPlex: Voice and Role Control for Full Duplex Conversational Speech Models PersonaPlex.
See also
Section titled “See also”- Introducing MAI-Image-2: for limitless creativity — the Microsoft AI launch post for the MAI-Image-2 component of this bundle
- Open foundation-model releases — Microsoft’s MAI Foundry cohort as a closed-API baseline against the open releases tracked there
- Cohere Transcribe — open-source speech-to-text model (announcement) — Cohere Transcribe: directly comparable open-source STT model launching in the same window
- Voxtral TTS — Mistral AI frontier open-weight text-to-speech model (announcement) — Voxtral TTS: open-weight TTS counterpart to the closed MAI-Voice-1
- Qwen3-TTS Family is Now Open Sourced: Voice Design, Clone, and Generation — Qwen3-TTS family: open counterpart to MAI-Voice-1 with full multi-backend serving
- PersonaPlex: Voice and Role Control for Full Duplex Conversational Speech Models — PersonaPlex: NVIDIA’s closed-API full-duplex conversational speech counterpart