project
MAI-Voice-2 - Microsoft's next-generation text-to-speech model
MAI-Voice-2 is Microsoft's next-generation text-to-speech (TTS) model, and is Microsoft's most expressive and natural-sounding speech synthesis model to date. Compared to its predecessors, it excels in fidelity, language coverage, speaker consistency, and emotional range...
What is MAI-Voice-2?
MAI-Voice-2 is Microsoft's next-generation text-to-speech (TTS) model, and is Microsoft's most expressive and natural-sounding speech synthesis model to date. Compared to its predecessor, it has been comprehensively improved in terms of fidelity, language coverage, speaker consistency, and emotional range. It supports 15+ languages and features fine-grained emotion control, zero-sample speech cloning, and code switching capabilities.
Main functions of MAI-Voice-2
-
Multilingual natural synthesisExpanding from English only to 15+ languages while maintaining the same naturalness and expressiveness.
-
Fine-grained emotion controlPrecisely control vocal emotions through emotional tags (such as sadness, whispering, excitement, confusion, etc.).
-
Zero-sample speech cloningThe target voice can be cloned with only 5-60 seconds of reference audio, and all languages are supported.
-
Speaker's identity is stableMaintain consistent speaker characteristics across long content, including audiobooks, podcasts, and lectures.
-
Natural code switchingIt supports natural mixing of Hindi-English, Spanish-English, and other languages without losing rhythm and identity consistency.
-
Role-playingSupports specific role styles such as motivational coaches and sports commentators.
Technical Principles of MAI-Voice-2
- Self-developed speech basic model architectureMAI-Voice-2 is built on Microsoft's in-house developed speech model and employs an end-to-end neural network speech synthesis architecture. The model can holistically understand input text, automatically adapting to intonation, emotion, and speaking style, generating human-like speech without requiring extensive manual parameter tuning by developers. Its architecture is similar to Azure Neural HD speech, achieving generational improvements in expressiveness, language coverage, and speaker consistency.
- Multilingual unified modelingMAI-Voice-2 expands upon the English monolingual model of MAI-Voice-1 into a unified multilingual speech synthesis system supporting 15+ languages. The model undergoes in-depth optimization for different phonological systems, including tonal languages, pitch-stressed languages, stress-timing languages, and syllable-timing languages, ensuring that each language achieves the same output quality as English in terms of naturalness and expressiveness.
- Zero-sample speech cloning (Voice Prompting)The model supports zero-shot speech cloning, requiring only 5–60 seconds of reference audio to extract speaker identity features and transfer them to the target language, without the need for fine-tuning or retraining for specific speakers. Based on voice prompting technology, the system extracts speaker embeddings through a reference audio encoder, maintaining consistency in timbre, intonation, and prosodic features during synthesis.
How to use MAI-Voice-2
-
Azure Foundry access: Directly call the MAI-Voice-2 API through the Azure Foundry platform.
-
Custom brand voiceYou can create a custom sound by uploading a 5-60 second reference audio file, without the need for retraining or fine-tuning.
-
Emotional labeling controlAdd emotion tags to the request to adjust the emotional style of the output speech.
-
Authorization applicationThe voice cloning function requires authorization, and the system only supports licensed voices for use in production environments.
MAI-Voice-2's core advantages
-
Leading in sound qualityIn blind testing, users preferred the previous generation MAI-Voice-1 in 72% of cases.
-
It's hard to tell the real from the fake.The speakers are extremely similar, making it difficult to distinguish the synthesized speech from the real person's recording.
-
Safety and complianceA system-level mandatory consent mechanism ensures that only authorized voice clones are allowed in the production environment, preventing unauthorized misuse.
-
Long text stabilityMaintaining consistent speaker identity and voice quality throughout hours of content.
-
Low-barrier cloningNo professional recording studio or large amount of training data is needed; sound can be reproduced in just a few seconds of audio.
MAI-Voice-2 project address
- Project official websitehttps://microsoft.ai/news/mai-voice-2expressive-speech-in-10-languages/
Comparison of MAI-Voice-2 with similar competing products
| Comparison Dimensions | MAI-Voice-2 | Gemini 3.1 Flash TTS |
|---|---|---|
| Developer | Microsoft (AI) | Google DeepMind |
| Release time | June 2026 | April 2026 (Public Preview) |
| Language support | 15+ languages, including code switching (Hindi-English, Spanish-English). | 70+ languages, covering a wider range of languages |
| Preset Sounds | The exact number was not disclosed, with a focus on brand customization. | 30 named sounds (Kore, Puck, Charon, etc.) |
| Emotional control | Fine-grained SSML tags (sadness, whispering, excitement, confusion, etc.) | 200+ inline audio tags ([sigh],[laughing],[whispering] (etc.), supports natural language prompts |
| Voice cloning | 5–60 seconds zero-shot, full language support | Not supported |
| talkative people | No explicit support | A single API call natively supports two-person conversations. |
| Long text stability | Optimized for audiobooks, podcasts, and lectures, ensuring highly stable speaker performance. | Quality may drift after a few minutes; it is recommended to process in blocks. |
| Safety and Compliance | System-level mandatory consent prevents the production and use of unauthorized audio. | All outputs are watermarked with SynthID, subject to the Terms of Service. |
| Sound quality ranking | 72% prefer MAI-Voice-1, which is difficult to distinguish from a real person. | Artificial Analysis TTS Ranking Elo 1211 (Second) |
Application scenarios of MAI-Voice-2
-
Smart AssistantProvides a brand-specific voice for Copilot, applications, devices, and customer service centers.
-
Entertainment contentCreate character voices and narration for games, podcasts, audiobooks, and AR/VR.
-
AccessibilityProvides text-to-speech for visually impaired users and speech alternatives for people with speech impairments.
-
Education and TrainingProvides instructor and virtual character voices for online courses and simulation scenarios.
-
Content creationCreators can convert text into personalized audio content without needing a recording studio.