AB
AiBoss
project

MAI-Voice-2 - Microsoft's next-generation text-to-speech model

MAI-Voice-2 is Microsoft's next-generation text-to-speech (TTS) model, and is Microsoft's most expressive and natural-sounding speech synthesis model to date. Compared to its predecessors, it excels in fidelity, language coverage, speaker consistency, and emotional range...

What is MAI-Voice-2?

MAI-Voice-2 is Microsoft's next-generation text-to-speech (TTS) model, and is Microsoft's most expressive and natural-sounding speech synthesis model to date. Compared to its predecessor, it has been comprehensively improved in terms of fidelity, language coverage, speaker consistency, and emotional range. It supports 15+ languages and features fine-grained emotion control, zero-sample speech cloning, and code switching capabilities.

Main functions of MAI-Voice-2

  • Multilingual natural synthesisExpanding from English only to 15+ languages while maintaining the same naturalness and expressiveness.
  • Fine-grained emotion controlPrecisely control vocal emotions through emotional tags (such as sadness, whispering, excitement, confusion, etc.).
  • Zero-sample speech cloningThe target voice can be cloned with only 5-60 seconds of reference audio, and all languages are supported.
  • Speaker's identity is stableMaintain consistent speaker characteristics across long content, including audiobooks, podcasts, and lectures.
  • Natural code switchingIt supports natural mixing of Hindi-English, Spanish-English, and other languages without losing rhythm and identity consistency.
  • Role-playingSupports specific role styles such as motivational coaches and sports commentators.

Technical Principles of MAI-Voice-2

  • Self-developed speech basic model architectureMAI-Voice-2 is built on Microsoft's in-house developed speech model and employs an end-to-end neural network speech synthesis architecture. The model can holistically understand input text, automatically adapting to intonation, emotion, and speaking style, generating human-like speech without requiring extensive manual parameter tuning by developers. Its architecture is similar to Azure Neural HD speech, achieving generational improvements in expressiveness, language coverage, and speaker consistency.
  • Multilingual unified modelingMAI-Voice-2 expands upon the English monolingual model of MAI-Voice-1 into a unified multilingual speech synthesis system supporting 15+ languages. The model undergoes in-depth optimization for different phonological systems, including tonal languages, pitch-stressed languages, stress-timing languages, and syllable-timing languages, ensuring that each language achieves the same output quality as English in terms of naturalness and expressiveness.
  • Zero-sample speech cloning (Voice Prompting)The model supports zero-shot speech cloning, requiring only 5–60 seconds of reference audio to extract speaker identity features and transfer them to the target language, without the need for fine-tuning or retraining for specific speakers. Based on voice prompting technology, the system extracts speaker embeddings through a reference audio encoder, maintaining consistency in timbre, intonation, and prosodic features during synthesis.

How to use MAI-Voice-2

  • Azure Foundry access: Directly call the MAI-Voice-2 API through the Azure Foundry platform.
  • Custom brand voiceYou can create a custom sound by uploading a 5-60 second reference audio file, without the need for retraining or fine-tuning.
  • Emotional labeling controlAdd emotion tags to the request to adjust the emotional style of the output speech.
  • Authorization applicationThe voice cloning function requires authorization, and the system only supports licensed voices for use in production environments.

MAI-Voice-2's core advantages

  • Leading in sound qualityIn blind testing, users preferred the previous generation MAI-Voice-1 in 72% of cases.
  • It's hard to tell the real from the fake.The speakers are extremely similar, making it difficult to distinguish the synthesized speech from the real person's recording.
  • Safety and complianceA system-level mandatory consent mechanism ensures that only authorized voice clones are allowed in the production environment, preventing unauthorized misuse.
  • Long text stabilityMaintaining consistent speaker identity and voice quality throughout hours of content.
  • Low-barrier cloningNo professional recording studio or large amount of training data is needed; sound can be reproduced in just a few seconds of audio.

MAI-Voice-2 project address

  • Project official websitehttps://microsoft.ai/news/mai-voice-2expressive-speech-in-10-languages/

Comparison of MAI-Voice-2 with similar competing products

Comparison Dimensions MAI-Voice-2 Gemini 3.1 Flash TTS
Developer Microsoft (AI) Google DeepMind
Release time June 2026 April 2026 (Public Preview)
Language support 15+ languages, including code switching (Hindi-English, Spanish-English). 70+ languages, covering a wider range of languages
Preset Sounds The exact number was not disclosed, with a focus on brand customization. 30 named sounds (Kore, Puck, Charon, etc.)
Emotional control Fine-grained SSML tags (sadness, whispering, excitement, confusion, etc.) 200+ inline audio tags ([sigh],[laughing],[whispering] (etc.), supports natural language prompts
Voice cloning 5–60 seconds zero-shot, full language support Not supported
talkative people No explicit support A single API call natively supports two-person conversations.
Long text stability Optimized for audiobooks, podcasts, and lectures, ensuring highly stable speaker performance. Quality may drift after a few minutes; it is recommended to process in blocks.
Safety and Compliance System-level mandatory consent prevents the production and use of unauthorized audio. All outputs are watermarked with SynthID, subject to the Terms of Service.
Sound quality ranking 72% prefer MAI-Voice-1, which is difficult to distinguish from a real person. Artificial Analysis TTS Ranking Elo 1211 (Second)

Application scenarios of MAI-Voice-2

  • Smart AssistantProvides a brand-specific voice for Copilot, applications, devices, and customer service centers.
  • Entertainment contentCreate character voices and narration for games, podcasts, audiobooks, and AR/VR.
  • AccessibilityProvides text-to-speech for visually impaired users and speech alternatives for people with speech impairments.
  • Education and TrainingProvides instructor and virtual character voices for online courses and simulation scenarios.
  • Content creationCreators can convert text into personalized audio content without needing a recording studio.