AB
AiBoss
project

Higgs Audio V2 - An open-source large-scale voice model capable of simulating multi-person interactive scenarios.

Higgs Audio V2 is an open-source speech model developed by Li Mu and his team at Boson AI. Trained on over 10 million hours of audio data, it features multilingual dialogue generation, automatic prosody adjustment, voice cloning, and singing capabilities...

What is Higgs Audio V2?

Higgs Audio V2 is an open-source speech model developed by Li Mu and his team at Boson AI. Trained on over 10 million hours of audio data, it features multilingual dialogue generation, automatic prosody adjustment, speech cloning, and vocal synthesis. The model can simulate natural and fluent multi-person conversations, automatically matching speaker emotions and intonation, and supports low-latency real-time voice interaction. It supports zero-sample speech cloning; users only need to provide a short speech sample to replicate the voice features of a specific person, and it can even synthesize singing voices. Higgs Audio V2 can generate both speech and background music simultaneously, providing powerful support for audio content creation.

Main features of Higgs Audio V2

  • Multilingual Dialogue GenerationIt supports multilingual dialogue generation, can simulate multi-person interaction scenarios, and automatically matches the speaker's emotions and energy levels to make the dialogue natural and fluent.
  • Automatic rhythm adjustmentIn long text reading, it can automatically adjust the speaking speed, pauses and intonation according to the content without manual intervention, generating natural and fluent speech.
  • Voice cloning and singing synthesisUsers only need to provide a short voice sample, and the model can achieve zero-sample voice cloning, replicating the voice characteristics of a specific person, and even making the cloned voice hum a melody.
  • Real-time voice interactionIt supports low-latency response, can understand user emotions and make emotional expressions, and provides a near-human interactive experience.
  • Voice and background music are generated simultaneously.It can generate both voice and background music simultaneously, enabling the creation process of "writing a song and singing it out".

Technical Principles of Higgs Audio V2

  • AudioVerse datasetWe developed an automated annotation process that combines multiple speech recognition models, sound event classification models, and our self-developed audio understanding model to clean and annotate 10 million hours of audio data.
  • Unified audio segmenterA unified audio segmenter was trained from scratch, capable of capturing both semantic and acoustic features.
  • DualFFN architectureIt significantly enhances the ability of large language models to model acoustic tokens with almost no increase in computational overhead.
  • Zero-sample speech cloningThe model incorporates contextual learning, enabling zero-shot speech cloning and matching of speaking styles using simple cues (such as short reference audio samples).

Higgs Audio V2 project address

  • Github repositoryhttps://github.com/boson-ai/higgs-audio
  • Experience the demo onlinehttps://huggingface.co/spaces/smola/higgs_audio_v2

Application scenarios of Higgs Audio V2

  • Real-time voice interactionSuitable for scenarios such as virtual anchors and real-time voice assistants, providing natural interaction with low latency and emotional expression.
  • Audio content creationIt can generate natural dialogue and narration, providing strong support for audiobooks, interactive training, and dynamic storytelling.
  • Entertainment and creative fieldsThe voice cloning feature can replicate the voice of a specific person, opening up new possibilities in the entertainment and creative fields.