AB
AiBoss
project

SoulX-Singer - A singing voice synthesis model jointly developed by Soul App and universities.

SoulX-Singer is an open-source, industrial-grade, zero-shot singing synthesis model developed by Soul App in collaboration with Tianjin University and Northwestern Polytechnical University. The model is trained on 42,000 hours of high-quality multilingual singing data and supports MIDI scores and F0...

What is SoulX-Singer?

SoulX-Singer is an industrial-grade, zero-shot singing voice synthesis model open-sourced by Soul App in collaboration with Tianjin University and Northwestern Polytechnical University. The model is trained on 42,000 hours of high-quality multilingual singing voice data, supports dual-mode control of MIDI scores and F0 melodies, and enables precise pitch and rhythm control, cross-language timbre cloning, and lyrics editing. SoulX-Singer employs an advanced Flow Matching architecture and a two-stage training strategy, comprehensively outperforming existing open-source solutions in key metrics such as pitch accuracy, singer similarity, and subjective listening experience, providing a reliable infrastructure for AI music creation and virtual singer applications.

SoulX-Singer's main functions

  • Zero-sample singing cloningInput any singer's reference audio, and a high-quality singing voice with that timbre can be generated without additional training.
  • Dual-mode control synthesisIt can precisely control pitch and rhythm through MIDI scores, and can also switch from humming to singing through F0 melodies.
  • Multilingual singing synthesisSupports high-quality vocal generation in Mandarin, English, and Cantonese.
  • Cross-language timbre transfer: Transferring the vocal characteristics of a singer in one language to the singing of songs in other languages.
  • Real-time lyrics editingWhile maintaining the melody and singing style, the lyrics can be flexibly modified.

The technical principles of SoulX-Singer

  • Flow Matching Generation FrameworkIt adopts flow matching instead of the traditional diffusion model, and achieves more efficient and stable audio generation by directly learning the transmission path of the probability distribution.
  • Audio Infilling MechanismThe singing synthesis is modeled as a conditional waveform completion task, and the target audio is predicted by using context fragments, which naturally ensures long-term coherence and timbre consistency.
  • Explicit multimodal alignmentBy using a length adjuster to force alignment of the temporal relationship between lyrics, MIDI notes, and acoustic features, rhythmic deviations and pronunciation ambiguity caused by implicit alignment are eliminated.
  • Progressive two-stage trainingShort segment training builds musical score comprehension, while long segment training captures long-term breath control, ultimately balancing local precision with overall naturalness.

SoulX-Singer's project address

  • GitHub repository: https://github.com/Soul-AILab/SoulX-Singer
  • HuggingFace model libraryhttps://huggingface.co/Soul-AILab/SoulX-Singer
  • arXiv technical paper: https://arxiv.org/pdf/2602.07803

Application scenarios of SoulX-Singer

  • Virtual singer creationThe model can quickly create virtual idols with unique voices, reducing the signing and recording costs for real singers.
  • AI Cover Songs and Fan CreationsUsers can use any singer's voice to cover popular songs, achieving creative adaptations across languages and styles.
  • Music-assisted creationSongwriters can quickly generate demos using MIDI input to verify the matching effect between melody and lyrics.
  • Audio content productionIt can generate high-quality singing or chanting content in batches for audiobooks, podcasts, game dubbing and other scenarios.
  • Personalized entertainmentOrdinary users can upload their own voices to generate a personalized AI singer to perform any song.