AB
AiBoss
project

SoulX-Podcast - Soul's multi-speaker speech synthesis model

SoulX-Podcast is a multi-speaker text-to-speech (TTS) model developed by Soul AI Lab, specifically designed for generating long podcast conversations. The model has 1.7B parameters and supports Mandarin, English, and various Chinese dialects (such as Sichuanese...).

What is SoulX-Podcast?

SoulX-Podcast is a multi-speaker text-to-speech (TTS) model developed by Soul AI Lab, specifically designed for generating long podcast conversations. The model has 1.7B parameters and supports Mandarin, English, and various Chinese dialects (such as Sichuanese, Henan dialect, and Cantonese). It features cross-dialect prompting, generating target dialect speech based on Mandarin prompts. The model supports paralinguistic control (such as laughter and sighs) to enhance the realism of the synthesized speech. SoulX-Podcast can generate coherent conversations exceeding 90 minutes, maintaining stable timbre and emotional continuity, making it suitable for podcasts, audiobooks, and other similar scenarios.

Main functions of SoulX-Podcast

  • More people supportSupports dialogue generation between multiple speakers and can naturally switch between different speakers' voices, making it suitable for podcasts, audiobooks, and other scenarios.
  • Multilingual and dialect supportSupports Mandarin, English, and various Chinese dialects (such as Sichuanese, Henanese, and Cantonese), and has cross-dialect prompting function, which can generate target dialect speech through Mandarin prompts.
  • Secondary language controlSupports non-verbal information (such as laughter, sighs, throat clearing, etc.) to enhance the realism of speech synthesis and make the generated speech more natural and vivid.
  • Long Dialogue GenerationIt can generate coherent dialogues of over 90 minutes, maintaining a stable tone and emotional continuity, making it suitable for generating long podcast content.
  • Zero-sample speech cloningIt supports zero-sample speech cloning, enabling the generation of high-quality personalized speech even without a target speaker's speech sample.

The technical principles of SoulX-Podcast

  • Basic model architectureBased on the Qwen3-1.7B architecture, it is a powerful pre-trained language model that has been fine-tuned to adapt to multi-speaker dialogue generation tasks.
  • Multi-speaker modelingBy introducing speaker embedding technology, the model can distinguish the speech features of different speakers and switch speakers naturally during the generation process.
  • Cross-dialect generationUsing the Dialect-Guided Prompting (DGP) method, the model can generate speech in the target dialect based on Mandarin prompts, supporting zero-shot generation of multiple dialects.
  • Secondary language controlBy adding specific sub-language markers (such as...) to the text input <|laughter|>,<|sigh|> (etc.), the model can add corresponding non-linguistic information to the generated speech to enhance the realism of the speech.
  • Long-form generation stabilityBy optimizing the model's attention mechanism and decoder structure, we ensure stable timbre and emotional continuity in the generation of long dialogues, avoiding timbre drift and emotional inconsistency.
  • Data processing and trainingThe model is trained using large-scale multi-speaker dialogue data. The data processing workflow includes speech enhancement, audio segmentation, speaker logging, text transcription, and quality filtering to ensure that the model can learn rich dialogue features.

SoulX-Podcast project address

  • Project official website: https://soul-ailab.github.io/soulx-podcast/
  • GitHub repository: https://github.com/Soul-AILab/SoulX-Podcast
  • HuggingFace model libraryhttps://huggingface.co/collections/Soul-AILab/soulx-podcast
  • arXiv technical paper: https://arxiv.org/pdf/2510.23541

Application scenarios of SoulX-Podcast

  • Podcast ProductionThe model can generate coherent dialogues of over 90 minutes, making it suitable for creating podcast content in various fields such as technology, culture, and entertainment.
  • audiobooksThe model can generate dialogues between multiple characters, making audiobooks more vivid and interesting, and is suitable for long content such as novels and stories.
  • Educational contentGenerate multi-role dialogues to enhance the interactivity and fun of educational content such as language learning and historical storytelling.
  • Entertainment and GamesGenerate natural multi-character voices for games, animations, and videos, enhancing the immersive experience of the content.
  • Corporate TrainingGenerate simulated dialogues to help employees train in communication skills and customer service.