AB
AiBoss
project

VoxCPM1.5 - Wallfacer's open-source end-to-end speech synthesis model

VoxCPM 1.5 is an advanced end-to-end text-to-speech (TTS) model from Wallfacer AI, focusing on context-aware speech generation and realistic voice cloning. The model directly generates speech from text using an end-to-end diffusion autoregressive architecture...

What is Vox CPM1.5?

VoxCPM 1.5 is an advanced end-to-end text-to-speech (TTS) model from Mianbi AI, focusing on context-aware speech generation and realistic voice cloning. The model generates continuous speech directly from text using an end-to-end diffusion autoregressive architecture, supporting 44.1kHz high sampling rate audio cloning for more refined results. Simultaneously, the model's generation efficiency is doubled, requiring only 6.25 tokens to generate 1 second of audio, and stability is enhanced with reduced artifacts. VoxCPM 1.5 offers deep customization capabilities, supporting LoRA and full fine-tuning, helping developers create personalized speech models.

Main functions of VoxCPM1.5

  • High sampling rate audio cloningIt supports a 44.1kHz sampling rate and can clone more detailed sounds from high-quality audio.
  • High-efficiency speech synthesisThe model generation efficiency has been improved, requiring only 6.25 tokens to generate 1 second of audio, doubling the speed and increasing the quality.
  • Context-aware speech generationIt automatically adjusts the tone and style based on the text content to generate natural and fluent speech.
  • Deep customization capabilityAdded LoRA and full fine-tuning scripts to support developers in personalized training and optimization.
  • Enhance stabilityReduce audio artifacts and optimize the speech generation effect for long texts.

Technical Principles of VoxCPM1.5

  • Tokenizer-Free ArchitectureVoxCPM 1.5 employs an unmarked end-to-end architecture that directly generates continuous speech signals from text, avoiding the limitations imposed by discrete tokenization in traditional TTS.
  • diffusion autoregressive modelThe autoregressive architecture based on the diffusion model achieves high-quality speech synthesis by progressively generating a continuous representation of the speech signal.
  • Hierarchical language modelingBy combining the MiniCPM-4 language model, hierarchical modeling is used to achieve implicit decoupling between semantics and acoustics, thereby improving the naturalness and expressiveness of speech.
  • FSQ constraints: Utilize techniques such as Flow Matching to optimize the stability of speech generation and ensure high-quality output of speech synthesis.
  • High-efficiency real-time synthesisIt supports streaming synthesis with an RTF as low as 0.15, enabling low-latency real-time speech synthesis on consumer-grade GPUs.

VoxCPM1.5 project address

  • GitHub repositoryhttps://github.com/OpenBMB/VoxCPM
  • HuggingFace model libraryhttps://huggingface.co/openbmb/VoxCPM1.5

Application Scenarios of VoxCPM1.5

  • Smart HomeIt provides natural and smooth voice interaction for devices such as smart speakers and smart home appliances, enhancing the user experience.
  • audiobooksIt can quickly convert text content into high-quality speech for the production of audiobooks and podcasts.
  • Language learningThe voice cloning function helps learners practice language pronunciation by mimicking the pronunciation of different languages.
  • Game character voice actingGenerate personalized voices for characters in the game to enhance the immersive experience.
  • Brand promotionThe voice cloning function generates the voice of a brand ambassador for use in advertising and promotion.