VoxCPM1.5 - Wallfacer's open-source end-to-end speech synthesis model
VoxCPM 1.5 is an advanced end-to-end text-to-speech (TTS) model from Wallfacer AI, focusing on context-aware speech generation and realistic voice cloning. The model directly generates speech from text using an end-to-end diffusion autoregressive architecture...
What is Vox CPM1.5?
VoxCPM 1.5 is an advanced end-to-end text-to-speech (TTS) model from Mianbi AI, focusing on context-aware speech generation and realistic voice cloning. The model generates continuous speech directly from text using an end-to-end diffusion autoregressive architecture, supporting 44.1kHz high sampling rate audio cloning for more refined results. Simultaneously, the model's generation efficiency is doubled, requiring only 6.25 tokens to generate 1 second of audio, and stability is enhanced with reduced artifacts. VoxCPM 1.5 offers deep customization capabilities, supporting LoRA and full fine-tuning, helping developers create personalized speech models.
Main functions of VoxCPM1.5
-
High sampling rate audio cloningIt supports a 44.1kHz sampling rate and can clone more detailed sounds from high-quality audio.
-
High-efficiency speech synthesisThe model generation efficiency has been improved, requiring only 6.25 tokens to generate 1 second of audio, doubling the speed and increasing the quality.
-
Context-aware speech generationIt automatically adjusts the tone and style based on the text content to generate natural and fluent speech.
-
Deep customization capabilityAdded LoRA and full fine-tuning scripts to support developers in personalized training and optimization.
-
Enhance stabilityReduce audio artifacts and optimize the speech generation effect for long texts.
Technical Principles of VoxCPM1.5
-
Tokenizer-Free ArchitectureVoxCPM 1.5 employs an unmarked end-to-end architecture that directly generates continuous speech signals from text, avoiding the limitations imposed by discrete tokenization in traditional TTS.
-
diffusion autoregressive modelThe autoregressive architecture based on the diffusion model achieves high-quality speech synthesis by progressively generating a continuous representation of the speech signal.
-
Hierarchical language modelingBy combining the MiniCPM-4 language model, hierarchical modeling is used to achieve implicit decoupling between semantics and acoustics, thereby improving the naturalness and expressiveness of speech.
-
FSQ constraints: Utilize techniques such as Flow Matching to optimize the stability of speech generation and ensure high-quality output of speech synthesis.
-
High-efficiency real-time synthesisIt supports streaming synthesis with an RTF as low as 0.15, enabling low-latency real-time speech synthesis on consumer-grade GPUs.
VoxCPM1.5 project address
- GitHub repositoryhttps://github.com/OpenBMB/VoxCPM
- HuggingFace model libraryhttps://huggingface.co/openbmb/VoxCPM1.5
Application Scenarios of VoxCPM1.5
-
Smart HomeIt provides natural and smooth voice interaction for devices such as smart speakers and smart home appliances, enhancing the user experience.
-
audiobooksIt can quickly convert text content into high-quality speech for the production of audiobooks and podcasts.
-
Language learningThe voice cloning function helps learners practice language pronunciation by mimicking the pronunciation of different languages.
-
Game character voice actingGenerate personalized voices for characters in the game to enhance the immersive experience.
-
Brand promotionThe voice cloning function generates the voice of a brand ambassador for use in advertising and promotion.