AB
AiBoss
project

VoxCPM - A speech generation model jointly developed by Facewall Intelligence and Tsinghua University

VoxCPM is a 0.5B parameter speech generation model jointly developed by Wallfacer and Tsinghua University Shenzhen International Graduate School. It has achieved industry-leading levels in speech synthesis naturalness, timbre similarity, and prosodic expressiveness. VoxC...

What is VoxCPM?

VoxCPM is a 0.5B parameter speech generation model jointly developed by Wallfacer and Tsinghua University Shenzhen International Graduate School. It has achieved industry-leading levels in naturalness, timbre similarity, and prosodic expressiveness in speech synthesis. VoxCPM employs an end-to-end diffusion autoregressive architecture, directly generating continuous speech representations from text, overcoming the limitations of traditional discrete word segmentation. Through hierarchical language modeling and finite-state quantization constraints, it achieves implicit decoupling between semantics and acoustics, significantly improving the expressiveness and generation stability of speech. VoxCPM supports zero-sample voice cloning, requiring only a reference audio clip to accurately replicate the speaker's timbre, accent, emotional intonation, and other features, generating highly realistic speech. It boasts extremely high inference efficiency, achieving a real-time factor (RTF) as low as 0.17 on an NVIDIA RTX 4090 GPU, meeting the needs of real-time applications. VoxCPM supports bilingual (Chinese and English) voice replication, can synthesize formula and symbol audio, and enables custom pronunciation correction.

Main functions of VoxCPM

  • Context-aware speech generationVoxCPM possesses deep text understanding capabilities, inferring appropriate prosody based on the text's semantics and generating highly expressive and fluent natural speech. It can autonomously adjust its speaking style according to the text content, generating highly fitting personalized voice expressions based on a massive 1.8 million-hour bilingual corpus.
  • Zero-sample speech cloningWith just a short reference audio clip, VoxCPM can achieve accurate zero-sample speech cloning. It can perfectly replicate the speaker's timbre, capturing subtle features such as accent, emotional intonation, rhythm, and pauses to create a highly faithful and natural voice imitation.
  • High-efficiency synthesisVoxCPM supports streaming synthesis, and on consumer-grade NVIDIA RTX 4090 GPUs, its real-time factor (RTF) is as low as 0.17, easily meeting the needs of real-time applications.
  • Multilingual supportVoxCPM is primarily trained on English and Chinese, and can generate high-quality bilingual (English and Chinese) speech, suitable for various language environments and application scenarios.
  • Flexible text input methodsVoxCPM supports multiple text input methods, including plain text input and phoneme input. Users can choose different input modes as needed for more precise pronunciation control.
  • powerful voice processing capabilitiesVoxCPM can process complex text content, including formulas, symbols, and other special texts, and generate corresponding speech output. It supports custom pronunciation correction, allowing users to achieve specific pronunciation needs through phoneme substitution.

VoxCPM Technical Principles

  • End-to-end diffusion autoregressive architectureVoxCPM employs an end-to-end diffusion autoregressive architecture to directly generate continuous speech representations from text, overcoming the limitations of traditional discrete word segmentation and enabling more natural handling of speech continuity.
  • Hierarchical Language Modeling and FSQ ConstraintsBy employing Hierarchical Language Modeling and Finite State Quantization (FSQ) constraints, VoxCPM achieves implicit semantic-acoustic decoupling, significantly enhancing the expressiveness and generation stability of speech.
  • Local audio coding module (LocEnc Module)This module is responsible for encoding the input text, extracting its semantic information, and converting it into an intermediate representation suitable for speech generation.
  • Text-Semantic Language Model (TSLM)TSLM is responsible for modeling the semantics of text, generating semantic representations related to the text content, and providing a semantic foundation for subsequent speech generation.
  • Residual Acoustic Language Model (RALM)RALM further refines the acoustic features and adds acoustic details based on TSLM, making the generated speech more natural and realistic.
  • Local Diffusion Generation Module (LocDiT Module)The LocDiT module generates continuous speech features through a diffusion process, fusing semantic and acoustic information to ultimately generate high-quality speech waveforms.
  • Causal VAE codecIt is used to compress the original audio waveform to a low frame rate latent space and reconstruct the generated speech representation back into the waveform signal, ensuring that the generated speech has good quality and stability.

VoxCPM project address

  • Github repositoryhttps://github.com/OpenBMB/VoxCPM/
  • Hugging Face Model Libraryhttps://huggingface.co/openbmb/VoxCPM-0.5B
  • Experience the demo onlinehttps://huggingface.co/spaces/OpenBMB/VoxCPM-Demo

Application scenarios of VoxCPM

  • voice assistantVoxCPM provides intelligent voice assistants with natural and fluent speech synthesis capabilities, enabling them to interact with users in a more human-like voice and improve the user experience.
  • audiobooksIt can convert text content into high-quality speech, making it suitable for producing audiobooks, audio novels, etc., bringing users a more vivid auditory experience.
  • Voice broadcastIt can be used in scenarios such as weather forecasts, news broadcasts, and traffic information broadcasts to generate clear and natural voice broadcast content, improving the efficiency and accuracy of information transmission.
  • Voice cloningVoxCPM's zero-sample voice cloning capability can be used to create personalized voices, such as giving virtual characters and intelligent customer service unique voice features, enhancing their realism and recognizability.
  • EducationIn language learning, online education, and other scenarios, VoxCPM can generate standard speech examples to help learners better imitate and learn pronunciation.
  • Entertainment industryIn the entertainment fields such as games, animation, and film, VoxCPM can generate voices for various characters, enriching the expressiveness and appeal of the content.