AB
AiBoss
News

Zhipu Open Source GLM-TTS: Controllable Pronunciation Speech Synthesis Based on Multi-Reward Reinforcement Learning

Zhipu AI has released and open-sourced its industrial-grade speech synthesis system, GLM-TTS. Employing a two-stage generation paradigm, it supports 3-second voice replication and multi-dialect cloning. The character error rate (CER) reaches 0.89% after reinforcement learning optimization, achieving state-of-the-art (SOTA) performance among open-source models. Key technological breakthroughs include multi-reward fusion reinforcement learning, refined pronunciation control (Phoneme-in), and a self-developed 2D-Vocos vocoder, significantly improving emotional expression and pronunciation accuracy.