News
Xiaohongshu open-sources a unified speech generation and editing model, FireRedTTS3.
Xiaohongshu's FireRed team has open-sourced FireRedTTS3, a unified speech generation and editing model. Based on RedAE semantically enhanced continuous representation and the LLM-DiT architecture, the model can achieve zero-sample timbre cloning, natural language voice design, and precise speech editing of specified regions for 24 languages and 21 Chinese dialects. The model has achieved state-of-the-art results on four public benchmark datasets, providing efficient solutions for audio content, game voice-over, and brand customer service scenarios.