AB
AiBoss
News

Alibaba's Tongyi dual-mode speech model, Fun-CosyVoice 3.5, and Fun-AudioGen-VD, have been released.

Tongyi Labs has released two speech generation models, Fun-CosyVoice3.5 and Fun-AudioGen-VD, pioneering the FreeStyle command control paradigm. Users can describe details such as tone, emotion, and scene through natural language, without relying on fixed labels. Fun-CosyVoice3.5 supports multilingual replication and fine-grained expression control, adding four new less common languages including Thai and Indonesian, and reducing the misreading rate of rare characters to 5.3%. Fun-AudioGen-VD achieves end-to-end sound design, generating character-specific timbres and simulating environmental acoustic effects.