AB
AiBoss
News

Meituan's LongCat team launches LongCat-AudioDiT speech synthesis model

Meituan's LongCat team has launched the LongCat-AudioDiT speech synthesis model, achieving state-of-the-art performance in zero-sample timbre cloning. The model directly generates signals in the waveform latent space, abandoning the traditional Mel-spectral intermediate representation and avoiding information loss. LongCat-AudioDiT proposes two key technologies: Dual Constraint Alignment (DCA) and Adaptive Projection Guidance (APG), to address the training-inference mismatch problem and alleviate oversaturation.