News
Meituan open-sources LongCat-Audio-Codec, a high-efficiency speech codec that helps implement real-time interaction.
Meituan's LongCat team has open-sourced its LongCat-Audio-Codec speech codec solution. Designed specifically for large speech language models (Speech LLM), it employs a parallel extraction mechanism of semantic and acoustic dual tokens to balance the semantic and acoustic features of speech, solving the problem of balancing semantic and acoustic information in traditional solutions. The low-latency streaming decoder supports real-time interaction, meeting the needs of scenarios such as in-vehicle voice assistants and real-time translation.