News
Meituan's open-source digital human video model, LongCat-Video-Avatar 1.5
Meituan's LongCat team has officially open-sourced the LongCat-Video-Avatar 1.5 digital human video model, moving from open-source state-of-the-art (SOTA) to commercial-grade applications. The model features an upgraded Whisper-large audio encoder, a high-quality multi-scene data system, and introduces frame-by-frame GRPO preference alignment, resulting in significant improvements in lip-sync, physical plausibility, long video stability, and multi-person interaction. The model utilizes DMD distillation for an 8-step generation process, improving efficiency by approximately 15 times.