AB
AiBoss
project

SkyReels V4 - Kunlun Tech's AI multimodal video foundation model

SkyReels V4 is a video foundation model developed by Kunlun Tech. It is the world's first AI video model to support multimodal input, joint audio and video generation, and unified generation/repair/editing. The model employs a dual-stream MMDiT architecture and can generate 108...

What is SkyReels V4?

SkyReels V4 is a video foundation model launched by Kunlun Tech. It is the world's first AI video model to support multimodal input, joint audio and video generation, and unified generation/repair/editing. The model adopts a dual-stream MMDiT architecture and can generate 1080p/32FPS/15-second cinematic-quality synchronized audio and video. It ranks second globally on the Artificial Analysis leaderboard, surpassing mainstream models such as Google Veo 3.1 and OpenAI Sora 2, and supports multimodal control of text, images, video, and audio, as well as professional-grade video repair and editing.

Main features of SkyReels V4

  • Multimodal precision controlIt supports various input combinations such as text, images, video clips, masks, and audio references, enabling the preservation of the main image, the transfer of timbre, and the replacement of actions.
  • Professional-grade video restorationThrough regional intelligent repair and reference-guided repair, the main body of the video can be accurately replaced, attributes modified, or background changed to ensure visual consistency before and after editing.
  • Full-dimensional video editingSupports partial editing (adding or deleting objects, modifying textures), intelligent element removal (watermarks/subtitles/logos), and global style migration and scene attribute adjustment.
  • High-quality audio generationThe model incorporates multilingual speech synthesis, sound effect generation, and background music adaptation, and supports synchronized singing of emotional speech and lyrics, with outstanding performance in Chinese speech.

The technical principles of SkyReels V4

  • Dual-stream MMDiT architectureIt adopts a symmetrical dual-stream design, with video and audio branches sharing an MLLM text encoder, and achieves full-network deep audio-visual synchronization through a bidirectional cross-attention mechanism; it uses RoPE frequency scaling technology to solve the problem of audio and video time scale mismatch, and combines it with a joint stream matching loss function to fundamentally solve the problems of lip-sync and sound effect alignment.
  • Unified splicing frameworkIt innovatively introduces a dual-dimensional paradigm that combines channel stitching and temporal stitching, unifying diverse tasks such as generation, repair, and editing into repair issues under specific mask configurations, achieving one-stop coverage of video operations across all scenarios, and enabling end-to-end creation without switching tools.
  • Efficient generation strategyThe model adopts a joint generation strategy of "low-resolution full sequence + high-resolution key frame" and, together with the video sparse attention mechanism, reduces the attention computation cost by about 3 times, making the generation of 1080p high-resolution long-duration videos practical.

SkyReels V4 project address

  • arXiv technical paper: https://arxiv.org/pdf/2602.21818

Application scenarios of SkyReels V4

  • Advertising and MarketingThe model can quickly generate product promotional videos, supports multiple style switching and batch editing, and improves advertising production efficiency.
  • Content creationThe model supports visualization of short video scripts, intelligent editing and repair of Vlogs, and simultaneous multi-language dubbing, lowering the barrier to creation.
  • Film and television productionUsed for early concept visualization, shot expansion, post-production restoration and local editing, accelerating the film and television industrialization process.
  • Education and TrainingThe model supports the generation of teaching videos, visualization of courseware, and automatic synchronization of multilingual subtitles, facilitating the production of online education content.