SkyReels-V1 - Kunlun Tech's first open-source video generation model for AI short drama creation.
SkyReels-V1 is Kunlun Wanwei's first open-source video generation model for AI-powered short drama creation. Based on fine-tuning of tens of millions of high-quality film and television data sets, it achieves film-level generation of micro-expressions and body movements, supporting 33 levels of detail...
What is SkyReels-V1?
SkyReels-V1 is Kunlun Wanwei's first open-source video generation model for AI-powered short drama creation. Based on fine-tuning of tens of millions of high-quality film and television datasets, it achieves film-level generation of micro-expressions and body movements, supporting 33 subtle facial expressions and over 400 natural motion combinations to highly reproduce realistic emotional expressions. The model supports both text-to-video and image-to-video generation, achieving state-of-the-art (SOTA) performance among open-source video generation models. SkyReels-V1 leverages the self-developed inference framework SkyReels-Infer to significantly improve inference efficiency, supports multi-GPU parallelism and low-memory optimization, and efficiently generates high-quality videos on consumer-grade graphics cards.
Main functions of SkyReels-V1
- High-quality film and television-grade video generationIt supports generating video content with cinematic lighting effects, delicate facial expressions, and natural body movements. Every frame boasts high-quality cinematic quality in composition, actor positioning, and camera angles.
- Fine control of facial expressions and movementsIt supports 33 kinds of delicate character expressions and more than 400 natural motion combinations, and can generate micro-expressions such as laughing, roaring, surprised, and crying.
- Text-based video and image-based videoIt supports two generation methods: Text-to-Video and Image-to-Video.
- Multi-scenario supportIt supports processing single-person shots and multi-person compositions, and supports complex scenes and emotional expressions.
The technical principles of SkyReels-V1
- Self-developed data cleaning and labeling pipelineThe model is trained using high-quality film and television data (such as Hollywood movies and TV series), and based on its self-developed data cleaning and annotation pipeline, it performs fine-grained annotations on character expressions, actions, scenes, etc., to improve the model's ability to understand human performances.
- Multi-stage pre-training and fine-tuning:
- Phase 1Model domain adaptation pre-training adapts the base model to the human-centric video domain.
- Phase 2Transform the text-to-video model into an image-to-video model and pre-train it on the same dataset.
- Phase 3Fine-tuning on a high-quality subset ensures high performance of the model in complex video generation tasks.
- Multimodal understanding and generationBy combining multimodal understanding of character expressions, actions, scenes, and plot, behavioral semantic units and character spatial position perception technology are constructed to achieve accurate character performance generation.
- Efficient reasoning optimization:
- Employing FP8 quantization, parameter-level offload, and optimized attention mechanisms (such as SageAttn), it significantly reduces memory usage and improves inference speed.
- It supports multi-GPU parallel inference and further improves generation efficiency based on distributed computing.
SkyReels-V1 project address
- GitHub repository:https://github.com/SkyworkAI/SkyReels-V1
- HuggingFace model library:https://huggingface.co/collections/Skywork/skyreels-v1
Application scenarios of SkyReels-V1
- AI Short Dramas and Film ProductionIt enables the low-cost generation of high-quality short dramas and film and television special effects, simplifies the production process, and improves efficiency.
- Virtual contentCreate vivid virtual anchors, virtual idols, and other characters, providing natural expressions and movements.
- Advertising and MarketingQuickly generate brand advertising videos to meet diverse marketing needs.
- Education and Training: Create engaging instructional videos to support language learning, historical reenactments, and scientific demonstrations.
- social mediaGenerate personalized short videos to meet users' content creation and sharing needs.