StableAvatar - An audio-driven video generation model launched by Fudan University
StableAvatar is an innovative audio-driven virtual avatar video generation model developed by Fudan University, Microsoft Research Asia, and others. The model utilizes an end-to-end video diffusion transformer, combined with a time-step-aware audio adapter and native audio...
What is StableAvatar?
StableAvatar is an innovative audio-driven virtual avatar video generation model developed by Fudan University, Microsoft Research Asia, and others. Through an end-to-end video diffusion transformer, combined with a time-step-aware audio adapter, native audio guidance mechanism, and a dynamically weighted sliding window strategy, the model can generate high-quality virtual avatar videos of unlimited length. The model solves the problems of identity consistency, audio synchronization, and video smoothness encountered by existing models in long video generation, significantly improving the naturalness and coherence of the generated videos, and is suitable for scenarios such as virtual reality and digital human creation.
Main functions of StableAvatar
-
High-quality long video generationSupports the generation of high-quality virtual avatar videos exceeding 3 minutes, maintaining identity consistency and audio synchronization.
-
No post-processing requiredIt can directly generate videos without using any post-processing tools (such as face-swapping tools or facial restoration models).
-
Diverse applicationsIt supports the generation of animations for various virtual characters, including full-body, half-body, multi-character, and cartoon characters, and is suitable for scenarios such as virtual reality, digital human creation, and virtual assistants.
The technical principles of StableAvatar
-
Time-step sensing audio adapter:By employing time-step-aware modulation and cross-attention mechanisms, audio embeddings interact with latent representations and time-step embeddings, reducing the accumulation of errors in the latent distribution.This enables diffusion models to more effectively capture the joint distribution of audio and latent features.
-
Audio native boot mechanism:Instead of the traditional classification free guidance (CFG), it directly manipulates the sampling distribution of the diffusion model, guiding the generation process toward the joint audio-latent distribution.Using the joint audio-latent prediction that evolves continuously during the denoising process of the diffusion model as a dynamic guiding signal, the naturalness of audio synchronization and facial expressions is enhanced.
-
Dynamic weighted sliding window strategy:When generating long videos, a dynamic weighted sliding window strategy is used to fuse latent representations and logarithmic interpolation is used to dynamically allocate weights, thereby reducing the discontinuity in transitions between video segments and improving the smoothness of the video.
StableAvatar's project address
- Project official websitehttps://francis-rings.github.io/StableAvatar/
- GitHub repositoryhttps://github.com/Francis-Rings/StableAvatar
- HuggingFace model libraryhttps://huggingface.co/FrancisRing/StableAvatar
- arXiv technical paper: https://arxiv.org/pdf/2508.08248
Application scenarios of StableAvatar
- Virtual Reality (VR) and Augmented Reality (AR)By generating high-quality virtual avatar videos, it provides users with a more realistic and natural virtual reality and augmented reality experience, enhancing the user's sense of immersion.
- Virtual assistants and customer serviceGenerate natural facial expressions and movements for virtual assistants and customer service representatives, and provide real-time animated responses to voice commands to enhance the user experience.
- Digital Human CreationIt can quickly generate digital human videos with highly consistent and natural movements, supporting various formats such as full-body, half-body, multi-person, and cartoon characters to meet the needs of different scenarios.
- Film and television productionIt is used to generate high-quality virtual character animations, reduce the time and cost of special effects production, and improve the efficiency and quality of film and television production.
- Online education and trainingThis feature generates animated videos of virtual teachers or trainers for online education platforms, displaying natural facial expressions and movements based on the audio content, thereby enhancing the interactivity and fun of teaching.