project
JoyVASA - JD Health's open-source, audio-driven digital head project
JoyVASA is an open-source, audio-driven digital head project from JD Health International. Based on diffusion model technology, it generates facial and head movements synchronized with the audio signal. JoyVASA can achieve synchronized lip movements...
What is JoyVASA?
JoyVASA is an open-source, audio-driven digital head project from JD Health International. Based on diffusion model technology, it generates facial dynamics and head movements synchronized with audio signals. JoyVASA can achieve lip-sync and facial expression control for humans and has been extended to the animation generation of animal heads, showing broad application potential in multilingual support and cross-species animation.
JoyVASA's main functions
- Audio-driven facial animationIt generates synchronized facial animations based on the input audio signal, including lip movements and facial expression changes.
- Lip synchronizationAchieving realistic dialogue effects through precise matching of audio and lip movements.
- Facial expression controlControlling and generating specific facial expressions enhances the expressiveness of animation.
- Animal facial animationJoyVASA can generate facial animations for animals, expanding its application range.
- Multilingual supportBased on training on a mixed dataset containing Chinese and English data, JoyVASA supports multilingual animation generation.
- High-quality video generationThe project can generate high-resolution and high-quality animated videos, enhancing the viewing experience.
JoyVASA's technical principles
- Decoupled facial representationJoyVASA uses a decoupled facial representation framework to separate dynamic facial expressions from static 3D facial representations, generating longer videos.
- diffusion modelThe project uses a diffusion model to generate motion sequences directly from audio cues, and the motion sequences are independent of the character's identity.
- Two-stage training:
- Phase 1Separate static facial features and dynamic motion features. Static features capture facial identity features, while dynamic features encode dynamic elements such as facial expressions, scaling, rotation, and translation.
- Phase TwoTrain a diffusion transformer to generate motion features from audio features.
- Audio feature extractionThe audio features of the input speech are extracted using the wav2vec2 encoder and used as a condition for generating motion sequences.
- Motion sequence generationBased on a diffusion model, audio-driven motion sequences are sampled in a sliding window. These motion sequences include facial expressions and head movements.
JoyVASA's project address
- Project official website:jdh-algo.github.io/JoyVASA
- GitHub repository:https://github.com/jdh-algo/JoyVASA
- HuggingFace model library:https://huggingface.co/jdh-algo/JoyVASA
- arXiv technical paper:https://arxiv.org/pdf/2411.09209
JoyVASA Application Scenarios
- Virtual AssistantIn smart homes, customer service, and technical support, it provides realistic facial animations and expressions for virtual assistants, enhancing the user interaction experience.
- Entertainment and MediaUsed to generate or enhance character facial expressions and movements, reducing the need for traditional motion capture. Provides game characters with more natural facial expressions and animations, enhancing game immersion.
- social mediaUsers can use JoyVASA to generate their own virtual avatars for use in video chats or content creation on social media platforms.
- Education and trainingIn online education platforms, virtual teachers can be created to provide a more engaging learning experience. In fields such as medicine and the military, simulated human reactions and expressions can be used for professional training.
- Advertising and MarketingCreate compelling virtual spokespeople for advertising and to enhance brand appeal.