JoyHallo - JD.com's audio-driven video generation AI digital human model
JoyHallo is an open-source AI digital human model from JD.com, designed specifically for Mandarin Chinese. It can generate realistic speaking videos based on audio. It is particularly well-suited for handling complex lip movements and intonation in Mandarin and has the ability to generate videos across languages.
What is JoyHallo?
JoyHallo is an open-source AI digital human model developed by JD.com, designed specifically for Mandarin Chinese. It can generate realistic speaking videos based on audio. It is particularly well-suited for handling the complex lip movements and intonation of Mandarin and has the ability to generate videos across languages. JoyHallo provides an open-source dataset and model training method, allowing users to generate speaking videos in both Mandarin and English. The project uses the Chinese wav2vec2 model for audio feature embedding and employs a semi-decoupled structure to improve inference speed, resulting in a 14.3% improvement.
JoyHallo's main functions
- Audio-driven video generationJoyHallo can generate corresponding videos based on audio input, especially Mandarin videos.
- Cross-language generation capabilitiesIn addition to Mandarin, JoyHallo can generate English videos, demonstrating its cross-language video generation capabilities.
- Lip synchronizationThe model can accurately synchronize lip movements in audio and video, improving the realism of the video.
- Facial expression generationGenerate corresponding facial expressions based on the emotions and tone of voice in the audio.
JoyHallo's technical principles
- semi-decoupled structureThis technology is used to improve the accuracy of lip movement prediction in audio-driven video generation. It enables more precise modeling by integrating and then separating key facial animation components such as lips, expressions, and head pose.
- Feature embeddingEmbedding audio features using the Chinese wav2vec2 model helps the model better understand and generate facial movements synchronized with the audio.
- Cross-attention mechanismIn the semi-decoupled structure, the cross-attention module processes the integrated features and captures correlations.
- Convolutional NetworksIn the decoupling phase, convolutional networks are used to separate different features, allowing the model to focus on the specific details of each feature.
- DatasetJoyHallo is trained on the jdh-Hallo dataset, a Mandarin video dataset containing various ages and speaking styles, covering everyday conversations and professional medical topics.
JoyHallo's project address
- Project official website:jdh-algo.github.io/JoyHallo
- GitHub repository:https://github.com/jdh-algo/JoyHallo
- HuggingFace model library:https://huggingface.co/jdh-algo/JoyHallo-v1
- arXiv technical paper:https://arxiv.org/pdf/2409.13268
JoyHallo's application scenarios
- Virtual streamerJoyHallo generates videos of virtual anchors in fields such as news broadcasting, weather forecasting, and sports commentary, providing 24/7 program production.
- Online EducationIn fields such as language learning and online courses, JoyHallo generates virtual avatars of teachers, providing a more vivid teaching experience.
- Customer ServiceIn the area of customer service, JoyHallo generates virtual customer service representatives to provide more friendly and professional customer service.
- Entertainment industryJoyHallo generates facial animations for characters in fields such as film, games, and animation production, improving production efficiency and reducing costs.
- social mediaUsers can use JoyHallo to create their own virtual avatars and post video content on social media, increasing interactivity and fun.
- Advertising productionIn the advertising industry, JoyHallo generates customized ad videos, enhancing the appeal and personalization of advertisements.