AB
AiBoss
project

JoyHallo - JD.com's audio-driven video generation AI digital human model

JoyHallo is an open-source AI digital human model from JD.com, designed specifically for Mandarin Chinese. It can generate realistic speaking videos based on audio. It is particularly well-suited for handling complex lip movements and intonation in Mandarin and has the ability to generate videos across languages.

What is JoyHallo?

JoyHallo is an open-source AI digital human model developed by JD.com, designed specifically for Mandarin Chinese. It can generate realistic speaking videos based on audio. It is particularly well-suited for handling the complex lip movements and intonation of Mandarin and has the ability to generate videos across languages. JoyHallo provides an open-source dataset and model training method, allowing users to generate speaking videos in both Mandarin and English. The project uses the Chinese wav2vec2 model for audio feature embedding and employs a semi-decoupled structure to improve inference speed, resulting in a 14.3% improvement.

JoyHallo's main functions

  • Audio-driven video generationJoyHallo can generate corresponding videos based on audio input, especially Mandarin videos.
  • Cross-language generation capabilitiesIn addition to Mandarin, JoyHallo can generate English videos, demonstrating its cross-language video generation capabilities.
  • Lip synchronizationThe model can accurately synchronize lip movements in audio and video, improving the realism of the video.
  • Facial expression generationGenerate corresponding facial expressions based on the emotions and tone of voice in the audio.

JoyHallo's technical principles

  • semi-decoupled structureThis technology is used to improve the accuracy of lip movement prediction in audio-driven video generation. It enables more precise modeling by integrating and then separating key facial animation components such as lips, expressions, and head pose.
  • Feature embeddingEmbedding audio features using the Chinese wav2vec2 model helps the model better understand and generate facial movements synchronized with the audio.
  • Cross-attention mechanismIn the semi-decoupled structure, the cross-attention module processes the integrated features and captures correlations.
  • Convolutional NetworksIn the decoupling phase, convolutional networks are used to separate different features, allowing the model to focus on the specific details of each feature.
  • DatasetJoyHallo is trained on the jdh-Hallo dataset, a Mandarin video dataset containing various ages and speaking styles, covering everyday conversations and professional medical topics.

JoyHallo's project address

JoyHallo's application scenarios

  • Virtual streamerJoyHallo generates videos of virtual anchors in fields such as news broadcasting, weather forecasting, and sports commentary, providing 24/7 program production.
  • Online EducationIn fields such as language learning and online courses, JoyHallo generates virtual avatars of teachers, providing a more vivid teaching experience.
  • Customer ServiceIn the area of customer service, JoyHallo generates virtual customer service representatives to provide more friendly and professional customer service.
  • Entertainment industryJoyHallo generates facial animations for characters in fields such as film, games, and animation production, improving production efficiency and reducing costs.
  • social mediaUsers can use JoyHallo to create their own virtual avatars and post video content on social media, increasing interactivity and fun.
  • Advertising productionIn the advertising industry, JoyHallo generates customized ad videos, enhancing the appeal and personalization of advertisements.