AB
AiBoss
project

JoyGen - JD.com and HKU launch audio-driven 3D speaking face video generation framework

JoyGen, developed by JD Technology and the University of Hong Kong, is an audio-driven 3D speaking face video generation framework focused on achieving accurate lip-to-audio synchronization and high-quality visual effects. JoyGen combines audio features and facial depth...

What is JoyGen?

JoyGen, developed by JD Technology and the University of Hong Kong, is an audio-driven 3D speaking face video generation framework focused on achieving accurate lip-audio synchronization and high-quality visual effects. JoyGen combines audio features and facial depth maps to drive lip movement generation, using a single-step UNet architecture for efficient video editing. During training, JoyGen was tested on the open-source HDTF dataset with a high-quality dataset containing 130 hours of Chinese video, validating its superior performance. Experimental results demonstrate that JoyGen achieves industry-leading levels in both lip-audio synchronization and visual quality, providing a new technological solution for speaking face video editing.

JoyGen's main functions

  • Lips synchronized with audioAudio-driven lip motion generation technology ensures that the lip movements of people in the video accurately correspond to the audio content.
  • High-quality visual effectsThe generated video has realistic visuals, including natural facial expressions and clear lip details.
  • Video editing and optimizationIt allows for the editing and optimization of lip movements based on existing videos without the need to regenerate the entire video.
  • Multilingual supportIt supports video generation in different languages such as Chinese and English, adapting to various application scenarios.

JoyGen's technical principles

  • Phase 1:
    • Audio-driven lip movement generation 3D reconstruction modelThe 3D reconstruction model extracts identity coefficients from the input facial image, which are used to describe the facial features of the person.
    • Audio to Motion ModelBased on an audio-to-motion model, audio signals are converted into expression coefficients, which are used to control lip movements.
    • Depth map generationThe system combines identity and expression coefficients to generate a 3D mesh of the face, and uses differentiable rendering technology to generate a facial depth map for subsequent video compositing.
  • Phase Two:
    • Visual appearance synthesis - single-step UNet architectureThis paper describes a method that integrates audio features and depth map information into the video frame generation process using a single-step UNet network. UNet maps the input image to a low-dimensional latent space based on the encoder, and combines audio features and depth map information to generate lip movements.
    • Cross-attention mechanismThe audio features interact with image features based on a cross-attention mechanism to ensure that the generated lip movements are highly consistent with the audio signal.
    • Decoding and OptimizationThe generated latent representation is then reconstructed into image space by the decoder to generate the final video frames. Optimization is performed in both the latent and pixel spaces using the L1 loss function to ensure high-quality and synchronized video generation.
  • Dataset supportJoyGen is trained using a high-quality dataset containing 130 hours of Chinese video to ensure that the model can adapt to various scenarios and language environments.

JoyGen's project address

JoyGen's application scenarios

  • Virtual anchors and live streamingCreate virtual anchors to deliver news broadcasts, e-commerce live streams, etc., and generate realistic lip movements in real time based on the input audio to enhance the viewer experience.
  • Animation ProductionIn the field of animation and film, it can quickly generate lip animation synchronized with dubbing, reducing the workload of animators and improving production efficiency.
  • Online EducationIt generates a virtual teacher avatar and enables lip movements synchronized with the teaching voice, making the teaching videos more vivid and enhancing students' learning interest.
  • Video content creationIt helps creators quickly generate high-quality videos of people speaking with their faces, such as virtual character short dramas and funny videos, thus enriching the creative forms.
  • Multilingual video generationIt supports multiple languages, quickly converting videos in one language to other languages, and lip movements are synchronized with the new language audio, facilitating the international dissemination of content.