AB
AiBoss
project

EMO2 - Audio-driven avatar video generation technology launched by Alibaba Research Institute

EMO2 (End-Effector Guided Audio-Driven Avatar Video Generation) is an audio-driven avatar video generation technology developed by Alibaba Research Institute of Intelligent Computing. Its full name is "End-Effector Guided Audio-Driven Avatar Video Generation..."

What is EMO2?

EMO2 (End-Effector Guided Audio-Driven Avatar Video Generation) is an audio-driven avatar video generation technology developed by Alibaba Research Institute. It generates expressive, dynamic videos from an audio input and a static portrait photograph. Its core innovation lies in combining audio signals with hand gestures and facial expressions, synthesizing video frames through a diffusion model to generate natural and fluid animations. This includes high-quality visual effects, high-precision audio synchronization, and a rich variety of motions.

EMO2's main functions

  • Audio-driven dynamic avatar generationEMO2 can generate expressive animated avatar videos from an audio input and a static portrait photo.
  • High-quality visual effectsBased on a diffusion model, video frames are synthesized and combined with hand gestures to generate natural and smooth facial expressions and body movements.
  • High-precision audio synchronization: Ensure that the generated video and audio input are highly synchronized in time to enhance the overall naturalness.
  • Diverse Action GenerationIt supports complex and fluid hand and body movements, making it suitable for a variety of scenarios.

EMO2's technical principles

  • Audio-driven motion modelingEMO2 uses an audio encoder to convert the input audio signal into feature embeddings, capturing the emotion, rhythm, and semantic information in the audio.
  • End effector guidanceThis technology pays particular attention to the generation of hand gestures (end-effectors) because there is a strong correlation between hand gestures and audio signals. The model first generates hand poses and then integrates them into the overall video generation process to ensure the naturalness and consistency of the movements.
  • Diffusion Model and Feature FusionEMO2 employs a diffusion model as its core generation framework. During the diffusion process, the model combines features from the reference image, audio features, and multi-frame noise to generate high-quality video frames through repeated denoising operations.
  • Frame encoding and decodingIn the frame encoding stage, ReferenceNet extracts facial features from the input static image, which are then combined with audio features and fed into the diffusion process. Finally, the model decodes and generates videos with rich expressions and natural movements.

EMO2 project address

Application scenarios of EMO2

  • Virtual reality and animationIt can be used to generate expressive and natural speaking avatar animations.
  • Cross-language and culturalIt supports voice input in multiple languages and can generate animations for characters of different styles.
  • Role-playing and gamesIt allows you to apply specific characters to movie and game scenes.