EDTalk - Shanghai Jiao Tong University and NetEase jointly launch a highly efficient, decoupled emotional speaking avatar synthesis model.
EDTalk is an audio-driven lip-sync model jointly developed by Shanghai Jiao Tong University and NetEase, enabling independent control of lip movements, head posture, and emotional expressions. Simply upload an image, an audio clip, and a reference video to drive...
What is EDTalk?
EDTalk is an audio-driven lip-sync model jointly developed by Shanghai Jiao Tong University and NetEase, enabling independent control of lip movements, head posture, and emotional expressions. Simply upload an image, an audio clip, and a reference video to drive the person in the image to speak, supporting custom emotions such as happiness, anger, and sadness. EDTalk decomposes facial dynamics into three independent latent spaces representing lip movements, posture, and expressions through three lightweight modules. Each space is represented by a set of learnable basis vectors, whose linear combination defines a specific action. This efficient decoupled training mechanism improves training efficiency and reduces resource consumption, allowing even beginners to quickly get started and explore innovative applications.
EDTalk's main functions
- Audio-driven lip synchronizationEDTalk can drive the people in the uploaded pictures and audio to speak, achieving lip-sync.
- Customized Emotional ExpressionEDTalk supports custom emotions, such as happiness, anger, and sadness, and the facial expressions of the characters in the synthesized video are highly consistent with the emotions in the audio.
- Audio-to-Motion moduleEDTalk's Audio-to-Motion module can automatically generate lip movements synchronized with the audio rhythm and contextual expressions based on audio input.
- Supports video and audio inputEDTalk can generate accurate emotional speaking avatars with both video and audio input.
EDTalk's technical principles
- High-efficiency decoupling frameworkEDTalk decomposes facial dynamics into three distinct latent spaces using three lightweight modules, representing mouth shape, head posture, and emotional expression. This decoupling technique allows for independent control of these facial movements without interference.
- Learnable basis vector representationsEach latent space is represented by a set of learnable basis vectors, and linear combinations of these basis vectors define specific actions. This design allows EDTalk to flexibly synthesize speaker headshot videos with specific lip shapes, head postures, and facial expressions.
- Orthogonality and efficient training strategiesTo ensure independence and accelerate training, EDTalk enforces orthogonality between basis vectors and designs an efficient training strategy that assigns action responsibility to each space without relying on external knowledge.
EDTalk project address
- Project official website:https://tanshuai0219.github.io/EDTalk/
- Github repository:https://github.com/tanshuai0219/EDTalk
- arXiv technical paper:https://arxiv.org/pdf/2404.01647
Application scenarios of EDTalk
- Personalized customization of personal digital assistantsEDTalk can be used to create personalized digital assistants by synthesizing dynamic facial videos that match the user's voice, thus enhancing the interactive experience.
- Film and television post-productionIn film and television production, EDTalk can be used for character dialogue synthesis, generating lip movements and facial expressions that match the character's emotions through audio-driven methods, thereby enhancing the character's expressiveness.
- Development of interactive teaching assistants for educational softwareEDTalk can be applied to educational software to create interactive teaching assistants and enhance the learning experience through emotional expression.
- Remote communicationIn the field of remote communication, EDTalk can provide a more realistic and emotionally resonant video communication experience, improving communication effectiveness.
- Virtual Reality InteractionIn virtual reality environments, EDTalk can be used to generate virtual characters with emotional expression, enhancing the user's sense of immersion.