EmoTalk3D - A 3D digital human framework jointly launched by Huawei and Fudan University
EmoTalk3D is a 3D digital human framework jointly developed by Huawei Noah's Ark Lab, Nanjing University, and Fudan University. The core technology lies in its ability to synthesize 3D talking avatars with rich emotional expression. EmoTalk3D can capture and reproduce...
What is EmoTalk3D?
EmoTalk3D is a 3D digital human framework jointly developed by Huawei Noah's Ark Lab, Nanjing University, and Fudan University. The core technology lies in its ability to synthesize 3D talking avatars with rich emotional expression. EmoTalk3D can capture and reproduce human lip movements, facial expressions, and even more subtle emotional details such as wrinkles and other micro-facial movements when speaking. Through a mapping framework called "Speech-to-Geometry-to-Appearance," EmoTalk3D achieves the prediction from audio features to 3D geometric sequences, and then the synthesis of the 3D avatar's appearance.
Main functions of EmoTalk3D
- Emotional Expression SynthesisIt can synthesize 3D avatar animations with corresponding emotional expressions based on the input audio signal, including but not limited to various emotional states such as joy, sadness, and anger.
- Lip synchronizationHighly accurate lip movements synchronized with speech; the 3D avatar's lip movements match the actual pronunciation when speaking.
- Multi-view renderingIt supports rendering 3D avatars from different angles, ensuring high quality and consistency when viewed from different perspectives.
- Dynamic detail captureIt can capture and reproduce facial micro-expressions and dynamic details during speech, such as wrinkles and subtle changes in expression.
- Controllable Emotional RenderingUsers can control the emotional expression of the 3D avatar as needed, achieving real-time adjustment and control of emotions.
- High fidelityEmoTalk3D uses advanced rendering technology to generate high-resolution, highly realistic 3D avatars.
The technical principles of EmoTalk3D
- Dataset creation (EmoTalk3D Dataset):Multi-view video data was collected, including emotion annotations and 3D facial geometry information for each frame.The dataset comes from multiple subjects, each of whom recorded multi-perspective videos under different emotional states.
- Audio feature extraction:A pre-trained HuBERT model is used as an audio encoder to convert the input speech into audio features.Emotional tags are extracted from audio features using an emotion extractor.
- Speech-to-Geometry Network (S2GNet):Using audio features and sentiment tags as input, it predicts dynamic 3D point cloud sequences. Based onThe Gated Cyclic Unit (GRU) serves as the core architecture, generating 4D mesh sequences.
- 3D geometry to appearance mapping:Based on the predicted 4D point cloud, the appearance of a 3D avatar is synthesized using Geometry-to-Appearance Network (G2ANet).The appearance is decomposed into normal Gaussian (static appearance) and dynamic Gaussian (wrinkles, shadows, etc. caused by facial movements).
- 4D Gaussian Model:3D Gaussian Splatting technology is used to represent the appearance of 3D avatars.Each 3D Gaussian is parametrically represented by its position, scale, rotation, and transparency.
- Dynamic detail synthesis:Predict dynamic details, such as wrinkles and subtle facial expressions, using FeatureNet and RotationNet networks.
- Head integrity:For non-facial areas (such as hair, neck, and shoulders), an optimization algorithm is used to construct the model starting from uniformly distributed points.
- Rendering module:By fusing dynamic Gaussian and normalized Gaussian, a 3D avatar animation with a free viewpoint is rendered.
- Emotional control:By manually setting emotional tags and varying their time series, the emotional expression of the generated avatars can be controlled.
EmoTalk3D project address
-
Project official website:https://nju3dv.github.io/projects/EmoTalk3D
-
arXivTechnical Papers:https://arxiv.org/abs/2408.00297
Application Scenarios of EmoTalk3D
- Virtual assistants and customer serviceAs an intelligent customer service representative or virtual assistant, it provides a more natural and emotionally rich interactive experience.
- Film and video productionGenerate realistic characters and animations in movies, television, and video games to enhance the visual experience.
- Virtual Reality (VR) and Augmented Reality (AR)Provide immersive experiences in VR and AR applications, enabling more realistic interactions with users.
- Social media and live streamingUsers can create and customize their own 3D avatars using EmoTalk3D for use on social media platforms or live streams.
- Advertising and MarketingCreate appealing 3D characters for advertising or brand promotion.