AB
AiBoss
project

DICE-Talk - An emotional dynamic portrait generation framework jointly launched by Fudan University and Tencent YouTu.

DICE-Talk is a novel emotional dynamic portrait generation framework jointly developed by Fudan University and Tencent YouTu Lab. It supports the generation of dynamic portrait videos with vivid emotional expression while maintaining identity consistency. DICE-Talk introduces emotional...

What is DICE-Talk?

DICE-Talk is a novel emotional dynamic portrait generation framework developed by Fudan University in collaboration with Tencent YouTu Lab. It supports the generation of dynamic portrait videos with vivid emotional expression while maintaining identity consistency. DICE-Talk introduces an emotion association enhancement module, capturing the relationships between different emotions based on an emotion database to improve the accuracy and diversity of emotion generation. The framework is designed with an emotion discrimination objective, ensuring emotional consistency during the generation process based on emotion classification. Experiments on the MEAD and HDTF datasets demonstrate that DICE-Talk outperforms existing technologies in terms of emotional accuracy, lip-sync accuracy, and visual quality.

Main functions of DICE-Talk

  • Emotional dynamic portrait generationGenerate dynamic portrait videos with specific emotional expressions based on input audio and reference images.
  • Identity preservationWhen generating emotional videos, preserve the identity features of the input reference image to avoid leakage or confusion of identity information.
  • High-quality video generationThe generated videos achieved a high level of visual quality, lip-sync, and emotional expression.
  • Generalization abilityIt can adapt to unfamiliar combinations of identities and emotions and has a good generalization ability.
  • User ControlUsers input specific emotional goals and control the emotional expression of the generated video, achieving a high degree of user customization.
  • Multimodal inputIt supports multiple input modalities, including audio, video, and reference images.

DICE-Talk's Technical Principles

  • Decoupling identity and emotionThis approach jointly models audio and visual sentiment cues based on a cross-modal attention mechanism, representing sentiment as an identity-independent Gaussian distribution. A sentiment embedding engine is trained using contrastive learning (such as InfoNCE loss) to ensure that features of the same sentiment cluster in the embedding space, while features of different sentiments are dispersed.
  • Enhanced emotional connectionThe sentiment database is a learnable module that stores feature representations of various emotions. It learns relationships between emotions using vector quantization and attention-based feature aggregation. The sentiment database stores features of individual emotions and learns the connections between emotions, helping the model generate other emotions more effectively.
  • Emotional discrimination targetDuring the generation process of the diffusion model, an emotion discriminator is used to ensure the emotional consistency of the generated videos. The emotion discriminator and the diffusion model are jointly trained to ensure that the generated videos are consistent with the target emotion in terms of emotional expression, maintaining visual quality and lip synchronization.
  • Diffusion Model FrameworkStarting with Gaussian noise, the target video is generated through progressive denoising. Based on a variational autoencoder (VAE), video frames are mapped to a latent space, where Gaussian noise is progressively introduced. A diffusion model is then used to gradually remove the noise, generating the target video. During denoising, the diffusion model utilizes a cross-modal attention mechanism, combining reference images, audio features, and emotional features to guide video generation.

DICE-Talk project address

Application Scenarios of DICE-Talk

  • Digital Humans and Virtual AssistantsIt endows digital humans and virtual assistants with rich emotional expression, making their interactions with users more natural and vivid, thus enhancing the user experience.
  • Film and television productionIn film and television special effects and animation production, it can quickly generate dynamic portraits with specific emotions, improving production efficiency and reducing production costs.
  • Virtual Reality and Augmented RealityIn VR/AR applications, virtual characters that interact emotionally with users are generated, enhancing immersion and emotional resonance.
  • Online Education and TrainingCreate instructional videos with emotional feedback to make learning content more vivid and interesting, and improve learning outcomes.
  • Mental health supportDevelop emotionally resonant virtual characters for use in psychotherapy and emotional support, helping users better express and understand emotions.