AB
AiBoss
project

Sonic - An audio-driven portrait animation framework jointly developed by Tencent and Zhejiang University

Sonic is an audio-driven portrait animation framework developed by Tencent and Zhejiang University. It generates realistic facial expressions and movements based on global audio perception. Sonic utilizes context-enhanced audio learning and a motion decoupling controller to extract audio...

What is Sonic?

Sonic is an audio-driven portrait animation framework developed by Tencent and Zhejiang University. It generates realistic facial expressions and movements based on global audio perception. Sonic utilizes context-enhanced audio learning and a motion decoupling controller to extract long-term temporal audio knowledge within audio segments and independently control head and facial expression movements, enhancing local audio perception capabilities. Sonic employs a time-aware positional offset fusion mechanism to extend local audio perception to the global level, resolving jitter and abrupt changes in long video generation. Sonic outperforms existing state-of-the-art methods in video quality, lip-sync accuracy, motion diversity, and temporal coherence, significantly improving the naturalness and coherence of portrait animations and supporting fine-tuning by users.

Sonic's main functions

  • Realistic lip synchronizationPrecisely align the audio with lip movements to ensure that the spoken content is highly consistent with the mouth shape.
  • Rich facial expressions and head movementsGenerate diverse and natural facial expressions and head movements, making animations more vivid and expressive.
  • Long-term stable generationWhen processing long videos, it can maintain stable output, avoid jitter and abrupt changes, and ensure overall consistency.
  • User adjustabilityIt allows users to adjust and control head movements, facial expression intensity, and lip synchronization effects based on parameters, providing a high degree of customizability.

Sonic's technical principles

  • Context-enhanced audio learningThis approach extracts long-term audio knowledge from audio segments, transforming information such as intonation and speech rate in the audio signal into prior knowledge of facial expressions and lip movements. The Whisper-Tiny model extracts audio features and combines these features with a spatial cross-attention layer based on multi-scale understanding to guide the generation of spatial frames.
  • Motion decoupling controllerThis feature decouples head movements and facial expressions, controlling them with independent parameters to enhance the diversity and naturalness of animation. It supports user-defined exaggerated movements, controlling the amplitude of head and facial movements by adjusting motion-bucket parameters.
  • Time-aware position offset fusionA time-aware sliding window strategy extends the perception of audio segments from local to global, addressing jitter and abrupt changes in long video generation. At each time step, the model processes the audio segment from a new position, gradually fusing global audio information to ensure the continuity of the long video.
  • Global audio driverSonic relies entirely on audio signals to drive animation generation, avoiding the dependence on visual signals (such as motion frames) found in traditional methods, thus improving the naturalness and temporal consistency of the generated animation. The audio signal, as a global signal, provides implicit prior information for facial expressions and head movements, making the generated animation more consistent with the audio content.

Sonic's experimental results

  • Quantitative comparison:
    • On the HDTF and CelebV-HQ datasets, Sonic outperforms existing state-of-the-art methods on multiple evaluation metrics, including FID (Fréchet Inception Distance), FVD (Fréchet Video Distance), lip synchronization accuracy (Sync-C, Sync-D), and video smoothness.
    • Sonic's FID and FVD scores are significantly lower than other methods, indicating that it generates higher quality videos with better consistency with real data.
  • Qualitative comparisonSonic can generate more natural and diverse facial expressions and head movements, and it is particularly robust when dealing with complex backgrounds and portraits of different styles.

Sonic's generation effect

  • Comparison with open source methodsSonic can generate richer facial expressions that better match the audio, promoting more natural head movements.
  • Comparison with closed-source methods:
    • Compared with EMO
      • Sonic performs better in terms of the naturalness of facial expressions and the realism of the reflection in the glasses.
      • In singing performances, Sonic demonstrates more precise pronunciation and a wider range of movements.
    • Compared with Immediate Dream:
      • In the anime example, Sonic's lip movements and appearance are closer to the original input, and he also blinks.
      • In long video generation, Sonic is not limited by motion frames, thus avoiding artifacts at the end of the video.

Sonic's project address

Sonic's application scenarios

  • Virtual Reality (VR)Generate realistic facial expressions and lip movements for virtual characters to enhance immersion.
  • Film and television productionQuickly generate lip-sync and facial animations for characters, improving production efficiency.
  • Online EducationTransforming teachers' voices into vivid animations enhances the fun of learning.
  • Game developmentGenerate natural facial expressions and movements for game characters to enhance realism.
  • social mediaUsers can combine voice and photos to create personalized animated videos to share.