AB
AiBoss
project

Playmate - A face animation generation framework developed by the FunWan Technology team.

Playmate is a face animation generation framework developed by the Guangzhou Quwan Technology team. Based on a 3D implicit space guided diffusion model, the framework uses a two-stage training process to precisely control facial expressions and head postures based on audio and commands, generating...

What is Playmate?

Playmate is a face animation generation framework developed by the Guangzhou Quwan Technology team. Based on a 3D implicit space guided diffusion model, the framework uses a two-stage training process to precisely control facial expressions and head postures according to audio and instructions, generating high-quality dynamic portrait videos. Playmate utilizes motion decoupling and emotion control modules to achieve fine-grained control over the generated videos, significantly improving video quality and the flexibility of emotional expression. Playmate represents a major advancement in audio-driven portrait animation, providing precise control over emotions and postures, generating dynamic portraits in various styles, and has broad application prospects.

Playmate's main functions

  • audio driverWith just a still photo and an audio clip, a corresponding dynamic portrait video can be generated, achieving natural lip-sync and facial expression changes.
  • Emotional controlGenerate dynamic videos with specific emotions based on specified emotional conditions (such as anger, disgust, contempt, fear, happiness, sadness, surprise, etc.).
  • Attitude controlIt supports the generation of postures based on the driving image control results, enabling various head movements and poses.
  • Independent controlIt enables independent control over facial expressions, lip movements, and head posture.
  • Diverse stylesIt can generate dynamic portraits in various styles, including realistic faces, animations, artistic portraits, and even animals, making it widely applicable.

Playmate's technical principles

  • 3D Implicit Space Guided Diffusion ModelBased on 3D implicit spatial representation, facial attributes (such as expression, lip shape, head posture, etc.) are decoupled. Based on an adaptive normalization strategy, the decoupling accuracy of motion attributes is further improved, ensuring that the generated video is more natural in terms of expression and posture.
  • Two-stage training framework:
    • Phase 1Train an audio conditional diffusion transformer to generate motion sequences directly from audio cues. Based on a motion decoupling module, achieve accurate decoupling of facial expressions, lip movements, and head posture.
    • Phase TwoAn emotion control module is introduced to encode emotional conditions into the latent space, enabling fine-grained emotion control over the generated video.
  • Emotional control moduleThe emotion control module is implemented based on DiT (Diffusion Transformer Blocks). Using a two-DiT block structure, emotional conditions are integrated into the generation process, enabling fine-grained control over emotions. A Classifier-Free Guidance (CFG) strategy is employed, adjusting CFG weights to balance the quality and diversity of the generated videos.
  • Efficient diffusion model trainingAudio features are extracted using a pre-trained Wav2Vec2 model, and audio and motion features are aligned based on a self-attention mechanism. Gaussian noise is progressively added to the target motion data based on forward and reverse Markov chains, and a diffusion transformer is used to predict and remove the noise, generating the final motion sequence.

Playmate's project address

Playmate application scenarios

  • Film and television productionGenerate virtual character animations, enhance special effects, and replace characters, reducing manual production costs and improving the realism of special effects.
  • Game developmentIt helps generate virtual characters, create interactive storylines, and produce NPC animations, enhancing game interactivity and immersion.
  • Virtual Reality (VR) and Augmented Reality (AR)To achieve natural facial expressions and lip-syncing in virtual character interactions, virtual meetings, and virtual social interactions, thereby enhancing the user experience.
  • Interactive MediaIt can be used in live streaming, video conferencing, virtual anchors, and interactive advertising to make content more vivid and interesting and enhance interactivity.
  • Education and trainingIt can be used for virtual teacher generation, simulation training, and language learning to make teaching content more attractive to students and provide a realistic training environment.