AB
AiBoss
project

FantasyTalking - Alibaba and Beijing University of Posts and Telecommunications jointly launch a framework for generating controllable digital humans from static portraits.

FantasyTalking is a novel framework jointly proposed by Alibaba's AMAP team and Beijing University of Posts and Telecommunications for generating realistic, animable virtual avatars from single static portraits. Based on a pre-trained video diffusion transformer model,...

What is FantasyTalking?

FantasyTalking is a novel framework jointly proposed by Alibaba's AMAP team and Beijing University of Posts and Telecommunications for generating realistic, animable virtual avatars from single static portraits. Based on a pre-trained video diffusion transformer model, it employs a two-stage audiovisual alignment strategy. The first stage establishes coherent global motion through a segment-level training scheme, while the second stage refines lip movements at the frame level using lip tracking masks to ensure precise synchronization with the audio signal. The framework introduces a cross-attention module for facial focus to maintain facial consistency and a motion intensity modulation module to control the intensity of facial expressions and body movements.

Main features of FantasyTalking

  • Lip-syncIt can accurately recognize and synchronize the lip movements of virtual characters with the input voice, making the lip movements of the characters when they speak completely consistent with the voice content, thus enhancing the realism and credibility of the characters.
  • Facial motion generationBased on the voice content and emotional information, corresponding facial movements are generated, such as blinking, frowning, and smiling, making the virtual character's expressions richer and more vivid.
  • Full-body motion generationIt can generate full-body movements and postures, such as walking, running, and jumping, according to the needs of the scene and plot, making the virtual character more natural and fluid in the animation.
  • Exercise intensity controlThrough the motion intensity modulation module, users can explicitly control the intensity of facial expressions and body movements, enabling controllable manipulation of portrait movements, not just lip movements.
  • Multiple styles supportedIt supports various styles of virtual avatars, including realistic and cartoon styles, and can generate high-quality dialogue videos.
  • Multiple posture supportSupports the generation of realistic speaking videos with various body ranges and orientations, including close-up portraits, half-body, full-body, and front and side poses.

The technical principles of FantasyTalking

  • Two-stage audiovisual alignment strategy
    • Fragment-level trainingIn the first stage, through a segment-level training scheme, the model captures the weak correlation between audio and the entire scene (including reference portraits, contextual objects, and background), establishing a global audiovisual dependency and achieving overall feature fusion. This enables the model to learn audio-related nonverbal cues (such as eyebrow movements and shoulder actions) and strongly audio-synchronized lip dynamics.
    • Frame-level trainingIn the second stage, the model focuses on refining visual features that are highly correlated with the audio at the frame level, particularly lip movements. By using lip tracking masks, the model ensures that lip movements are precisely aligned with the audio signal, improving the quality of the generated video.
  • Identity preservationTraditional reference network methods often restrict the large-scale natural variations of characters and backgrounds in videos. FantasyTalking employs a face-focused cross-attention module, centrally modeling facial regions and decoupling identity preservation from motion generation through a cross-attention mechanism. This is more lightweight, freeing it from constraints on the natural movement of backgrounds and characters, ensuring that the character's identity is maintained throughout the generated video sequence.
  • Exercise intensity adjustmentFantasyTalking introduces a motion intensity modulation module that allows explicit control over the intensity of facial expressions and body movements. This enables users to manipulate portrait movement, not just lip movements. By adjusting the motion intensity, more natural and diverse animations can be generated.
  • Based on a pre-trained video diffusion transformer modelFantasyTalking, based on the Wan2.1 video diffusion transformer model and leveraging its spatiotemporal modeling capabilities, generates high-fidelity, coherent speaking portrait videos. The model effectively captures the relationship between audio signals and lip movements, facial expressions, and body motions, producing high-quality dynamic portraits.

FantasyTalking's project address

Application scenarios of FantasyTalking

  • Game developmentIn game development, FantasyTalking can be used to generate dialogue and combat animations for game characters. It can generate precise lip-sync, rich facial expressions, and natural full-body movements based on voice content, making game characters more vivid and realistic, enhancing the game's visual effects and player immersion.
  • Film and television productionIn film and television production, it can be used to generate performance animations and special effects animations for virtual characters. FantasyTalking allows for the rapid generation of virtual characters with complex expressions and movements, reducing the manpower and time costs of traditional animation production and adding more creativity and imagination to film and television works.
  • Virtual Reality and Augmented RealityIn virtual reality (VR) and augmented reality (AR) applications, FantasyTalking can generate interactive and guided animations for virtual characters.
  • Virtual streamerFantasyTalking can be used to generate animated videos of virtual anchors. It supports various styles of virtual avatars and can be used in various scenarios such as news broadcasting, live e-commerce, and online education, offering high practicality and flexibility.
  • Smart EducationIn the field of smart education, FantasyTalking can generate animated videos of virtual teachers or virtual teaching assistants.