AB
AiBoss
project

AniTalker - An open-source lip-syncing video generation framework from Shanghai Jiao Tong University

AniTalker is an AI-powered lip-syncing video generation framework developed by researchers from the X-LANCE Lab at Shanghai Jiao Tong University and AISpeech. It can convert a single static image of a person and input audio into lifelike animation...

What is AniTalker?

AniTalker, developed by researchers from the X-LANCE Lab at Shanghai Jiao Tong University and AISpeech, is an AI-powered lip-syncing video generation framework capable of converting a single static portrait and input audio into lifelike animated dialogue videos. The framework captures complex facial dynamics, including subtle expressions and head movements, through a self-supervised learning strategy. AniTalker leverages general motion representation and identity decoupling techniques to reduce reliance on labeled data, while combining a diffusion model and variance adapter to generate diverse and controllable facial animations, achieving effects similar to Alibaba's EMO and Tencent's AniPortrait.

AniTalker's main functions

  • Animation of static portraitsAniTalker can convert any single facial portrait into a dynamic video in which the person can speak and change facial expressions.
  • Audio synchronizationThis framework can synchronize the input audio with the character's lip movements and speech rhythm to achieve a natural dialogue effect.
  • Facial motion captureBeyond just lip-syncing, AniTalker can also simulate a range of complex facial expressions and subtle muscle movements.
  • Diverse animation generationBy using a diffusion model, AniTalker can generate diverse facial animations with random variations, increasing the naturalness and unpredictability of the generated content.
  • Real-time facial animation controlUsers can guide the generation of animations in real time through control signals, including but not limited to head posture, facial expressions, and eye movements.
  • Voice-driven animation generationThe framework supports generating animations directly using voice signals, without requiring additional video input.
  • Long video continuous generationAniTalker can continuously generate long animated videos, making it suitable for long conversations or speeches.

AniTalker's official website entrance

How AniTalker works

  • Motion representation learningAniTalker uses a self-supervised learning method to train a general motion encoder capable of capturing facial dynamics. This process involves selecting source and target images from a video and learning motion information by reconstructing the target image.
  • Decoupling Identity and MovementTo ensure that motion representations do not contain identity-specific information, AniTalker employs metric learning and mutual information minimization techniques. Metric learning helps the model distinguish the identity information of different individuals, while mutual information minimization ensures that the motion encoder focuses on capturing motion rather than identity features.
  • Hierarchical Aggregation Layer (HAL)The Hierarchical Aggregation Layer (HAL) is introduced to enhance the motion encoder's ability to understand motion changes at different scales. HAL integrates information from different stages of the image encoder through average pooling layers and weighted sum layers.
  • Motion generationAfter training the motion encoder, AniTalker can generate motion representations based on user-controlled drive signals. This includes both video-driven and voice-driven pipelines.
    • Video driver pipeline: Use video sequences that drive the speaker to generate animations for source images, thereby accurately replicating the driving poses and facial expressions.
    • Voice-driven pipelineUnlike video-driven methods, voice-driven methods generate video based on voice signals or other control signals, synchronized with the input audio.
  • Diffusion model and variance adapterIn its voice-driven approach, AniTalker uses a diffusion model to generate motion latent sequences and introduces attribute operations using a variance adapter, resulting in diverse and controllable facial animations.
  • Rendering moduleFinally, the final animated video is rendered frame by frame using an image renderer based on the generated potential motion sequence.
  • Training and optimizationThe training process of AniTalker includes multiple loss functions, such as reconstruction loss, perceptual loss, adversarial loss, mutual information loss, and identity metric learning loss, to optimize model performance.
  • Control attribute characteristicsAniTalker allows users to control head pose and camera parameters, such as head position and face size, to generate animations with specific attributes.

AniTalker Application Scenarios

  • Virtual assistants and customer serviceAniTalker can generate realistic virtual faces for use as virtual assistants or online customer service, providing a more natural and friendly interactive experience.
  • Film and video productionIn film post-production, AniTalker can be used to generate or edit actors' facial expressions and movements, especially for scenes that could not be captured in the original performance.
  • Game developmentGame developers can use AniTalker to create realistic facial animations for game characters, enhancing the game's immersion and the characters' expressiveness.
  • videoconferenceIn video conferencing, AniTalker can generate virtual faces for participants, especially in situations where privacy needs to be protected or fun needs to be added.
  • social mediaUsers can use AniTalker to create personalized virtual avatars and communicate and share on social media.
  • News BroadcastAniTalker can generate virtual news anchors for automated news broadcasting, especially when multilingual broadcasting is required.
  • Advertising and MarketingBusinesses can use AniTalker to generate engaging virtual avatars for advertising or brand endorsement.