AB
AiBoss
project

OmniAvatar - An audio-driven full-body video generation model jointly developed by Zhejiang University and Alibaba.

OmniAvatar is an audio-driven full-body video generation model jointly developed by Zhejiang University and Alibaba Group. Based on the input audio and text prompts, the model generates natural and realistic full-body animated videos, with character movements perfectly synchronized with the audio...

What is OmniAvatar?

OmniAvatar is an audio-driven full-body video generation model jointly developed by Zhejiang University and Alibaba Group. Based on input audio and text prompts, the model generates natural and realistic full-body animated videos with perfect synchronization between character movements and audio, and rich facial expressions. The model is based on a pixel-level multi-level audio embedding strategy and the LoRA training method, effectively improving lip-sync accuracy and the naturalness of full-body movements. It supports features such as character-object interaction, background control, and emotion control, and is widely used in podcasts, interactive videos, virtual scenes, and other fields.

OmniAvatar's main functions

  • Natural Lip SynchronizationIt can generate lip movements that are perfectly synchronized with the audio, maintaining high accuracy even in complex scenarios.
  • Full-body animation generationIt supports generating natural and fluid full-body movements, making animations more vivid and realistic.
  • Text controlIt allows for precise control of video content based on text prompts, including character actions, background, and emotions, enabling highly customized video generation.
  • Interaction between people and objectsIt supports generating scenes where characters interact with surrounding objects, such as picking up items or operating equipment, thus expanding the scope of applications.
  • Background controlThe background can be changed according to text prompts to adapt to various scene requirements.
  • Emotional controlIt controls the expression of characters' emotions, such as happiness, sadness, and anger, based on text prompts, thereby enhancing the expressiveness of the video.

OmniAvatar's technical principles

  • Pixel-level multi-level audio embedding strategyThis method maps audio features into the model's latent space and embeds them at the pixel level, allowing audio features to more naturally influence the generation of full-body movements, thus improving the accuracy of lip synchronization and the naturalness of full-body movements.
  • LoRA training methodThis method fine-tunes pre-trained models using Low-Rank Adaptation (LoRA) technology. By introducing low-rank decomposition into the model's weight matrix, the number of training parameters is reduced while preserving the model's original capabilities, thus improving training efficiency and generation quality.
  • Long video generation strategyTo generate long videos, OmniAvatar uses a reference image embedding and frame overlay strategy. Reference image embedding ensures consistency in the identities of people in the video, while frame overlay guarantees temporal continuity and avoids abrupt changes in action.
  • Video generation based on diffusion modelBased on diffusion models, this method gradually removes noise to generate video. It produces high-quality video content and performs exceptionally well when processing long data sequences.
  • Transformer architectureBased on the diffusion model, the Transformer architecture is introduced to better capture long-term dependencies and semantic consistency in videos, further improving the quality and coherence of generated videos.

OmniAvatar project address

  • Project official websitehttps://omni-avatar.github.io/
  • GitHub repository: https://github.com/Omni-Avatar/OmniAvatar
  • HuggingFace model libraryhttps://huggingface.co/OmniAvatar/OmniAvatar-14B
  • arXiv technical paper: https://arxiv.org/pdf/2506.18866

Application scenarios of OmniAvatar

  • Virtual content creationIt can be used to generate virtual avatars for podcasters, video bloggers, etc., reducing production costs and enriching the forms of content presentation.
  • Interactive social platformIn virtual social scenarios, it provides users with personalized virtual avatars, enabling natural interaction through actions and expressions.
  • Education and training sectorIt generates virtual teacher avatars to explain teaching content based on audio input, thereby enhancing the fun and appeal of teaching.
  • Advertising and MarketingGenerate virtual spokesperson images, customize images and actions according to brand needs, and achieve precise advertising.
  • Games and Virtual RealityIt can quickly generate virtual game characters with natural movements and expressions, enriching game content and enhancing the realism of the virtual reality experience.