MEMO - An audio-driven generative portrait speaking video framework that maintains identity consistency and expressiveness.
MEMO (Memory-Guided EMOtionaware diffusion) is an audio-driven portrait animation framework developed by Skywork AI, Nanyang Technological University, and the National University of Singapore. It is used to generate images with consistent identity and expressiveness...
What is MEMO?
MEMO (Memory-Guided EMOtionaware diffusion) is an audio-driven portrait animation framework developed by Skywork AI, Nanyang Technological University, and the National University of Singapore for generating speaking videos with identity consistency and expressiveness. MEMO is built around two core modules: a memory-guided temporal module and an emotion-aware audio module. The memory-guided module enhances identity consistency and motion smoothness by storing longer-term motion information, while the emotion-aware module uses a multimodal attention mechanism to improve audio-video interaction and refine facial expressions based on emotions in the audio. MEMO demonstrates superior overall quality, audio-lip synchronization, identity consistency, and expression-emotion alignment compared to state-of-the-art methods across various image and audio types of speaking videos.
The main functions of MEMO
- Audio-driven portrait animationMEMO generates synchronized, identity-consistent speaking video based on the input audio and reference images.
- Diverse content generationIt supports the generation of spoken videos in various image styles (such as portraits, sculptures, and digital art) and audio types (such as speeches, singing, and rapping).
- Multilingual supportIt can handle audio input in multiple languages, including English, Mandarin, Spanish, Japanese, Korean, and Cantonese.
- Generate expressive videosGenerate speaking videos with corresponding facial expressions based on the emotional content of the audio.
- Long video generation capabilityIt can generate long-duration speaking videos with minimal error accumulation.
MEMO's technical principles
- Memory-guided time module:
- Memory state: Develop memory state storage to provide information from longer past contexts to guide temporal modeling.
- Linear attentionBased on the linear attention mechanism, long-term motion information is used to improve the coherence of facial movements and reduce error accumulation.
- Emotion-sensing audio module:
- Multimodal attentionIt can process both video and audio inputs simultaneously, enhancing the interaction between the two.
- Audio Emotion DetectionIt dynamically detects emotional cues in audio, integrates emotional information into the video generation process, and refines facial expressions.
- end-to-end framework:
- Reference NetProvides identity information for use in spatial and temporal modeling.
- Diffusion NetThe core innovations include a time-based module for memory guidance and an audio module for emotional perception.
- Data processing flowThis includes steps such as scene transition detection, face detection, image quality assessment, and audio-lip synchronization detection to ensure data quality.
- Training strategyThe training is divided into two phases: robust training for facial domain adaptation and emotional decoupling, using modified flow loss.
MEMO's project address
- Project official website:memoavatar.github.io
- GitHub repository:https://github.com/memoavatar/memo
- HuggingFace model library:https://huggingface.co/memoavatar/memo
- arXiv technical paper:https://arxiv.org/pdf/2412.04448
MEMO application scenarios
- Virtual assistants and chatbotsGenerate realistic videos of virtual assistants or chatbots to make interactions with users more natural and personal.
- Entertainment and social mediaIn the entertainment industry, creating dynamic video content featuring virtual idols, game characters, or social media influencers.
- Education and trainingGenerate educational videos in which the image of the teacher or trainer changes dynamically according to the teaching content, enhancing the interactivity and appeal of the learning experience.
- News and MediaIn news broadcasting, generate anchor videos, especially when multilingual broadcasting is required, quickly generate anchor videos in the corresponding language.
- Advertising and MarketingCreate customized advertising videos, with product spokespeople tailored to different markets and target audiences.