AB
AiBoss
project

InfiniteTalk - Meituan's open-source digital human video generation framework

InfiniteTalk is a new digital human-driven technology launched by Meituan's Visual Intelligence Department. Through a sparse frame video dubbing paradigm, it requires only a small number of keyframes to drive the digital human to generate natural and smooth videos, solving the problem of speech dubbing in traditional technologies...

What is InfiniteTalk?

InfiniteTalk is a new digital human-driven technology launched by Meituan's Visual Intelligence Department. Through a sparse frame video dubbing paradigm, it requires only a small number of keyframes to drive the generation of natural and smooth videos from digital humans, solving the problem of disconnect between lip movements, facial expressions, and body movements in traditional technologies. InfiniteTalk makes digital human videos more immersive and natural, with high generation efficiency and low cost. The InfiniteTalk paper, code, and weights are open-source, providing important reference for the development of digital human technology.

InfiniteTalk's main functions

  • High-efficiency driving of virtual humansWith only a few keyframes, it can accurately drive virtual human to generate natural and smooth videos, achieving perfect synchronization of lip movements, facial expressions, and body movements.
  • Diverse scene adaptationIt is applicable to various scenarios such as virtual anchors, customer service, and actors, providing efficient and low-cost virtual human solutions for different industries.
  • High-efficiency video generationHigh-quality videos can be generated quickly by using sparse frame driving and time interpolation techniques, significantly reducing production costs and time.

InfiniteTalk's technical principles

  • Sparse frame video dubbing paradigmBased on a sparse frame-driven approach, only a small number of keyframes are needed to capture changes in a person's lip movements, facial expressions, and actions. Keyframes contain key information about the person's movements and facial expressions. Through appropriate temporal interpolation, intermediate frames can be generated to achieve a complete video sequence. Advanced temporal interpolation algorithms are used to appropriately fill the time intervals between keyframes. Simultaneously, fusion techniques are used to naturally transition the actions, expressions, and lip movements from keyframes into intermediate frames, generating coherent video content.
  • Multimodal fusion and optimizationThis involves fusing text, audio, and visual information. For example, speech recognition technology extracts speech content from audio and combines it with text information to more accurately control the virtual human's lip movements and facial expressions. Based on optimization algorithms in deep learning, the virtual human's movements, facial expressions, and lip movements are fine-tuned to ensure high consistency with the input audio and text, thereby enhancing the naturalness and realism of the video.
  • High-efficiency computing architectureWe construct lightweight deep learning models to reduce computational resource consumption while ensuring model performance. We use parallel computing techniques to process multiple tasks in the video generation process in parallel, further improving the speed and efficiency of video generation.

InfiniteTalk project address

  • Project official websitehttps://meigen-ai.github.io/InfiniteTalk/
  • GitHub repositoryhttps://github.com/MeiGen-AI/InfiniteTalk
  • HuggingFace model libraryhttps://huggingface.co/MeiGen-AI/InfiniteTalk
  • arXiv technical paper: https://arxiv.org/pdf/2508.14033

InfiniteTalk application scenarios

  • Virtual streamerIt provides virtual anchors for news, variety shows, live broadcasts, and other programs, enabling 24-hour uninterrupted broadcasting and improving program efficiency and entertainment.
  • Film and television productionIn film and television production, it is used for the rapid generation and motion capture of virtual characters, reducing production costs and time.
  • Game developmentIt helps generate virtual characters in games, improves the naturalness and smoothness of character movements, and enhances the immersion of the game.
  • Online EducationCreate virtual teachers to provide students with personalized teaching services, such as online Q&A and course explanations, to improve teaching effectiveness.
  • Training simulationVirtual scenario simulations used in corporate training, such as customer service training and sales training, allow employees to practice and learn in a virtual environment.