AB
AiBoss
project

InfinityHuman - An AI-powered digital human video generation model jointly developed by ByteDance and Zhejiang University.

InfinityHuman is a commercial-grade long-sequence audio-driven human video generation model launched by a joint team from ByteDance and Zhejiang University, opening a new chapter in the practical application of AI digital humans.

What is InfinityHuman?

InfinityHuman, a commercial-grade long-term audio-driven human video generation model developed by a joint team from ByteDance and Zhejiang University, opens a new chapter in the practical application of AI digital humans. Based on a coarse-to-fine framework, the model generates low-resolution motion representations and gradually generates high-resolution long-term videos through a pose-guided refiner. The model introduces a hand-specific reward mechanism to optimize the naturalness and synchronization of hand movements, effectively solving common problems in existing methods such as identity drift, unstable visuals, and stiff hand movements. InfinityHuman demonstrates outstanding performance on the EMTD and HDTF datasets, providing new possibilities for applications in virtual broadcasting, education, customer service, and other fields.

Main functions of InfinityHuman

  • Long-duration video generationIt can generate high-resolution, long-duration human body animation videos while maintaining visual consistency and stability.
  • Natural hand movementsIt generates natural, accurate, and voice-synchronized hand gestures through a hand-specific reward mechanism.
  • Identity ConsistencyBy using a pose-guided refiner and the first frame as visual anchors, cumulative errors are reduced, and long-term consistency of character identity is maintained.
  • Lip-sync: Ensure that the lip movements of the characters in the generated video are highly synchronized with the audio to enhance realism.
  • Diverse character stylesIt supports the generation of characters in different styles to meet the needs of various application scenarios.

InfinityHuman's technical principles

  • Low-resolution action representation generationThe model generates low-resolution pose representations synchronized with the audio through audio-driven generation, which is equivalent to "laying the groundwork" to ensure that the global rhythm, movement and lip movements are aligned in advance.
  • Pose-Guided RefinerBased on the generation of low-resolution motion representations, the model uses a pose-guided refiner to gradually generate high-resolution video.
    • pose sequencePose sequences serve as stable intermediate representations, resisting temporal degradation and maintaining visual consistency.
    • visual anchorThe first frame serves as a visual anchor point, continuously referencing and correcting the identity and image to reduce accumulated errors.
    • Hand reward mechanismBy training with high-quality hand movement data, a hand-specific reward mechanism is introduced to optimize the naturalness of hand movements and their synchronization with speech.
  • Multimodal conditional fusionThe model integrates multiple modal information, including reference images, text prompts, and audio, to ensure the visual and auditory consistency and naturalness of the generated video.

InfinityHuman's project address

  • Project official websitehttps://infinityhuman.github.io/
  • arXiv technical paperhttps://arxiv.org/pdf/2508.20210

Application scenarios of InfinityHuman

  • Virtual streamerVirtual anchors can deliver news broadcasts and host programs naturally and smoothly, enhancing the viewing experience and reducing labor costs.
  • Online EducationAI teachers can make corresponding gestures while explaining knowledge, making the teaching process more vivid and engaging, and improving students' learning interest and concentration.
  • Customer serviceDigital customer service representatives can respond naturally during voice communication, breaking away from the mechanical feel of traditional customer service and improving customer satisfaction.
  • Film and television productionIt can quickly generate high-quality, long-running character animations in animated films, TV series, and other film and television works, reducing the workload of manual drawing and post-production restoration.
  • Virtual socialIt gives virtual characters in virtual reality (VR) and augmented reality (AR) natural movements and expressions, making virtual social interaction more realistic and immersive, and enhancing the interactivity between users.