AB
AiBoss
project

MTVCrafter - A human portrait animation generation framework jointly developed by the Chinese Academy of Sciences, China Telecom, and other institutions.

MTVCrafter is a novel human image animation framework developed by the Computer Vision and Pattern Recognition Laboratory of the Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences, and the Artificial Intelligence Research Institute of China Telecom, among other institutions. It is based on original 3D motion sequences...

What is MTVCrafter?

MTVCrafter is a novel human image animation framework developed by the Computer Vision and Pattern Recognition Laboratory of the Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences, and the Artificial Intelligence Research Institute of China Telecom, among other institutions. It generates high-quality animations based on original 3D motion sequences. The framework directly models 3D motion data using 4D Motion Tagging (4DMoT), avoiding the limitations of traditional methods that rely on 2D rendering of pose images. It introduces a motion-aware video diffusion Transformer (MV-DiT), using unique 4D motion attention and positional encoding to effectively utilize 4D motion tags as the context for animation generation. MTVCrafter achieved an FID-VID score of 6.98 in the TikTok benchmark, outperforming the second-place method by 65%, demonstrating strong generalization ability and robustness.

Main functions of MTVCrafter

  • High-quality animation generationIt directly models 3D motion sequences to generate high-quality, natural, and coherent human animation videos.
  • Strong generalization abilitySupports generalization to unseen movements and characters, including single and multiple characters, full-body and half-body characters, covering a variety of styles (such as anime, pixel art, ink painting and realistic styles).
  • Precise motion controlJiyu 4D motion tagging and motion attention mechanisms enable precise control over motion sequences, ensuring the accuracy and consistency of animation.
  • Maintaining identity consistencyDuring animation generation, maintain the identity characteristics of the reference image to avoid identity drift or distortion.

MTVCrafter's technical principles

  • 4D Motion Marker (4DMoT)4DMoT uses an encoder-decoder architecture, processing temporal (frame) and spatial (joint) dimensions of data based on 2D convolution and residual blocks. It uses a vector quantizer to map continuous motion features to a discrete label space. The labels are represented in a unified space, which facilitates subsequent animation generation.
  • Motion-aware video diffusion Transformer (MV-DiT)This paper designs a 4D motion attention mechanism that combines 4D motion markers with visual markers (such as video frames). Based on 4D Rotational Position Encoding (RoPE), it recovers the spatiotemporal relationships lost due to markerization and flattening. A motion-aware classifier is introduced for free guidance, improving generation quality and generalization ability by learning joint representations of unconditional and conditional generation. A simple yet effective repetition and stitching strategy is used to combine reference images with latent variables from noisy video to ensure identity consistency.

MTVCrafter project address

Application scenarios of MTVCrafter

  • Digital Human AnimationIt generates natural and fluid movements and expressions for digital humans such as virtual anchors, customer service representatives, and idols.
  • Virtual try-onCombine user photos and clothing to generate dynamic try-on effects, enhancing the shopping experience.
  • Immersive contentGenerate virtual character animations in VR and AR that are synchronized with the user's movements, enhancing immersion.
  • Film and television special effectsIt can quickly generate high-quality character animations, reduce production costs, and enhance special effects.
  • social mediaThis allows users to create personalized animations by combining photos and actions, increasing the fun of the content.