AB
AiBoss
project

3DV-TON - A video virtual try-on framework launched by Alibaba DAMO Academy in collaboration with Zhejiang University and other institutions.

3DV-TON (Textured 3D-Guided Consistent Video Try-on via Diffusion Models) is a video virtual try-on based on diffusion models, jointly developed by Alibaba DAMO Academy, Lakeside Lab, and Zhejiang University...

What is 3DV-TON?

3DV-TON (Textured 3D-Guided Consistent Video Try-on via Diffusion Models) is a video virtual try-on framework based on diffusion models, jointly developed by Alibaba DAMO Academy, Lakeside Lab, and Zhejiang University. It addresses the issue of poor generation results from existing methods when handling complex clothing patterns and diverse human poses. The framework uses generated, animable, textured 3D meshes as explicit frame-level guidance, ensuring excellent visual quality and temporal consistency in the generated try-on videos. 3DV-TON incorporates the high-resolution benchmark dataset HR-VVT, advancing research in video try-on technology.

Main functions of 3DV-TON

  • High-fidelity visual effectsIt accurately reproduces clothing details and generates realistic try-on effects.
  • Time Consistency: Ensure that the clothing texture in the video maintains a consistent motion between different frames to avoid artifacts or distortions.
  • Adapt to complex scenariosIt supports handling diverse clothing types, complex human postures, and dynamic scenes.
  • Provide benchmark datasets: Introducing the high-resolution video try-on benchmark dataset HR-VVT to promote research and evaluation in related fields.

3DV-TON Technical Principles

  • Textured 3D guidanceSingle-image 3D reconstruction technology generates animable textured 3D meshes. Synchronizing the 3D mesh with the pose of the original video provides explicit frame-level guidance for the diffusion model, ensuring consistency in appearance and motion of the generated try-on results.
  • Dynamic 3D Guided PipelineSelect keyframes for initial 2D image try-on, then reconstruct an animated, textured 3D mesh. Optimize SMPL-X parameters to ensure precise alignment of the 3D mesh with the human pose.
  • Rectangular masking strategyTo prevent the leakage of clothing information and avoid artifacts in dynamic human and clothing movements, this method combines clothing images and try-on images as references to provide contextual information and enhance the generation effect.
  • Diffusion Model ArchitectureBased on Stable Diffusion, the UNet architecture is extended to support pseudo-3D structures. Realistic motion generation is achieved through temporal module integration, reducing reliance on explicit optical flow or deformation operations.
  • Training strategyThe model is trained using a combination of image and video data, balancing image quality and temporal consistency based on randomly selected data types. A classifier free-guided (CFG) strategy is employed to randomly omit certain conditional inputs, enhancing the model's robustness.

3DV-TON project address

Application scenarios of 3DV-TON

  • Online shoppingIt helps users virtually try on clothes, improving the shopping experience and reducing returns.
  • Fashion DesignQuickly showcase clothing design effects to assist in design and marketing.
  • Virtual fitting roomSave time and effort on trying on clothes in physical stores.
  • Film and games: Assists in character costume design and customization, improving production efficiency.
  • social mediaIt provides users with fun tools for creating and sharing try-on videos.