AB
AiBoss
project

DreamFit - A virtual fitting framework launched by ByteDance in collaboration with Tsinghua University and Sun Yat-sen University

DreamFit is a virtual fitting framework developed by ByteDance in collaboration with Tsinghua University Shenzhen International Graduate School and Sun Yat-sen University Shenzhen Campus. It's specifically designed for generating human images centered around lightweight clothing. Based on adaptive attention and L...

What is DreamFit?

DreamFit is a virtual try-on framework developed by ByteDance in collaboration with Tsinghua University Shenzhen International Graduate School and Sun Yat-sen University Shenzhen Campus. It's specifically designed for generating human images centered around lightweight clothing. The framework significantly reduces model complexity and training costs, improving the quality and consistency of generated images through optimized text prompts and feature fusion. DreamFit generalizes to various clothing styles and prompts to generate high-quality human images. DreamFit supports seamless integration with community control plugins, lowering the barrier to entry.

DreamFit's main functions

  • Plug and playIt is easy to integrate with community control plugins, lowering the barrier to entry for users.
  • High-quality generationBased on rich prompts from a large multimodal model, highly consistent images are generated.
  • Posture controlSupports specifying a person's pose and generating images that match that pose.
  • Multi-themed clothing migrationThis technique combines multiple clothing elements into a single image, making it suitable for scenarios such as e-commerce clothing displays.

DreamFit's technical principles

  • Lightweight encoder (Anything-Dressing Encoder)Based on the LoRA layer, a pre-trained diffusion model (such as Stable Diffusion's UNet) is extended into a lightweight clothing feature extractor. By training only the LoRA layer instead of the entire UNet, the model complexity and training cost are greatly reduced.
  • Adaptive AttentionTwo trainable linear projection layers are introduced to align reference image features with potential noise. Based on an adaptive attention mechanism, reference image features are seamlessly injected into UNet, ensuring that the generated image is highly consistent with the reference image.
  • Pre-trained multimodal models (LMMs)During the inference phase, LMMs are used to rewrite the text prompts for user input, increasing the fine-grained description of the reference image and reducing the difference in text prompts between the training and inference phases.

DreamFit's project address

DreamFit application scenarios

  • Virtual try-onConsumers can virtually try on clothes online, saving time and costs and enhancing their shopping experience.
  • Fashion DesignDesigners can quickly generate clothing renderings, accelerating the design process and improving work efficiency.
  • Personalized advertisingGenerate customized ads based on user preferences to improve ad appeal and conversion rates.
  • Virtual Reality (VR) / Augmented Reality (AR)It provides a virtual try-on experience, enhancing user immersion and interactivity.
  • Social media content creationGenerate personalized images to attract more attention and enhance the diversity and appeal of your content.