AB
AiBoss
project

DCEdit - A dual-layer control image editing method jointly developed by Beijing Jiaotong University and Meitu.

DCEdit is a novel two-layer control image editing method jointly developed by Beijing Jiaotong University and Meitu 2MT Lab. DCEdit is based on the Precise Semantic Localization (PSL) strategy, using visual and text self-attention to optimize cross-attention...

What is DCEdit?

DCEdit is a novel two-layer control image editing method jointly developed by Beijing Jiaotong University and Meitu 2MT Lab. Based on the Precise Semantic Localization (PSL) strategy, DCEdit optimizes the cross-attention map using visual and textual self-attention, providing more accurate region cues to guide image editing. DCEdit introduces a two-layer control mechanism (DLC), incorporating region cues simultaneously in the feature layer and latent space layer to achieve finer editing control. DCEdit requires no additional training or fine-tuning and performs excellently in background preservation and editing accuracy when applied to existing diffusion transformer (DiT) based editing methods.

DCEdit's main functions

  • Precise semantic positioningIt precisely locates the semantic regions in an image that need to be edited, while preserving details of the background and other unedited areas.
  • Two-layer control mechanismBy incorporating regional cues into both the feature layer and the latent space layer, fine-grained control over the editing process can be achieved, thereby improving the editing effect.
  • Supports complex image editingSuitable for high-resolution, complex background real-world images, it supports various editing tasks, such as changing colors, replacing objects, adding or deleting objects, etc.

DCEdit's technical principles

  • Precise Semantic Locator (PSL)This approach combines visual and textual self-attention to optimize the cross-attention map. The visual self-attention matrix captures affinity relationships within an image, while the textual self-attention matrix decouples semantic entanglement. Based on the reweighting of the visual self-attention matrix and the inverse operation of the textual self-attention matrix, the cross-attention map is optimized to more accurately reflect the target semantic region. The optimized cross-attention map serves as a region cue, guiding the editing process and ensuring that editing effects are focused on the target region.
  • Two-layer control mechanism (DLC)In the feature layer, based on a soft fusion mechanism, optimized cross-attention maps selectively retain features activated by the edited text, avoiding the loss of editing effects caused by direct feature replacement. In the latent space layer, based on a diffusion fusion method, binarized cross-attention maps preserve background information, preventing background regions from being incorrectly edited. The inversion process maps the source image to initial noise, and a two-layer control mechanism is applied during sampling to generate the edited image.
  • RW-800 StandardIncludes high-resolution real-world images to ensure the diversity and complexity of test data. Provides detailed text descriptions to support complex editing tasks.

DCEdit project address

Application scenarios of DCEdit

  • Advertising and MarketingQuickly modify elements in advertising images (such as colors, backgrounds, logos, etc.) to improve production efficiency.
  • Film and EntertainmentIt allows for convenient adjustment of props, costumes, or backgrounds in film and television scenes, saving time and costs.
  • Social media and content creationQuickly modify images based on the theme to enhance the appeal and diversity of your content.
  • Product Design and DevelopmentIt can quickly generate different product design schemes and accelerate the development process.
  • Education and TrainingCreate personalized learning materials to help students better understand the teaching content.