Qwen-Image-Edit - An all-in-one image editing model launched by Alitongyi
Qwen-Image-Edit is a versatile image editing model built on the Qwen-Image architecture with 20 billion parameters. The model possesses both semantic and visual editing capabilities, enabling low-level visual editing (such as adding, deleting...).
What is Qwen-Image-Edit?
Qwen-Image-Edit is a versatile image editing model built on the Qwen-Image architecture with 20 billion parameters. The model possesses dual editing capabilities, encompassing both semantic and visual aspects. It can perform low-level visual editing (such as adding, deleting, and modifying elements) and high-level visual semantic editing (such as IP creation, object rotation, and style transfer). The model supports precise editing of both Chinese and English text, allowing modification of text within images while preserving the original font, size, and style. Qwen-Image-Edit performs exceptionally well in multiple public benchmark tests, demonstrating state-of-the-art (SOTA) performance. It can be experienced through Qwen Chat.
Qwen-Image-Edit-2509 is the latest monthly iteration of Qwen-Image-Edit from the Qwen team. The model supports multiple image inputs, enabling various combined editing such as "person + person" and "person + scene," significantly improving the consistency of single-image editing, including editing people, products, and text. The model natively supports ControlNet, allowing flexible use of image conditions such as depth maps and edge maps, making it suitable for various creative scenarios such as creating emojis, restoring old photos, and generating cartoon dolls.
Main functions of Qwen-Image-Edit
- Semantic editingIt supports modifying image content while maintaining the visual semantic consistency of the original image.
- Appearance editingIt supports precise modification of local areas of an image, such as adding, deleting, or modifying elements in the image while keeping other areas unchanged.
- Precise text editingIt supports bilingual text editing in Chinese and English, allowing users to add, delete, and modify text in images while preserving the original font, size, and style.
- Strong benchmark performanceIt performs exceptionally well in multiple public benchmark tests, possessing state-of-the-art (SOTA) performance, and can efficiently complete a variety of complex image editing tasks.
The technical principles of Qwen-Image-Edit
- Model ArchitectureQwen-Image-Edit is further trained based on the Qwen-Image model with 20 billion parameters, inheriting its powerful text rendering and image generation capabilities. Input images are simultaneously fed into two modules: Qwen2.5-VL handles visual semantic control, understanding the semantic content of the image and performing semantic-level editing; VAE Encoder handles visual appearance control, precisely processing the visual details of the image and enabling editing of local areas.
- Semantic and Appearance EditingThe Qwen2.5-VL module enables the model to understand the overall semantics of an image and modify content while maintaining semantic consistency. The VAE Encoder module allows the model to precisely process visual details of an image, enabling the addition, deletion, or modification of local regions.
- Text editingQwen-Image-Edit optimizes text rendering, enabling accurate recognition and editing of text within images. The model supports both Chinese and English, allowing for the addition, deletion, and modification of text while preserving the original font, size, and style.
- Chain editingThe model supports chained editing, allowing for fine-tuning of complex image content through incremental adjustments. Users can specify the areas to be modified, and the model progressively optimizes those areas until the desired effect is achieved.
Qwen-Image-Edit project address
- Project official website: https://qwenlm.github.io/blog/qwen-image-edit/
- GitHub repositoryhttps://github.com/QwenLM/Qwen-Image
- HuggingFace model libraryhttps://huggingface.co/Qwen/Qwen-Image-Edit
- Experience the demo onlinehttps://huggingface.co/spaces/Qwen/Qwen-Image-Edit
Application Scenarios of Qwen-Image-Edit
- Creative DesignIt can quickly generate and modify the appearance, clothing, and background of virtual characters, efficiently completing the diverse creation of original IPs.
- Advertising and Poster DesignYou can directly modify the text content and adjust the font, size, and color in the poster without redesigning, thus improving design efficiency.
- Film and Video ProductionIn post-production, it can be used to quickly adjust scene elements or character appearances, or to change the style of a video from realistic to an anime style.
- Education and TrainingQuickly generate and modify teaching images and charts, such as portraits of historical figures and diagrams of scientific experiments, to enhance teaching effectiveness.
- Personal ApplicationsQuickly adjust your personal photos, such as changing the background, adding decorative elements, and modifying clothing, to easily create personalized photos.