Qwen-Image - A text-based graph model open-sourced by Alibaba's Tongyi Qianwen.
Qwen-Image is an open-source 20-parameter MMDiT model from the Alibaba Tongyi Qianwen team. It is the first fundamental image generation model in the Tongyi Qianwen series, excelling in complex text rendering and precise image editing, and supporting multi-line...
What is Qwen-Image?
Qwen-Image is an open-source 20-parameter MMDit model from Alibaba's Tongyi Qianwen team. It's the first foundational image generation model in the Tongyi Qianwen series, excelling in complex text rendering and precise image editing. It supports multi-line layouts, paragraph-level text generation, and fine-grained detail rendering, achieving high-fidelity output in both Chinese and English. Qwen-Image demonstrates powerful capabilities in general image generation and editing tasks, supporting various artistic styles and advanced editing operations. Users can currently experience the model's performance through Qwen Chat's image generation function.
Qwen-Image's main functions
- Complex text renderingIt supports the generation of multi-line and paragraph text, can clearly present small text, and excels at rendering Chinese and English text.
- Precise Image EditingIt supports style transfer, object addition, deletion and modification, detail enhancement, text editing and character pose adjustment, while maintaining the naturalness and realism of the image.
- General Image GenerationIt supports multiple art styles and can generate creative images based on user descriptions.
Qwen-Image's technical principles
- Model ArchitectureBased on an advanced Multimodal Large Language Model (MLLM) as the text feature extraction module, it can accurately understand text semantics and transform them into features required for image generation. A Variational Autoencoder (VAE) is responsible for encoding the input image into a compact latent representation, which is then decoded during the inference phase, achieving efficient image processing and generation. The core of the model is the Multimodal Diffusion Transformer (MMDiT), which generates images by progressively removing noise and is guided by text features to ensure that the generated image is highly consistent with the text description.
- Data processingThrough large-scale data collection and annotation, a rich dataset covering natural, design, people, and synthetic data is constructed. Based on a multi-stage data filtering process, low-quality or unacceptable data is gradually removed to ensure high-quality and diverse data.
- Training strategyDuring training, flow matching is used as the pre-training objective, and ordinary differential equations (ODE) are used to achieve stable training dynamics while maintaining equivalence with the maximum likelihood objective. The model combines multi-task training paradigms of text-to-image (T2I), image-to-image (I2I), and text-to-image-to-image (TI2I), and achieves multi-task learning based on a shared latent space.
Qwen-Image's performance
- Overall performance:
- Leading in multiple benchmark testsQwen-Image achieved 12 state-of-the-art (SOTA) scores in multiple public benchmark tests, demonstrating its strong competitiveness in the field of image generation and editing.
- Beyond the head modelIn general image generation tests (such as GenEval, DPG, and OneIG-Bench) and image editing tests (such as GEdit, ImgEdit, and GSO), Qwen-Image outperforms open-source models such as Flux.1 and BAGEL, and also surpasses closed-source models such as ByteDance's SeedDream 3.0 and OpenAI's GPT Image 1 (High). Qwen-Image achieves a high level in both generation quality and editing capabilities.
- Text rendering capabilities:
- Text rendering benchmarkIn benchmark tests such as LongText-Bench, ChineseWord, and TextCraft, Qwen-Image performs exceptionally well, especially in Chinese text rendering, where it significantly outperforms existing state-of-the-art models such as SeedDream 3.0 and GPT Image 1 (High).
- Advantages of Chinese text renderingQwen-Image has unique advantages in processing Chinese text rendering, with optimized technologies in language understanding, font generation, and typesetting, making it better suited to the complexity and diversity of Chinese.
How to use Qwen-Image
- Visit QwenChatVisit the official Qwen Chat website.
- Select image generation functionIn the QwenChat interface, find and select the "Image Generation" function.
- Input text promptEnter a description of the image you want to generate in the text input box.
- Generate imageClick the "Generate" button, and Qwen-Image will generate an image based on the text prompts.
- View and download the generated imagesThe generated image is displayed on the interface, allowing users to view the effect and choose to download and save it locally.
Qwen-Image's project address
- GitHub repositoryhttps://github.com/QwenLM/Qwen-Image
- HuggingFace model libraryhttps://huggingface.co/Qwen/Qwen-Image
- Technical Papers: https://qianwen-res.oss-cn-beijing.aliyuncs.com/Qwen-Image/Qwen_Image.pdf
- Experience the demo onlinehttps://huggingface.co/spaces/Qwen/Qwen-Image
Application Scenarios of Qwen-Image
- Content creationIt can quickly generate high-quality images, posters, and PPT pages based on text descriptions, greatly improving the efficiency and visual effects of creative design and presentation production.
- Art and DesignThe model can easily achieve style transfer and creative painting, providing artists and designers with a wealth of inspiration and accelerating the creation process of artworks.
- Education and LearningBy generating teaching materials and language learning-related images, it helps teachers impart knowledge more vividly and assists learners in better understanding and memorization.
- Business and MarketingIn the business field, it can quickly generate attractive advertising images and brand promotion materials, effectively enhancing the appeal of advertisements and the market influence of brands.
- Entertainment and GamesUsed to generate images of characters, scenes, and props in games, as well as special effects and concept art in film and television production, accelerating the creation cycle of entertainment content.