UniWorld V2 - An image editing model jointly launched by RabbitShow Intelligence and Peking University
UniWorld V2 is a next-generation image editing model developed by RabbitShow Intelligence and the UniWorld team at Peking University. It employs the innovative UniWorld-R1 training framework and is the first to apply reinforcement learning policy optimization to image editing, through Diffu...
What is UniWorld V2?
UniWorld V2 is a next-generation image editing model developed by RabbitShow Intelligence and the UniWorld team at Peking University. It employs the innovative UniWorld-R1 training framework, applying reinforcement learning strategies to image editing for the first time, and achieving efficient training through DiffusionNFT technology. The model uses a multimodal large language model as the reward model, providing stable and fine-grained feedback, while introducing a low-variance group filtering mechanism to improve training stability. It can accurately understand and render complex Chinese fonts, supporting fine-grained spatial control, such as specifying the editing area through a frame, and achieving global lighting and shadow blending for more natural and harmonious images. It achieves leading results in industry benchmark tests such as GEdit-Bench and ImgEdit, comprehensively surpassing existing publicly available models.
Main functions of UniWorld V2
-
Precise rendering of Chinese fontsIt can understand and generate complex artistic Chinese fonts, such as "月满中秋" (Full Moon Mid-Autumn Festival), with clear effects and accurate semantics. Text modification can be achieved with just simple commands.
-
Refined space controlIt supports specifying the editing area through a frame, such as "moving the bird out of the red frame". The model can strictly adhere to the space constraints and complete highly difficult operations.
-
Global Lighting BlendA deep understanding of lighting instructions, such as "relighting the scene," allows objects to blend naturally into the scene, resulting in a high degree of light and shadow integration and a unified and harmonious image.
-
Instruction alignment and image quality improvementIt performs well in terms of instruction alignment and image quality, and users prefer its output results, especially in terms of instruction compliance.
-
Multi-model applicabilityThe framework is model-independent and can be applied to various basic models, such as Qwen-Image-Edit and FLUX-Kontext, significantly improving the performance of these models.
The technical principles of UniWorld V2
-
Innovative Training FrameworkUsing the UniWorld-R1 training framework, this study is the first to apply reinforcement learning policy optimization to image editing. It achieves policy optimization without likelihood estimation through Diffusion Negative-aware Finetuning (DiffusionNFT) technology, thereby improving training efficiency.
-
Multimodal reward modelUsing a multimodal large language model (MLLM) as the reward model, we can directly use the logarithmic value of its output to provide fine-grained feedback, avoiding the computational overhead and bias caused by complex reasoning and sampling.
-
Low variance group filtering mechanismTo address the issue of low-variance groups in reward normalization, a filtering strategy based on the reward mean and variance was designed to remove sample groups with high mean and low variance, thereby stabilizing the training process.
-
Model independenceThe framework is designed to be model-independent and can be applied to various basic image editing models, such as Qwen-Image-Edit and FLUX-Kontext, making it widely applicable.
UniWorld V2 project address
- Github repositoryhttps://github.com/PKU-YuanGroup/Uniworld
- arXiv technical paperhttps://arxiv.org/pdf/2510.16888
Application scenarios of UniWorld V2
-
Image Editing and DesignIt can precisely edit images according to user instructions, such as modifying text in the image, adjusting the position of objects, and changing the lighting and shadows of the scene. It is suitable for poster design, advertising creativity, visual arts and other fields.
-
Content creation and generationIt helps creators quickly generate image content that meets specific requirements, improving creative efficiency. It is suitable for scenarios that require a large number of image materials, such as video production, animation design, and game development.
-
Product Display and MarketingEnhance product presentation through image editing, such as adding special effects, adjusting backgrounds, and optimizing lighting and shadows, to increase product appeal. This is suitable for e-commerce product displays and brand promotion.
-
Education and TrainingAs a teaching tool, it helps students and learners better understand and master image editing skills. It can also be used to create educational image materials, such as textbook illustrations and teaching courseware.
-
Scientific research and experimentIn the field of scientific research, it can be used to generate simulated image data to assist in experimental design and result presentation, such as generating image samples under specific conditions in fields such as medical image processing and environmental science.