AB
AiBoss
project

Skywork UniPic 2.0 - Kunlun Wanwei's open-source unified multimodal model

Skywork UniPic 2.0 is a high-efficiency multimodal model open-sourced by Kunlun Wanwei, focusing on unified image generation, editing, and understanding capabilities. The model is based on the SD3.5-Medium architecture with 2B parameters, and utilizes pre-training and progressive dual-task training...

What is Skywork UniPic 2.0?

Skywork UniPic 2.0 is a high-efficiency multimodal model open-sourced by Kunlun Wanwei, focusing on unified image generation, editing, and understanding capabilities. Based on the 2B parameter SD3.5-Medium architecture, the model achieves synergistic optimization of generation and editing tasks through pre-training, progressive dual-task reinforcement strategies, and joint training, outperforming many high-parameter models. The model supports text-to-image generation, image editing, and multimodal understanding, featuring lightweight efficiency and flexible switching, helping developers quickly build multimodal applications.

Main features of Skywork UniPic 2.0

  • Image generationGenerates high-quality images based on user-input text descriptions, supporting various styles and scenes.
  • Image editingIt allows users to modify content and convert styles of existing images to meet diverse editing needs.
  • Multimodal understandingIt can understand image content and answer related questions, and supports the execution of complex commands and content modification.

The technical principles of Skywork UniPic 2.0

  • Architecture DesignBased on the SD3.5-Medium architecture with 2B parameters, it supports text-to-image generation and image editing tasks. By freezing the raw image editing module and combining it with multimodal models (such as Qwen2.5-VL-7B) and connectors, it constructs an integrated model for understanding, generation, and editing.
  • Pre-trainingThe model is pre-trained on large-scale, high-quality image generation and editing datasets, enabling it to possess basic generation and editing capabilities. Based on a text encoder and a VAE encoder, text and images are used as conditional inputs to enhance the model's multimodal understanding capabilities.
  • reinforcement learningBased on the Flow-GRPO framework, a progressive dual-task enhancement strategy is designed to optimize the generation and editing tasks separately, avoid mutual interference between tasks, and improve the overall performance of the model.
  • Joint trainingThe multimodal model is aligned with the raw image editing module via a connector and pre-trained. Based on the connector pre-training, the connector and raw image editing module are jointly trained to further improve the model's performance.

Skywork UniPic 2.0 project address

  • Project official websitehttps://unipic-v2.github.io/
  • GitHub repository: https://github.com/SkyworkAI/UniPic/tree/main/UniPic-2
  • HuggingFace model library: https://huggingface.co/collections/Skywork/skywork-unipic2-6899b9e1b038b24674d996fd
  • Technical Papers: https://github.com/SkyworkAI/UniPic/blob/main/UniPic-2/assets/pdf/UNIPIC2.pdf

Application scenarios of Skywork UniPic 2.0

  • Creative DesignQuickly generate ads, posters, or illustrations to help designers realize their creative ideas.
  • Content creationGenerate keyframes, characters, or scenes for video, animation, or game development, accelerating the creation process.
  • EducationGenerate relevant images or animations based on the teaching content to assist teaching and enhance students' learning interest.
  • EntertainmentGenerate personalized social media images or virtual reality scenes to enhance the user experience.
  • Business applicationsGenerate product concept images, packaging designs, or marketing promotional images to help commercial projects move forward quickly.