USO - ByteDance's unified framework for decoupling and reorganizing content and style
USO (Unified Style-Subject Optimized) is a unified framework for decoupling and recombining content and style, developed by ByteDance's UXO team. It allows for the free combination of any theme and style in any scenario to generate content with high...
What is USO?
USO (Unified Style-Subject Optimized) is a unified framework for content and style decoupling and recombining, developed by ByteDance's UXO team. It allows for the free combination of any theme and style in any scene to generate images with high subject consistency, strong style fidelity, and a natural, non-artificial look. USO utilizes a large-scale triplet dataset and a decoupling learning scheme to simultaneously align style features and separate content and style, further enhancing model performance by introducing Style Reward Learning (SRL). USO has released the USO-Bench benchmark test for comprehensively evaluating style similarity and subject fidelity. Experiments show that USO achieves state-of-the-art performance among open-source models in both subject consistency and style similarity.
Main functions of USO
-
Style and subject integrationIt can freely combine any theme with any style to generate images that retain the characteristics of the subject while conforming to the specified style, thus solving the problem of the difficulty in integrating style and subject.
-
High-fidelity generationWhen generating images, it can maintain a high degree of subject consistency and style fidelity, ensuring that the generated images are natural and of high quality.
-
Multi-scenario applicationsSuitable for a variety of scenarios, it can be widely used in fields such as artistic creation, advertising design, and game development.
-
Open source supportThe project is fully open source, including training code, inference scripts, model weights, and datasets, providing a wealth of resources for researchers and developers.
-
Leading performanceIt achieves top-tier performance in both subject consistency and style similarity dimensions compared to open-source models, with performance improvements achieved through a large-scale triplet dataset and a decoupled learning scheme.
-
BenchmarkingUSO-Bench benchmark test was released to comprehensively evaluate style similarity and subject fidelity, providing a unified standard for comparison of subsequent models.
USO's technical principles
-
Construction of large-scale triplet datasetsA triplet dataset containing content images, style images, and corresponding stylized images was created, providing a rich data foundation for model training.
-
Decoupling learning schemeBy employing two phases—style alignment training and content-style decoupling training—style features are aligned simultaneously while content and style are separated, avoiding feature crosstalk and achieving precise fusion.
-
Style Reward Learning (SRL)Introducing reward signals optimizes generation quality, balances style similarity and subject consistency, and further improves model performance.
-
Unified frameworkThis approach merges style-driven and subject-driven tasks into a single model framework, resolving the conflict between the two in traditional methods and achieving collaborative optimization of style and subject.
-
Two-stage training processThe first stage enables the model to reproduce styles through style alignment training; the second stage achieves joint conditional generation through content-style decoupling training; and finally, the entire training process is supervised through style reward learning.
USO's core value
- An innovative collaborative decoupling paradigm was proposed.This breaks the situation where style and subject generation tasks operate independently, proving that cross-task joint learning can achieve more thorough content-style decoupling and mutual promotion.
- A powerful unified generative model was built.USO is the first model to achieve state-of-the-art subject consistency and style similarity simultaneously within a single framework, and its effectiveness and versatility are impressive.
- Introduced reward-based learning reinforcementThe successful application of the reward-based learning paradigm to style generation provides an effective way to further improve the fine control and aesthetic quality of generative models.
- The first joint assessment benchmark was released.USO-Bench fills the gap in comprehensive evaluation in this field, providing a fair and comprehensive comparison platform for subsequent research.
USO's project address
- Project official websitehttps://bytedance.github.io/USO/
- Github repositoryhttps://github.com/bytedance/USO
- arXiv technical paper: https://arxiv.org/pdf/2508.18966
USO's model effect
-
Accurate style transferIt can accurately transfer different styles to new content, and the generated images retain the brushstrokes and colors of the original style without distorting the subject, resulting in a high degree of style similarity.
-
Main features preservedWhen styles change, it can lock in the main features, adapt to multiple styles, maintain the original appearance of people or objects, and maintain good consistency of the main subject.
-
Strong joint generation capabilityIt can simultaneously meet the dual requirements of style and subject, generating images that conform to the specified style while fully preserving the subject layout in one step, achieving a perfect fusion of style and subject.
-
High qualityIt achieved state-of-the-art (SOTA) results in subject-driven generation, style-driven generation, and joint style-subject-driven generation tasks, producing natural, realistic, and high-quality images.
-
Highly adaptableThe model is highly adaptable to different subjects and styles, and can handle a variety of content types, such as people, animals, and scenes, as well as various styles, such as oil painting, ink painting, and comics.
-
Quantitative comparisonOn USO-Bench, in both subject-driven and style-driven tasks, USO's various metrics (such as CLIP-I, DINO, CSD) all...Significantly betterAll existing open-source state-of-the-art (SOTA) models. In more challenging...Style-Subject Joint DrivingIn terms of tasks, USO also has a significant lead, demonstrating its powerful unified generation capabilities.
Application scenarios of USO
-
Artistic CreationArtists can use USO to apply different art styles to the same subject, quickly generating sketches or finished products in various styles, inspiring creative ideas and improving creative efficiency.
-
Advertising designAdvertising designers can use USO to quickly generate advertising images with specific styles and thematic characteristics based on different advertising themes and target audiences, thereby enhancing the attractiveness and relevance of their advertisements..
-
Game developmentGame developers can use USO to generate images of different styles for game characters and scenes, enriching the visual effects and enhancing the immersive experience. For example, it can transform the appearance of a game character from realistic to cartoonish.
-
Film and television productionIn film and television special effects production, USO can be used to quickly generate scenes or character designs with a specific style, assisting special effects artists in creative conception and effect preview. For example, it can be used to generate futuristic character designs for a science fiction film.
-
EducationIn art and design education, USO (Unique Artwork Objects) can serve as a teaching tool to help students better understand and master the characteristics of different art styles and how to apply these styles to their actual creations. For example, teachers can use USO to demonstrate how the same artwork is presented in different styles.