XVerse - A multi-subject controlled image generation model launched by ByteDance
XVerse is a new multi-agent controlled image generation model launched by ByteDance's intelligent creation team. The model achieves fine-grained control over the identities and semantic attributes (such as pose, style, and lighting) of multiple subjects in the text-to-image generation domain...
What is XVerse?
XVerse is a novel multi-subject controlled image generation model launched by ByteDance's Intelligent Creation Team. In the text-to-image generation domain, the model achieves fine-grained control over the identities and semantic attributes (such as pose, style, and lighting) of multiple subjects while maintaining high quality and consistency in the generated images. XVerse converts the reference image into a tag-specific text stream modulated offset, enabling precise and independent control over specific subjects without interfering with latent image variables or features. The model incorporates VAE-encoded image feature modules and regularization techniques to enhance detail preservation and generation quality. XVerse provides high fidelity and editability in multi-subject controlled image synthesis, offering powerful control over individual subject features and semantic attributes.
XVerse's main functions
- Multi-entity controlXVerse can control the identity and semantic attributes of multiple subjects simultaneously. For example, it can control the identity, posture, style, etc. of multiple people in an image at the same time, so as to achieve complex scene generation.
- High-fidelity image synthesisThe generated images have high fidelity, accurately reflecting the details and semantic information in the text description, while maintaining the overall quality and consistency of the images.
- Semantic attribute controlIt supports fine-grained control over semantic attributes (such as pose, style, and lighting), enabling flexible adjustments to image style and atmosphere.
- Powerful editabilityUsers can edit and adjust the generated images based on simple text prompts, enabling personalized image creation.
- Reduce artifacts and distortionBy introducing a VAE-encoded image feature module and regularization techniques, XVerse can significantly reduce artifacts and distortions in generated images, improving the naturalness and visual effect of the images.
XVerse's technical principles
- Text-stream Modulation MechanismThis method converts a reference image into a marker-specific text stream modulation offset, enabling precise control over a specific subject. The offset is added to the model's text embedding, allowing for fine-grained control over the generated image without interfering with latent variables or features of the image.
- VAE encoded image feature moduleTo enhance the detail preservation capabilities of generated images, XVerse introduces a VAE-encoded image feature module. This module acts as an auxiliary module, helping the model retain more detail and reduce artifacts and distortion during generation.
- Regularization techniquesBased on modulation injection that randomly retains one side, the model is forced to maintain consistency in non-modulation regions. Regularized subject-specific features are used as a data augmentation strategy for multi-subject datasets, improving the model's ability to distinguish and preserve subject features in multi-subject scenarios. By calculating the L2 loss of the text-image cross-attention map between the modulation model and the reference T2I branch, it is ensured that the modulation model retains the same attention pattern as the T2I branch, maintaining the consistency and editability of semantic interactions.
- Training dataXVerse is trained using a high-quality multi-agent control training dataset. The dataset utilizes Florence2 for image captioning and phrase localization, and SAM2 for accurate face extraction, constructing a high-quality training dataset encompassing various subjects and scenes. The training data covers a variety of scenarios, including human-object interactions, human-animal combinations, and complex multi-person scenes, enhancing the model's generalization ability.
XVerse's project address
- Project official websitehttps://bytedance.github.io/XVerse/
- GitHub repositoryhttps://github.com/bytedance/XVerse
- HuggingFace model libraryhttps://huggingface.co/ByteDance/XVerse
- arXiv technical paper: https://arxiv.org/pdf/2506.21416
Application scenarios of XVerse
- E-commerce ad generationIt can quickly generate advertising images of different people using the same product for e-commerce promotional activities, meeting the brand's personalized needs.
- Game character designGenerate multiple character concept art pieces with unique appearances and skills based on the game designer's descriptions, accelerating the character design process.
- Medical Education IllustrationsGenerate detailed anatomical and physiological diagrams of the human body to help medical students better understand human structure and function.
- Personal avatar customization on virtual social platformsUser input descriptions generate personalized virtual avatars, which can be used as avatars on virtual social platforms or as personal images in virtual reality.
- Urban planning scheme presentationGenerate virtual renderings of city parks to help citizens better understand the design plans of urban planners.