FlexIP - A personalized image generation and editing framework launched by Tencent.
FlexIP is a flexible subject attribute editing framework for image compositing proposed by Tencent, balancing identity preservation and personalized editing in image generation. The framework employs a dual-adaptor architecture to decouple identity preservation from personalized editing, through...
What is FlexIP?
FlexIP, proposed by Tencent, is a flexible subject attribute editing framework for image compositing, balancing identity preservation and personalized editing in image generation. The framework employs a dual-adaptor architecture, decoupling identity preservation from personalized editing, and ensuring identity integrity through high-level semantic concepts and low-level spatial details. A dynamic weight gating mechanism allows users to flexibly parameterize the balance between identity preservation and style personalization, transforming the traditional binary trade-off into a continuous control surface. FlexIP incorporates a multimodal data training strategy, optimizing the adapter's identity locking and deformation capabilities based on image and video data respectively, further enhancing generation robustness.
Main functions of FlexIP
- Dual adapter decoupling designFor the first time, the Preservation Adapter and Personalization Adapter are explicitly separated. The Preservation Adapter combines high-level semantic concepts with low-level spatial details to ensure identity integrity; the Personalization Adapter interacts with textual and visual CLS tokens, incorporating meaningful visual cues, placing text modifications within a coherent visual context, avoiding feature competition, and achieving more precise control.
- Dynamic weight gating mechanismBy dynamically balancing identity preservation and editing intensity through continuously adjustable parameters, the traditional binary trade-off is transformed into a continuous parameter control surface, supporting flexible control from fine adjustments to large deformations. Users can flexibly adjust the generated effect as needed.
- Modality-aware training strategyThe adapter weights are adaptively adjusted based on data characteristics (static images/video frames). Image data strengthens identity locking, and video data optimizes temporal deformation, thereby improving generation robustness.
- Cross-attention mechanismThe adapter enhances identity robustness by capturing multi-granular visual features (such as facial details) across attention.
- Dynamic interpolationThe weighted gating mechanism allows users to adjust the adapter contribution in real time, forming a continuous "control surface".
- Multimodal data trainingBy combining image and video data, the adapter's identity locking and transformation capabilities are optimized respectively.
FlexIP performance comparison
- Quantitative comparison
- Overall RankingOn the overall ranking (mRank) metric, FlexIP outperformed all other methods, indicating that it performed best across multiple key metrics.
- Personalization capabilitiesIn the personalization assessment, FlexIP scored 0.284 on CLIP-T, slightly lower than λ-Eclipse, but λ-Eclipse achieves this at the cost of sacrificing subject retention. FlexIP, on the other hand, achieves a high level of personalization while maintaining subject characteristics.
- Ability to maintain identityIn terms of identity preservation, FlexIP achieved high scores of 0.873 and 0.739 on CLIP-I and DINO-I respectively, significantly outperforming other methods and demonstrating its strong advantage in preserving image details and semantic consistency.
- Image qualityIn image quality assessment, FlexIP scored 0.598 on CLIP-IQA and 6.039 on aesthetics, indicating that the images it generates are not only of high quality but also have better aesthetic appeal.
- User ResearchIn practical user satisfaction evaluations, FlexIP performed exceptionally well in both flexibility (Flex) and identity preservation (ID-Pres). All 60 evaluators agreed that the images generated by FlexIP best matched the semantics of the text and best preserved the subject's features.
- Qualitative comparison
- FidelityImages generated by FlexIP exhibit excellent fidelity, faithfully reproducing the main features and details of the reference image, maintaining high quality and realism even during personalized editing.
- EditabilityFlexIP has a significant advantage in editability, capable of generating diverse editing results based on different text commands, meeting users' personalized needs in different scenarios.
- Identity ConsistencyIn terms of identity consistency, FlexIP can stably maintain subject features across different reference images, ensuring subject identity consistency even during significant deformation or stylistic editing, thus avoiding the identity mutation problem common in traditional methods.
- Comparison with existing methodsWhen qualitatively compared with five state-of-the-art methods, FlexIP generates images with significant improvements in fidelity, editability, and identity consistency, better meeting users' needs for personalized generation of high-fidelity images.
FlexIP project address
- Project official website:http://flexip-tech.github.io/flexip/#/
- arXiv technical paper:https://arxiv.org/pdf/2504.07405
Application scenarios of FlexIP
- Artistic CreationFlexIP allows artists to flexibly personalize images according to their needs while preserving the subject's identity.
- Advertising designIn the field of advertising design, FlexIP helps designers quickly generate image content that meets brand needs. Through a dynamic weight gating mechanism, designers can flexibly adjust the style, scene, and details of advertising images while maintaining the brand image.
- Film and television productionFlexIP can be used for visual effects and character design in film and television production. It allows for flexible adjustments to a character's appearance while maintaining consistency in the character's identity.
- Game developmentIn game development, FlexIP can be used for generating and editing characters and scenes. Developers can use this framework to quickly generate diverse character designs while maintaining the core characteristics of the characters.