Insert Anything - An image insertion framework jointly developed by Zhejiang University, Harvard University, and Nanyang Technological University.
Insert Anything is a context-editing-based image insertion framework jointly developed by researchers from Zhejiang University, Harvard University, and Nanyang Technological University. The framework seamlessly inserts objects from a reference image into a target scene...
What is Insert Anything?
Insert Anything is a context-editing-based image insertion framework jointly developed by researchers from Zhejiang University, Harvard University, and Nanyang Technological University. The framework seamlessly inserts objects from reference images into a target scene, supporting various practical applications such as artistic creation, real-life face replacement, film scene compositing, virtual try-on, accessory customization, and digital prop replacement. Trained on the AnyInsertion dataset containing 120K cue image pairs, Insert Anything can flexibly adapt to various insertion scenarios, providing powerful technical support for creative content generation and virtual try-on.
The main function of Insert Anything
- Multi-scenario supportIt supports handling various image insertion tasks, such as people insertion, object insertion, and clothing insertion.
- Flexible user controlSupports both mask-guided and text-guided control modes. Users can specify the insertion area and content based on manually drawn masks or entered text descriptions.
- High-quality outputSupports the generation of high-quality, high-resolution images while maintaining the detail and style consistency of inserted elements.
The technical principle of Insert Anything
- AnyInsertion datasetThe framework is trained on the large-scale dataset AnyInsertion, which contains 120K cue-image pairs and covers a variety of insertion tasks (such as people, objects and clothing insertion).
- Diffusion converter (DiT)This technology utilizes a multimodal attention mechanism based on DiT to simultaneously process text and image inputs. DiT can jointly model the relationships between text, masks, and image patches, supporting flexible editing control.
- Context editing mechanismBased on the polyptych format (such as mask-guided diptych and text-guided triptych), the reference image is combined with the target scene, allowing the model to capture contextual information and achieve a natural insertion effect.
- Semantic guidanceCombining image encoders (such as CLIP) and text encoders to extract semantic information provides advanced guidance for the editing process, ensuring that inserted elements are consistent with the style and semantics of the target scene.
- Adaptive pruning strategyWhen dealing with small targets, the cropping area is dynamically adjusted to ensure that the editing area receives sufficient attention, retains enough contextual information, and achieves high-quality detail preservation.
Insert Anything's project address
- Project official website:https://song-wensong.github.io/insert-anything/
- GitHub repository:https://github.com/song-wensong/insert-anything
- arXiv technical paper:https://arxiv.org/pdf/2504.15009
Application scenarios for Insert Anything
- Artistic CreationQuickly combine different elements to inspire creative ideas.
- Virtual try-onAllowing consumers to preview the appearance of clothing enhances the shopping experience.
- Film and television special effectsSeamlessly insert virtual elements to reduce shooting costs.
- Advertising designQuickly generate a variety of creative ads to enhance their appeal.
- Cultural Heritage RestorationVirtual restoration of cultural relics or architectural details aids in research and exhibition.