AB
AiBoss
project

IMAGPose - Nanjing University of Science and Technology launches a unified framework for pose-guided image generation

IMAGPose is a unified conditional framework for human pose-guided image generation, developed by Nanjing University of Science and Technology. It addresses the limitations of traditional methods in pose-guided human image generation, such as the inability to simultaneously generate multiple different poses...

What is IMAGPose?

IMAGPose is a unified conditional framework for human pose-guided image generation developed by Nanjing University of Science and Technology. It addresses the limitations of traditional methods in pose-guided human image generation, such as the inability to simultaneously generate multiple target images with different poses, limitations in generating target images from multi-view source images, and the loss of detailed information in human images due to the use of frozen image encoders.

Main functions of IMAGPose

  • Multi-scenario adaptationIMAGPose supports various user scenarios, including generating target images from a single source image, generating target images from multi-view source images, and generating multiple target images with different poses simultaneously.
  • fusion of details and semanticsBy combining low-level texture features with high-level semantic features through the Feature-Level Conditional Module (FLC), the problem of loss of detail information caused by the lack of a dedicated feature extractor for human images is solved.
  • Flexible image and pose alignmentThe Image-Level Conditioning (ILC) module aligns images and poses by injecting a variable number of source image conditions and introducing a masking strategy, adapting to a variety of flexible user scenarios.
  • Global and local consistencyThe Cross-View Attention (CVA) module introduces a cross-attention mechanism that decomposes global and local information, ensuring local fidelity and global consistency of the person image when using multi-source image cues.

IMAGPose Technical Principles

  • Feature-level conditional blocks (FLC)The FLC module addresses the problem of lost detail information due to the lack of a dedicated feature extractor for human images by combining low-level texture features extracted by a variational autoencoder (VAE) encoder with high-level semantic features extracted by an image encoder.
  • Image-level conditional module (ILC)The ILC module achieves image and pose alignment by injecting a variable number of source image conditions and introducing a masking strategy, adapting to a variety of flexible user scenarios.
  • Cross-view attention module (CVA)The CVA module introduces a cross-attention mechanism that combines global and local decomposition to ensure local fidelity and global consistency of the person image when using multi-source image cues.

IMAGPose project address

Application scenarios of IMAGPose

  • Virtual Reality (VR) and Augmented Reality (AR)IMAGPose can generate images of people in specific poses, allowing them to present themselves in different poses in a virtual environment, or generate multiple poses for virtual characters, enhancing immersion.
  • Film production and special effectsIn film production, IMAGPose can be used to generate various poses for characters, helping special effects teams quickly generate character images in different scenes and reducing the time and cost of manual modeling and animation.
  • E-commerce and fashionIMAGPose can be used to generate clothing display images in different poses. Merchants can generate renderings of models wearing clothing in various poses, providing consumers with a more comprehensive visual experience.
  • Pedestrian Re-IDImages generated by IMAGPose can be used to improve the performance of pedestrian re-identification tasks. By generating images of people in different poses, the diversity of the dataset can be increased, improving the robustness and accuracy of the model.
  • Virtual Photography and Artistic CreationArtists and photographers can use IMAGPose to generate creative images of people in poses for virtual photography or artistic creation, exploring more visual possibilities.