GenXD - A general-purpose 3D and 4D joint generative framework jointly developed by the National University of Singapore and Microsoft.
GenXD is a 3D-4D co-generation framework jointly developed by the National University of Singapore and Microsoft. It can generate high-quality 3D and 4D scenes from any number of conditional images. The framework uses a data processing workflow to extract camera data from video...
What is GenXD?
GenXD is a joint 3D-4D generation framework jointly developed by the National University of Singapore and Microsoft. It can generate high-quality 3D and 4D scenes from any number of conditional images. The framework uses a data processing workflow to extract camera pose and object motion intensity from videos, and trains a model based on this information and the large-scale 4D dataset CamVid-30K. GenXD decouples camera and object motion based on a multi-view temporal module, and supports conditional generation from multiple perspectives using masked latent conditions, enabling the handling of multiple 3D and 4D generation tasks in a single model.
Main functions of GenXD
- 3D and 4D scene generationGenXD can generate high-quality 3D and 4D scenes from single or multiple views, including dynamic and static content.
- Camera pose estimationBased on the Structure from Motion (SfM) technique, GenXD estimates the camera pose in the video, providing a foundation for generating videos that are consistent with the camera trajectory.
- Object motion estimationBased on depth estimation and keypoint tracking, GenXD identifies and simulates the motion of objects in videos.
- Multi-view timing moduleThe internal modules of the framework process multi-view and temporal information, decouple camera motion and object motion, and generate more realistic dynamic scenes.
- Masking latent conditionsGenXD supports condition generation using masked latent conditions, allowing models to accept any number of input views without changing the network structure.
GenXD's technical principles
- Data processing processGenXD uses a data processing workflow to extract camera pose and object motion information from videos, providing the necessary data for subsequent model training.
- Multi-view timing moduleGenXD's internal multi-view temporal module can process multi-view and temporal information, and use the α fusion strategy to perform seamless learning in 3D and 4D data.
- Masked latent conditional diffusion modelGenXD uses the masked latent conditional diffusion model (LDM) to generate images with different camera viewpoints and time steps, supporting single-view and multi-view generation.
- Decoupling camera and object motionBased on the multi-view timing module, GenXD separates camera motion and object motion, which is crucial for generating dynamic scenes.
- 3D and 4D data fusionGenXD combines 3D and 4D data during training, allowing the model to learn spatial and temporal information simultaneously, thereby improving the quality of the generated data.
- 3D representation optimizationImages generated by GenXD can be directly used to optimize 3D representations, such as 3D Gaussian point clouds (3D-GS) and Zip-NeRF, to achieve high-quality 3D scene reconstruction.
GenXD's project address
- Project official website:gen-x-d.github.io
- GitHub repository:https://github.com/HeliosZhao/GenXD
- arXiv technical paper:https://arxiv.org/pdf/2411.02319
Application scenarios of GenXD
- Video game developmentGenXD is used to generate 3D and 4D environments in games, providing a more realistic and dynamic game world.
- Film and visual effectsIn film production, GenXD creates complex 3D scenes and special effects, reducing the cost of actual shooting and post-production.
- Virtual Reality (VR) and Augmented Reality (AR)GenXD generates immersive 3D and 4D content, enhancing the user experience of VR and AR applications.
- Architecture and Urban Planning:Based on 3D models generated by GenXD, architects and urban planners can more intuitively showcase design concepts and planning schemes.
- Education and trainingGenXD creates simulation environments for use in education and professional training, such as simulated surgery and historical reenactment.