AB
AiBoss
project

ViewCrafter - A high-fidelity new view compositing technology jointly proposed by Peking University, the Chinese University of Hong Kong, and Tencent.

ViewCrafter is an advanced video diffusion model jointly proposed by Peking University, the Chinese University of Hong Kong, and Tencent. It can synthesize high-fidelity new views from single or small numbers of images. It combines the generative capabilities of video diffusion models with point-based 3D representations...

What is ViewCrafter?

ViewCrafter, a state-of-the-art video diffusion model proposed by Peking University, the Chinese University of Hong Kong, and Tencent, can synthesize high-fidelity new views from single or small numbers of images. It combines the generative capabilities of video diffusion models with point-based 3D representations, precisely controlling camera pose to generate high-quality video frames. Through iterative view synthesis strategies and camera trajectory planning, ViewCrafter can progressively expand 3D cues to generate a wider range of new views. It demonstrates strong generalization ability and performance on multiple datasets, providing new possibilities for applications such as immersive real-time rendering experiences and scene-level text-to-3D generation.

ViewCrafte's main functions

  • New View CompositionSynthesize new views from single or small numbers of images to expand the user's perspective.
  • 3D scene reconstruction: Reconstruct the 3D structure of the scene to provide the geometric basis for generating new views.
  • Content creationIt supports text descriptions or other creative inputs to generate 3D scenes, enhancing the flexibility of content creation.
  • Real-time renderingOptimizes 3D scene representation, enables real-time rendering, and is suitable for virtual reality and augmented reality applications.
  • Dataset generalization: Validate model performance on multiple datasets to ensure generalization ability in different scenarios.

ViewCrafte's technical principles

  • Point cloud reconstructionBased on dense stereo vision algorithms, depth information is extracted from input images to construct a 3D point cloud model of the scene.
  • Video diffusion modelGenerative models, particularly diffusion models, in deep learning are used to generate new views. These views are then progressively recovered from noisy images to obtain clearer ones.
  • Iterative View Composition: Continuously optimize the generation of new views. Each iteration includes generating a new view and updating the point cloud model.
  • Camera trajectory planningIt automatically plans the camera's movement trajectory, captures the scene from different angles, and generates a more comprehensive view.
  • 3D scene understandingBy combining point clouds and generative models, we can understand the 3D structure of a scene and generate a new view that is consistent with the original scene.

ViewCrafte's project address

Application Scenarios of ViewCrafte

  • Film and television productionGenerate new perspectives in special effects shots to enhance the visual effects of scenes in post-production.
  • Game developmentVideo games create realistic game environments and backgrounds, providing a more immersive gaming experience.
  • Virtual Reality (VR)In virtual reality applications, ViewCrafter generates 360-degree panoramic images to enhance the user's immersion.
  • Augmented Reality (AR)It seamlessly integrates virtual objects into the real world, providing a richer interactive experience.
  • Architectural VisualizationIt helps designers showcase architectural models from different perspectives and provides a more intuitive design evaluation.