AB
AiBoss
project

See3D - An open-source, label-free video-based 3D generative model from the Beijing Academy of Artificial Intelligence.

See3D (See Video, Get 3D) is a 3D generative model developed by the Beijing Academy of Artificial Intelligence. It learns from large-scale unlabeled internet videos to generate 3D content from them. This differs from traditional camera-based models...

What is See3D?

See3D (See Video, Get 3D) is a 3D generative model developed by the Beijing Academy of Artificial Intelligence (BAAI). It learns from large-scale unlabeled internet videos to generate 3D content. Unlike traditional 3D generative models that rely on camera parameters, See3D uses visual conditional techniques to generate multi-view images with controllable camera orientation and consistent geometry, solely based on visual cues in the video. This avoids the need for expensive 3D or camera annotation and efficiently learns 3D priors from internet videos. See3D supports the generation of 3D content from text, single-view, and sparse-view images, and can perform 3D editing and Gaussian rendering.

See3D's main functions

  • Generation from text, single view, and sparse view to 3DSee3D can generate 3D content based on text descriptions, images from a single perspective, or a small number of images.
  • 3D Editing and Gaussian RenderingThe model supports editing of the generated 3D content and uses Gaussian rendering technology to improve rendering effects.
  • Unlock the 3D Interactive WorldAfter inputting an image, an immersive and interactive 3D scene can be generated, allowing users to explore the real spatial structure in real time.
  • 3D Reconstruction Based on Sparse ImagesInput a small number of images (3-6), and the model can generate a detailed 3D scene.
  • Open World 3D GenerationBased on the text prompts, the model can generate artistic images, and then use these images to create virtual 3D scenes.
  • 3D generation based on a single viewInput a real-world image, and the model can generate a realistic 3D scene.

See3D's technical principles

  • Visual Conditioning TechnologySee3D does not rely on traditional camera parameters. Instead, it uses visual condition technology to generate multi-view images with controllable camera orientation and geometric consistency through visual cues in the video.
  • Large-scale unlabeled video learningSee3D can efficiently learn 3D priors from internet videos without relying on expensive 3D or camera annotations.
  • Dataset ConstructionThe team has built a high-quality, diverse, large-scale, multi-view image dataset, WebVi3D, which includes 320 million frames of images from 16 million video clips. The dataset can be continuously expanded through automated processes as the amount of internet video grows.
  • Multi-view diffusion model trainingSee3D introduces a new visual condition by adding time-dependent noise to masked video data to generate a pure 2D inductive visual signal. It supports scalable multi-view diffusion model (MVD) training, avoids dependence on camera conditions, and achieves the goal of "obtaining 3D through vision alone".
  • 3D Generative FrameworkSee3D's learned 3D priors enable a range of 3D creation applications, including single-view-based 3D generation, sparse view reconstruction, and 3D editing in open-world scenes, supporting the generation of long sequences of views under complex camera trajectories at both the object and scene levels.

See3D's project address

Application Scenarios of See3D

  • Game developmentAI-generated 3D models can be used to create characters, environments, and objects in games, improving development efficiency and reducing costs.
  • Architectural DesignIn architectural design, AI can generate building models, helping designers quickly conceive and modify design schemes.
  • e-commerceOnline retailers can use AI-generated 3D models to showcase products, enhancing the user shopping experience.
  • AR/VRIn the AR/VR field, AI-generated 3D models can be used to create realistic virtual environments and characters, enhancing the user's sense of immersion.
  • Movies and EntertainmentAI can help filmmakers create CG characters by replacing real-life actors, simplifying the special effects production process.
  • Industrial DesignAI-generated 3D models can be used to simulate the design of industrial products, accelerating the product development process.