AB
AiBoss
project

3DTown - Columbia University, in collaboration with Cybever AI and others, launches a framework for generating 3D town scenes from a single view.

3DTown is a framework developed by Columbia University in collaboration with Cybever AI and other institutions to generate 3D town scenes from a single top-down view. The framework is based on region-based generation and spatially aware 3D inpainting techniques, decomposing the input image into overlapping regions...

What is 3DTown?

3DTown is a framework developed by Columbia University in collaboration with Cybever AI and other institutions to generate 3D town scenes from a single top-down view. Based on region-based generation and spatially aware 3D inpainting techniques, the framework decomposes the input image into overlapping regions. A pre-trained 3D object generator then generates 3D content for each region. A mask-based corrective flow inpainting process fills in missing geometry while maintaining structural continuity. 3DTown supports the generation of coherent 3D scenes with high geometric quality and texture fidelity, performing exceptionally well in scene generation across various styles, outperforming existing state-of-the-art methods.

3DTown's main functions

  • Generate diverse 3D scenesSupports the generation of scenes with different styles and layouts, such as "Snow Town" and "Desert Town".
  • Maintain geometric and texture consistencyThe generated 3D scene is highly consistent with the input image in terms of geometry and texture.
  • Efficiently handle complex scenariosIt can effectively handle complex scenes and avoid geometric distortion and layout illusion.

3DTown's technical principles

  • RegionalizationThe input image is decomposed into overlapping regions, each of which generates 3D content independently. A pre-trained 3D object generator is used to generate 3D content for each region, improving local alignment and resolution. Based on region fusion, the generated regions are progressively merged into a coherent global 3D scene.
  • Spatial Awareness 3D RestorationInitialize a coarse 3D structure as a spatial prior using monocular depth estimation and landmark detection. Fill in missing geometry using a Masked Rectified Flow technique while maintaining the continuity of known content. A two-stage Masked Rectified Flow pipeline generates sparse structures and structured latent representations, ensuring global consistency.
  • Structured latent representationThis approach constructs 3D scenes based on structured latent representations, including location indices and latent feature vectors. A sparse structure generator and a structured latent generator are used to progressively generate latent representations of the 3D scene.
  • Modular designBased on modular design, the complex problem of 3D scene generation is broken down into multiple sub-problems, each of which is solved independently before being integrated.

3DTown project address

3DTown application scenarios

  • Virtual World ConstructionIt can quickly generate virtual towns or scenes, providing realistic environments for virtual reality (VR) and augmented reality (AR) applications.
  • Game developmentIt provides game designers with efficient tools to generate complex 3D game scenes from simple top-down views, saving time and costs.
  • Robot SimulationTo create realistic 3D scenes for robot training and improve the robot's navigation and interaction capabilities in complex environments.
  • Digital content creationIt helps artists and designers quickly generate 3D scene prototypes, accelerating the creative process and improving work efficiency.
  • Architecture and Urban PlanningGenerate 3D architectural models and urban layouts from conceptual sketches to assist in planning and design work, and facilitate the presentation and evaluation of schemes.