News
ByteDance releases visual spatial reconstruction model: Depth Anything 3
ByteDance's Seed team has open-sourced Depth Anything 3, a visual spatial reconstruction model that innovatively employs a single Transformer architecture to achieve spatial perception from any viewpoint. The model integrates tasks such as camera pose estimation and geometric reconstruction into a concise framework using a unified "depth-ray" representation method, achieving improvements of 35.7% and 23.6% respectively in camera pose accuracy and geometric reconstruction compared to the mainstream model VGGT.