Step1X-3D - A 3D asset generation framework open-sourced by Step1X and LightIllusions.
Step1X-3D is a high-fidelity, controllable 3D asset generation framework developed by StepFun in collaboration with LightIllusions. Based on a rigorous data processing process, it selects 2 million high-quality assets from over 5 million 3D assets...
What is Step1X-3D?
Step1X-3D is a high-fidelity, controllable 3D asset generation framework developed by StepFun in collaboration with LightIllusions. Based on a rigorous data processing workflow, it selects 2 million high-quality data points from over 5 million 3D assets to create a standardized dataset of geometric and texture attributes. Step1X-3D supports multimodal conditional inputs, such as text and semantic labels, and achieves flexible geometric control based on low-rank adaptive (LoRA) fine-tuning. Step1X-3D is driving the development of 3D generation technology.
Main functions of Step1X-3D
- High-fidelity and controllable 3D asset generationGenerates 3D assets with high-fidelity geometry and diverse texture maps, maintaining excellent alignment between surface geometry and texture mapping.
- Supports multiple input conditionsIt supports various input conditions, such as multiple views, bounding boxes, and skeletons, enabling more flexible 3D asset generation.
- open sourceIt provides open-source technical reports, inference code, model weights, and training code.
Step1X-3D Technical Principles
- Data processingBased on multi-dimensional filtering conditions, high-quality 3D assets are accurately selected. By using the number of turns technique, the success rate of mesh to SDF conversion is improved, ensuring the accuracy of geometric supervision.
- Geometric generationBy leveraging perceptron-based latent encoding and sharp-edge sampling strategies, high-fidelity TSDF representations are generated. Efficient diffusion model training is performed based on a rectifier-flow converter, ensuring the stability and efficiency of geometry generation.
- Texture generationBased on a pre-trained multi-view image generation model, combined with geometric guidance, a consistent texture across multiple views is generated. A texture space synchronization module is introduced to achieve latent space alignment, ensuring precise alignment between texture and geometry. Texture inpainting techniques are used to process artifacts in UV mapping, achieving seamless texture synthesis.
- ControllabilityBased on LoRA fine-tuning technology, it achieves flexible geometric control, supports symmetry, geometric detail level control, and is compatible with multimodal condition inputs, enhancing the controllability and diversity of the generated data.
Step1X-3D project address
- GitHub repository:https://github.com/stepfun-ai/Step1X-3D
- HuggingFace model library:https://huggingface.co/stepfun-ai/Step1X-3D
- arXiv technical paper:https://arxiv.org/pdf/2505.07747
- Experience the demo online:https://huggingface.co/spaces/stepfun-ai/Step1X-3D
Application Scenarios of Step1X-3D
- Game developmentGenerate high-fidelity 3D models, quickly create prototypes, support personalized content, and enhance visual effects and player experience.
- Film and television productionUsed in the generation of virtual scenes, characters, and special effects to accelerate the production process and improve visual quality.
- Virtual Reality (VR) and Augmented Reality (AR)Create immersive 3D environments and interactive content to enhance the user experience.
- Architectural DesignGenerate virtual building and interior design models to assist in urban planning and enhance the design presentation effect.
- Education and training: Construct virtual laboratories, historical and cultural heritage models, and skills training environments to provide an intuitive and interactive learning experience.