Step 3 - The latest multimodal inference model from Leap Star
Step 3 is the latest generation of fundamental large-scale model released by StepStar, designed specifically for the inference era, combining high performance with extreme cost-effectiveness. Utilizing the MoE architecture, it boasts 321B total parameters and 38B activation parameters, making it the first...
What is Step 3?
Step 3 is StepStar's latest generation of foundational large-scale model, designed specifically for the inference era, combining high performance with extreme cost-effectiveness. Utilizing the MoE architecture, it boasts 321B total parameters and 38B activation parameters, making it the first full-size, native multimodal inference model. It possesses powerful visual perception and complex reasoning capabilities, enabling efficient applications across multiple domains. Through the AFD distributed inference system and MFA attention mechanism, it achieves a significant improvement in inference efficiency. On domestically produced chips, its inference efficiency is up to 3 times that of similar models, and on NVIDIA Hopper architecture chips, its throughput is increased by over 70%, significantly reducing inference costs. Step 3 will be officially open-sourced on July 31st, providing the most powerful multimodal inference model for developers and enterprises worldwide.
The main functions of Step 3
-
Visual perceptionStep 3 can accurately identify and analyze complex information in images and videos, such as accurately reconstructing content in menu recognition with severe glare.
-
Complex ReasoningIt supports complex knowledge understanding across disciplines and cross-analysis of mathematical and visual information, such as automatically calculating the cost-sharing of split expenses by combining WeChat group chat records and shopping receipts.
-
Multimodal task processingAs a native multimodal model, Step 3 can handle tasks of multiple modalities such as language and vision, meeting the needs of diverse application scenarios.
-
Efficient ReasoningThrough system architecture innovation, Step 3 demonstrates outstanding inference efficiency. On domestically produced chips, its inference efficiency can reach up to that of DeepSeek-R1. 300%Throughput improvement exceeding [percentage missing] on NVIDIA Hopper architecture chips 70%.
-
Hardware friendlyStep 3: Adapts to multiple hardware platforms, including mainstream and domestic chips, which can significantly reduce inference costs and improve resource utilization.
Technical principles of Step 3
- MoE architectureStep 3 employs the MoE (Mixture of Experts) architecture, a highly efficient method for model parallelization. By decomposing the model into multiple "expert" modules and dynamically selecting the appropriate expert for computation based on the input, the MoE architecture can significantly reduce the waste of computing resources while maintaining high performance.
- AFD Distributed Inference SystemDistribute the attention and feedforward network (FFN) computation tasks in the model to the most suitable hardware to improve overall efficiency.
-
Attention calculationTasks that consume a great deal of memory bandwidth should be assigned to GPU clusters with high memory bandwidth.
-
FFN calculationTasks that consume extremely high computing power are assigned to GPU clusters with powerful computing capabilities.
-
-
MFA attention mechanism: Optimize arithmetic strength, adapt to the performance characteristics of mainstream and domestic chips, and achieve efficient inference across hardware platforms.
Step 3 Project Address
- Github repositoryhttps://github.com/stepfun-ai/Step3
Application scenarios of Step 3
- Smart Terminal AgentStep 3 can be applied to various IoT devices, such as smart home devices and smart wearable devices, providing intelligent voice assistants and visual recognition functions.
-
Finance and EconomicsStep 3 can be used in scenarios such as financial risk assessment, intelligent customer service, and market analysis. Through multimodal data processing, the model can more accurately analyze market trends and user needs.
-
Content creationStep 3 can assist content creators in generating creative copy, images, and video content. For example, it can combine visual and textual information to generate high-quality advertising copy or video scripts.
-
Visual recognitionStep 3 can handle complex visual tasks, such as reflective menu recognition, image classification, and object detection.
-
Complex ReasoningStep 3 supports understanding complex knowledge across different fields, such as automatically calculating the cost-sharing of split expenses by combining WeChat group chat records and shopping receipts.
-
Natural Language ProcessingStep 3: Performs well in natural language processing tasks, able to understand and generate high-quality text content.