Step Edge - A complete set of end-side models launched by Step Edge.
Step Edge is a complete suite of edge models launched by StepStar, including four major components: basic models, audio, GUI, and Gen, designed for mobile phones, automobiles, and other terminals.
What is Step Edge?
Step Edge is a complete suite of edge models launched by StepStar, comprising four main components: basic models, audio, GUI, and Gen, designed for mobile devices and automobiles. Through an edge-cloud collaborative architecture, the model enables the Agent to achieve a 0.1-second response time locally while ensuring full-modal data privacy. It achieved first place in 29 core evaluations and is equipped with a self-developed Step Inference NPU engine to optimize terminal inference.
Step Edge's main functions
-
Step Edge Basic ModelIt supports edge-side multimodal understanding and reasoning, including text, images, videos, OCR, spatial understanding, and tool invocation.
-
Step Edge AudioOn-device audio understanding and speech recognition, covering ASR and audio question answering in multiple languages including Chinese and English.
-
Step Edge GUI: A client-side GUI agent that supports visual perception, element location, and automated operations for mobile and desktop applications.
-
Step Edge GenIt supports edge image generation and editing, including text-to-image, image-to-image, and local editing, and is a leader in benchmarks such as OneIG and DPG-Bench.
Step Edge's technical principles
-
End-to-cloud collaborative architectureSimple, high-frequency tasks are processed locally on the device, while complex, long-chain inference is handled by the cloud, achieving an overall balance between speed, capability, and cost.
-
Full-modal local inferenceMultimodal data, including text, visual, and voice data, is processed locally on the terminal, ensuring that sensitive information does not leave the terminal and protecting privacy and security through local computing.
-
Step Inference NPU Engine: Self-developed terminal inference engine, which optimizes operators and memory for hardware such as mobile phones and automobiles, reducing end-to-end latency for input forms such as text, vision, and voice.
-
Low-latency Toolcall executionThe latency of local tool calls on the client side is as low as 0.1 seconds, supporting real-time response in high-frequency interaction scenarios and reducing dependence on cloud links.
How to use Step Edge
The Step Edge suite has not yet been officially launched for public testing or API release; it is currently in the release and unveiling phase.
Step Edge's project address
- Project official website: https://static.stepfun.com/blog/step-edge/
Step Edge's Competitive Comparison
| Comparison Dimensions | Step Edge | Qwen3-VL |
|---|---|---|
| Product Form | A complete set of edge-side models (basic + audio + GUI + Gen) | Open source multimodal large model series |
| End-side optimization | Deeply optimized for terminal scenarios, equipped with a self-developed NPU engine. | It supports edge deployment, but is primarily designed for general-purpose multimodal tasks. |
| GUI Agent | Leading in edge GUI model evaluations such as OSWorld | GUI automation needs to be achieved by integrating with external frameworks. |
| Audio capabilities | Independent Audio Model, ASR and Audio Understanding Integrated | The primary focus is on visual and text-based content; audio is not a core area. |
| Image generation | Built-in Gen model, supports on-device image generation and editing. | The focus is on understanding, with generation capabilities not being the primary concern. |
| Privacy Policy | Full-modal local processing, native edge-cloud collaboration | It can run on the client side, but the coordination strategy needs to be built by yourself. |
Step Edge Application Scenarios
-
Smartphone AssistantReal-time voice recognition and GUI automation on the device side enable offline completion of high-frequency tasks such as ordering takeout, sending messages, and checking photo albums.
-
In-vehicle intelligent cockpitLocal processing of voice commands and in-vehicle visual perception ensures driving data privacy and enables zero-latency control of navigation, air conditioning, and entertainment.
-
End-side image creationThe app allows users to create and edit photos locally on their phones without uploading them to the cloud, satisfying both their needs for instant creation and privacy protection.
-
Industrial terminal inspectionIn weak network factories or outdoor scenarios, the edge model can locally identify equipment malfunctions and audio anomalies, and report key information in real time.
-
IoT edge devicesDeployed in smart homes and wearable devices, it enables local voice wake-up, visual recognition, and simple decision-making, reducing cloud costs.