Sage - SenseTime's large-scale edge-side multimodal intelligent agent base model
Sage is a large-scale edge-side multimodal intelligent agent base model launched by SenseTime Jueying. It adopts the MoE architecture, with a total of 32B parameters and only 3B activation parameters. The model has been deployed on the NVIDIA Orin X platform.
What is Sage?
Sage is a large-scale edge-side multimodal intelligent agent platform model launched by SenseTime's Jueying platform. It adopts the MoE architecture, with a total of 32 bytes of parameters and only 3 bytes of activation parameters. The model has been deployed on the NVIDIA Orin X platform. In PinchBench benchmarks, it achieved a task completion rate of 94%, surpassing cloud-based flagship models such as Claude-Opus-4.6 and GPT-5.4. Sage incorporates SCOUT and ERL post-training technologies, supporting compound instruction parsing, multi-system linkage, and proactive perception, providing cloud-level agent capabilities for intelligent cockpits.
Sage's main functions
- Compound instruction parsingIt supports parsing multiple user intent commands at once and automatically linking in-vehicle systems such as air conditioning, navigation, and audio-visual systems to complete the task loop.
- Proactive sensing serviceBy combining sensors to perceive the status of passengers and road conditions in real time, it can proactively trigger scenario-based services such as child mode and route adjustment.
- Multimodal cockpit understandingLeveraging native in-vehicle data, it enables semantic and visual understanding of the cabin, accurately identifying the in-vehicle environment and user needs.
- Tool Invocation and Task ClosureIt supports long-link tool calls and multi-step inference, achieving a score of 80 on the τ2-bench benchmark, and possesses true agent execution capabilities.
- Real-time response at the edgeThe Orin X platform has a first-word response time of approximately 0.5 seconds, a single token latency of 0.03 seconds, and a generation throughput of 80 tk/s, without relying on the cloud.
Sage's technical principles
- MoE Sparse Activation ArchitectureThe total number of parameters is 32B, and the activation parameter is only 3B. It achieves efficient inference under limited computing power on the edge through a sparse activation mechanism.
- SCOUT graded collaborative learningIt adopts a hierarchical collaborative mechanism of "small model exploration and large model absorption", which saves about 60% of GPU hours in training complex tasks.
- ERL Erasable Reinforcement LearningThe model automatically identifies and erases erroneous steps in the inference chain, preventing the spread of bias and improving the completion rate of complex tasks by 20%.
- Integrated multimodal architectureA native training system that integrates visual, linguistic, and sensor data to build differentiated understanding capabilities for in-vehicle scenarios.
How to use Sage
The model is now deployed on the NVIDIA Orin X edge platform.
Key information and usage requirements for Sage
- Model architecture: MoE architecture, total parameters 32B, activation parameters 3B.
- Deployment platform: It has been deployed on the NVIDIA Orin X edge platform.
- Evaluation results: PinchBench achieved a best task completion rate of 94%, surpassing mainstream models such as Claude-Opus-4.6 and GPT-5.4.
- Hardware products: The SageBox, equipped with Sage, was launched during the Beijing Auto Show.
- Target audience: It is primarily targeted at automakers, Tier 1 suppliers, and edge intelligent agent developers.
- Network requirements: It can be deployed on the device side and run without relying on cloud network connections.
- Framework compatibility: It supports integration with mainstream agent frameworks such as OpenClaw and Hermes.
Sage's core advantages
- Leading end-side performanceUsing the 3B activation parameter, a 94% task completion rate was achieved on PinchBench, surpassing many high-parameter cloud flagship models.
- Ultimate efficiency ratioCompared to MiMo-v2-Pro (42B activation), Sage's activation computing power is only 1/14, its memory usage is about 1/31, and its performance is 6.6 percentage points higher.
- Training cost optimizationSCOUT technology saves approximately 60% of GPU hours for training complex tasks and reduces post-training costs.
- Reasoning and error correction abilityERL technology enables the automatic identification and erasure of erroneous steps during the inference process, preventing task failure at its source.
- Advantages of native in-vehicle dataThe Human Semantic Understanding test scored 91.5 points, leading other end-side models in its class by 32%, demonstrating a deep understanding of the cockpit environment.
- Mass production feasibilityIt has been verified and deployed on the Orin X platform and is ready for automotive-grade mass production.
Comparison of Sage's similar products
| Comparison Dimensions | Sage | Google Gemma 4 | MiMo-v2-Pro |
|---|---|---|---|
| Publisher | Shangtang Jueying | Millet | |
| Total number of parameters | 32B | Same order of magnitude end side | Over 1T |
| Activation parameter quantity | 3B | not disclosed | 42B |
| PinchBench completion rate | 94% | 83.9% | 87.4% |
| MMLU Pro | 75.8 | 69.2 | – |
| GPQA Diamond | 77.3 | 58.5 | – |
| τ2-bench | 80.7 | 42.1 | – |
| Human Semantic Understanding | 91.5 | 69.5 | – |
| Deployment Platform | Nvidia Orin X | End side | End side |
| Core positioning | End-side intelligent agent base | End-side general model | End-side inference model |
Sage application scenarios
- Intelligent cockpit interaction: Users issue compound commands in natural language, and Sage parses them all at once and coordinates multiple in-vehicle systems such as air conditioning, navigation, and music to complete a service loop.
- Child safety protection: When the rear-seat sensors detect a child in the car, Sage automatically activates child mode, locking the windows, switching to age-appropriate content, and limiting the volume.
- Intelligent Mobility Planning: By combining real-time traffic conditions with user needs, Sage proactively recommends the optimal route and provides alternative options during congestion, thus transforming from a passive response to a proactive service.
- Integrated cabin and driver service: As the core AI support for the cockpit-driver integration solution, Sage connects cockpit interaction and driving perception, providing intelligent agent services across all scenarios.