AB
AiBoss
project

Sage - SenseTime's large-scale edge-side multimodal intelligent agent base model

Sage is a large-scale edge-side multimodal intelligent agent base model launched by SenseTime Jueying. It adopts the MoE architecture, with a total of 32B parameters and only 3B activation parameters. The model has been deployed on the NVIDIA Orin X platform.

What is Sage?

Sage is a large-scale edge-side multimodal intelligent agent platform model launched by SenseTime's Jueying platform. It adopts the MoE architecture, with a total of 32 bytes of parameters and only 3 bytes of activation parameters. The model has been deployed on the NVIDIA Orin X platform. In PinchBench benchmarks, it achieved a task completion rate of 94%, surpassing cloud-based flagship models such as Claude-Opus-4.6 and GPT-5.4. Sage incorporates SCOUT and ERL post-training technologies, supporting compound instruction parsing, multi-system linkage, and proactive perception, providing cloud-level agent capabilities for intelligent cockpits.

Sage's main functions

  • Compound instruction parsingIt supports parsing multiple user intent commands at once and automatically linking in-vehicle systems such as air conditioning, navigation, and audio-visual systems to complete the task loop.
  • Proactive sensing serviceBy combining sensors to perceive the status of passengers and road conditions in real time, it can proactively trigger scenario-based services such as child mode and route adjustment.
  • Multimodal cockpit understandingLeveraging native in-vehicle data, it enables semantic and visual understanding of the cabin, accurately identifying the in-vehicle environment and user needs.
  • Tool Invocation and Task ClosureIt supports long-link tool calls and multi-step inference, achieving a score of 80 on the τ2-bench benchmark, and possesses true agent execution capabilities.
  • Real-time response at the edgeThe Orin X platform has a first-word response time of approximately 0.5 seconds, a single token latency of 0.03 seconds, and a generation throughput of 80 tk/s, without relying on the cloud.

Sage's technical principles

  • MoE Sparse Activation ArchitectureThe total number of parameters is 32B, and the activation parameter is only 3B. It achieves efficient inference under limited computing power on the edge through a sparse activation mechanism.
  • SCOUT graded collaborative learningIt adopts a hierarchical collaborative mechanism of "small model exploration and large model absorption", which saves about 60% of GPU hours in training complex tasks.
  • ERL Erasable Reinforcement LearningThe model automatically identifies and erases erroneous steps in the inference chain, preventing the spread of bias and improving the completion rate of complex tasks by 20%.
  • Integrated multimodal architectureA native training system that integrates visual, linguistic, and sensor data to build differentiated understanding capabilities for in-vehicle scenarios.

How to use Sage

The model is now deployed on the NVIDIA Orin X edge platform.

Key information and usage requirements for Sage

  • Model architecture: MoE architecture, total parameters 32B, activation parameters 3B.
  • Deployment platform: It has been deployed on the NVIDIA Orin X edge platform.
  • Evaluation results: PinchBench achieved a best task completion rate of 94%, surpassing mainstream models such as Claude-Opus-4.6 and GPT-5.4.
  • Hardware products: The SageBox, equipped with Sage, was launched during the Beijing Auto Show.
  • Target audience: It is primarily targeted at automakers, Tier 1 suppliers, and edge intelligent agent developers.
  • Network requirements: It can be deployed on the device side and run without relying on cloud network connections.
  • Framework compatibility: It supports integration with mainstream agent frameworks such as OpenClaw and Hermes.

Sage's core advantages

  • Leading end-side performanceUsing the 3B activation parameter, a 94% task completion rate was achieved on PinchBench, surpassing many high-parameter cloud flagship models.
  • Ultimate efficiency ratioCompared to MiMo-v2-Pro (42B activation), Sage's activation computing power is only 1/14, its memory usage is about 1/31, and its performance is 6.6 percentage points higher.
  • Training cost optimizationSCOUT technology saves approximately 60% of GPU hours for training complex tasks and reduces post-training costs.
  • Reasoning and error correction abilityERL technology enables the automatic identification and erasure of erroneous steps during the inference process, preventing task failure at its source.
  • Advantages of native in-vehicle dataThe Human Semantic Understanding test scored 91.5 points, leading other end-side models in its class by 32%, demonstrating a deep understanding of the cockpit environment.
  • Mass production feasibilityIt has been verified and deployed on the Orin X platform and is ready for automotive-grade mass production.

Comparison of Sage's similar products

Comparison Dimensions Sage Google Gemma 4 MiMo-v2-Pro
Publisher Shangtang Jueying Google Millet
Total number of parameters 32B Same order of magnitude end side Over 1T
Activation parameter quantity 3B not disclosed 42B
PinchBench completion rate 94% 83.9% 87.4%
MMLU Pro 75.8 69.2
GPQA Diamond 77.3 58.5
τ2-bench 80.7 42.1
Human Semantic Understanding 91.5 69.5
Deployment Platform Nvidia Orin X End side End side
Core positioning End-side intelligent agent base End-side general model End-side inference model

Sage application scenarios

  • Intelligent cockpit interaction: Users issue compound commands in natural language, and Sage parses them all at once and coordinates multiple in-vehicle systems such as air conditioning, navigation, and music to complete a service loop.
  • Child safety protection: When the rear-seat sensors detect a child in the car, Sage automatically activates child mode, locking the windows, switching to age-appropriate content, and limiting the volume.
  • Intelligent Mobility Planning: By combining real-time traffic conditions with user needs, Sage proactively recommends the optimal route and provides alternative options during congestion, thus transforming from a passive response to a proactive service.
  • Integrated cabin and driver service: As the core AI support for the cockpit-driver integration solution, Sage connects cockpit interaction and driving perception, providing intelligent agent services across all scenarios.