Xiaomi MiMo-V2-Omni - Xiaomi's All-Modal Agent Base Model
Xiaomi MiMo-V2-Omni is a multimodal agent base model launched by Xiaomi, which integrates text, vision and speech modalities and natively possesses perception, reasoning and execution capabilities.
What is Xiaomi MiMo-V2-Omni?
Xiaomi MiMo-V2-Omni is a multi-modal Agent foundation model launched by Xiaomi, integrating text, vision, and speech modalities, and natively possessing perception, reasoning, and execution capabilities. The model supports tool invocation, GUI operation, and autonomous planning for complex tasks, and in evaluations of audio understanding and image reasoning, it rivals the Gemini 3 Pro and Claude Opus 4.6. The model, previously anonymously tested under the codename "Healer Alpha," topped the OpenRouter invocation leaderboard and has now become Xiaomi's core AI infrastructure for the Agent era.
Main features of Xiaomi MiMo-V2-Omni
-
Full-modal perceptionThe model integrates text, vision, and audio modalities to achieve image understanding, video analysis, processing of 10+ hours of audio, and cross-modal joint reasoning.
-
Agent execution capabilitiesIt natively supports tool calls, GUI operations, and autonomous task planning, enabling the formulation of strategies, real-time corrections, and end-to-end delivery of complete results.
-
Complex scenario applicationsIt covers real-world digital environment interaction tasks such as web browsing, code engineering, and front-end development.
The technical principles of Xiaomi MiMo-V2-Omni
- Unified full-modal architectureIt builds a base model that integrates text, vision, and speech from the bottom up, and achieves native multimodal representation through a unified encoder and fusion layer, without post-modal splicing.
- Deep integration of perception and actionBreaking away from the limitations of traditional models that emphasize understanding over execution, end-to-end training integrates perception capabilities with action capabilities such as tool invocation and GUI operation, achieving a leap from understanding to control.
- Video pre-training and long contextIt employs an innovative video pre-training method to achieve joint audio and video understanding, supports ultra-long context modeling, and provides structural advantages for complex agent tasks.
Key information and usage requirements for Xiaomi MiMo-V2-Omni
- PublisherXiaomi Technology Team
- Release timeMarch 19, 2026
- Internal test codenameHealer Alpha (formerly anonymously listed on OpenRouter)
- Model sizeFull-modal fusion architecture (text + vision + audio)
- Context windowSupports long sequence modeling (refer to the Pro version of the same series, which reaches up to 1MB).
- Benchmark rankings: First in average score on PinchBench, tops OpenRouter call volume
- Access methodIt can be seamlessly integrated with mainstream agent frameworks such as OpenClaw through API calls from platforms such as OpenRouter.
- Hardware/EnvironmentDeployed in the cloud, no local configuration required; supports multimodal input (images, videos, audio files or streams).
Xiaomi MiMo-V2-Omni's core advantages
- Full-modal native fusionIt builds a unified architecture for text, vision, and audio from the ground up, enabling true cross-modal understanding and joint reasoning, rather than simply splicing them together.
- Sensing and Action IntegrationBreaking away from the limitations of "emphasizing understanding over execution," it natively integrates capabilities such as tool calls and GUI operations, forming a combined advantage of "the more accurate the perception, the more effective the action."
- Long context supportIt supports millions of context windows, providing a structural advantage when handling long videos, long audios, and complex multi-round agent tasks.
- Real-world scenario verificationThe Healer Alpha anonymous internal test topped the OpenRouter call volume and ranked first in PinchBench, proving its effectiveness through both market and benchmark testing.
- Seamless Ecosystem IntegrationIt can quickly integrate with mainstream agent frameworks such as OpenClaw, significantly reducing the threshold for deploying full-modal agents.
How to use Xiaomi MiMo-V2-Omni
Developers can visit https://platform.xiaomimimo.com to register and obtain an API key, and then call the API at a set price (input $0.4/million tokens, output $2/million tokens).
Xiaomi MiMo-V2-Omni Competitive Product Comparison
| Evaluation Dimensions | MiMo-V2-Omni | Gemini 3 Pro | Claude Opus 4.6 |
|---|---|---|---|
| MMAU-Pro (Audio Understanding) | 69.4 | 67.0 | – |
| MMMU-Pro (Image Understanding) | 76.8 | 81.0 | 73.9 |
| Video-MME (Video Comprehension) | 85.3 | 88.4 | – |
| CharXiv RQ (Chart Comprehension) | 80.1 | 81.4 | 77.4 |
| FutureOmni (Future Prediction) | 66.7 | 62.9 | 60.3 |
| MM-BrowserComp (Web Browser) | 52.0 | 37.2 | 59.3 |
| OmniGAIA (Multimodal Sensing) | 49.8 | 62.5 | 59.7 |
| Claw Eval (Complex Interactions) | 54.8 | 51.9 | 66.3 |
| PinchBench (Agent Integration) | 85.6 | 75.0 | 86.3 |
Application scenarios of Xiaomi MiMo-V2-Omni
- Multimodal content understandingThe model supports analysis of 10+ hours of long videos, complex chart parsing, and cross-modal information association reasoning, enabling joint deep understanding of audio and video.
- Intelligent agent task executionThe model can autonomously complete tasks such as web browsing, code engineering, and front-end development, and can generate exquisitely designed and fully functional web pages with zero samples.
- GUI automationIt allows direct control of the graphical interface, supporting strategy planning, real-time correction, and autonomous toolchain invocation in multi-turn dialogues.
- Enterprise-level long document processingThe model relies on a 256K context window to complete long document analysis, report generation, and decision support for automated office processes.