Spark X2-Flash - A large language model based on the MoE architecture launched by iFlytek.
Spark X2-Flash is a large language model based on the MoE architecture released by iFlytek. It has a total of 30B parameters, supports 256K ultra-long context, and is trained on a domestic computing power cluster based on Huawei Ascend 910B.
What is Spark X2-Flash?
Xinghuo X2-Flash is a large language model based on the MoE architecture released by iFlytek. It has a total of 30 bytes of parameters, supports 256KB of ultra-long context, and is trained on a domestic computing cluster powered by Huawei Ascend 910B. Designed specifically for the Agent era, the model performs close to trillion-parameter models in scenarios such as agent task execution, code generation, and deep research, while its token consumption cost is less than one-third of mainstream large-scale models. The model achieves efficient training and inference through techniques such as DSA sparse attention and MTP multi-token prediction. Its API is open and integrated with platforms such as AstronClaw and Loomy.
Main functions of Xinghuo X2-Flash
-
Intelligent agent task executionIt supports complex agent workflows such as in-depth research report generation, skill management and invocation, and system control and execution, with results approaching those of a trillion-parameter model.
-
Code generationIt can quickly generate complex skills (such as AI video-generated skills), including complete descriptions of skill structure, core functions, and use cases.
-
Long context processingIt supports a maximum context window of 256K, which can meet the consumption needs of hundreds of thousands or even millions of tokens in long-chain Agent tasks.
-
Multi-platform accessIt has been integrated with products such as AstronClaw and Loomy, and is compatible with mainstream agent frameworks such as OpenClaw and Claude Code.
-
API serviceThe Xingchen Coding Plan fully supports this model by providing API calls through the iFlytek Open Platform and Xingchen MaaS Platform.
Technical Principles of Spark X2-Flash
-
MoE architectureThe model employs a hybrid expert architecture with a total of 30 parameters, achieving higher efficiency while maintaining performance.
-
Domestic computing power trainingTraining was completed based on the Huawei Ascend 910B cluster, and deep optimization was achieved through operators that are compatible with domestic chips and distributed training strategies.
-
Intelligent agent data closed loop: Construct a verifiable large-scale intelligent agent data automatic synthesis platform, in which the agent autonomously builds the environment and tests the accuracy of the results to achieve efficient data synthesis and closed-loop.
-
Efficient Training of Long TextsIt is the first domestic computing power to combine DSA (sparse attention) and MTP (multi-token prediction), with the context extended to 256K, and the training efficiency increased from 20% to 90% compared with the A800 cluster of the same size.
-
Sampling decoding efficiency optimizationIn reinforcement learning training scenarios, through algorithmic and engineering innovations, sampling and decoding efficiency can be improved by more than 2 times, alleviating the computational barrier for RL training in long-interaction scenarios.
Key information and usage requirements of Spark X2-Flash
-
Model NameSpark X2-Flash
-
PublisheriFlytek / iFlytek Open Platform
-
Model ArchitectureMoE (Hybrid Expert), Total Parameters 30B
-
Context windowMaximum support 256K
-
Training computing powerHuawei Ascend 910B domestic cluster
-
Already connected to the platformAstronClaw, Loomy
-
API EntryiFlytek Open Platform, Xingchen MaaS Platform
-
Compatible frameworkMainstream agent frameworks such as OpenClaw and Claude Code
- Usage requirements:
-
Developers can call the API through the iFlytek Open Platform or the StarMaaS Platform.
-
The Starry Sky Coding Plan fully supports this model, and both new and existing users can switch to it independently.
-
The core advantages of Spark X2-Flash
-
Extremely high cost performanceThe performance of complex agent tasks is close to that of models with trillions of parameters, while token consumption is less than one-third of that of mainstream large-scale models.
-
Domestic computing power is independent and controllableTraining is based on Huawei Ascend 910B clusters, and it runs efficiently on local computing power architecture.
-
Extremely long contextA 256K context window meets the long-link requirements of complex agent workflows.
-
Breakthrough in training efficiencyThrough DSA+MTP technology, the training efficiency of domestically produced computing power has been increased from 20% to 90%.
-
Fast reasoning speedSampling and decoding efficiency is improved by more than 2 times, and the training time for reinforcement learning is significantly reduced.
-
Agent native optimizationDeeply compatible with mainstream agent frameworks such as OpenClaw, supporting automatic closed-loop data synthesis for intelligent agents.
-
Rapid Ecosystem AccessIt has been integrated with applications such as AstronClaw and Loomy, and developers can use it immediately.
Comparison of Spark X2-Flash with similar competing products
| Comparison Dimensions | Spark X2-Flash | DeepSeek-V3 | Qwen2.5-72B |
|---|---|---|---|
| Parameter size | 30B (MoE) | 671B MoE (37B per activation) | 72B (Dense) |
| Context window | 256K | 128K | 128K |
| Model Architecture | MoE | MoE | Dense architecture |
| Training computing power | Huawei Ascend 910B (domestic) | NVIDIA H800 Cluster | GPUs from NVIDIA, AMD, and other companies |
| Open source situation | Closed-source (API service) | Open source (can be deployed locally) | Open source (can be deployed locally) |
| Agent adaptation | Natively optimized, deeply compatible with OpenClaw and Claude Code | Strong versatility, Agent ecosystem relies on community/third parties | Strong versatility, Agent ecosystem relies on community/third parties |
| Task effect | Model with nearly a trillion parameters | Approaching GPT-40 level, with outstanding math/code skills. | Excellent overall capabilities and strong multilingual support. |
| Token Cost | Less than 1/3 the size of mainstream large-size models | API pricing is lower (about 1/10 of GPT-4o). | API pricing is lower (about 1/20th the price of GPT-4o). |
| Core positioning | The cost-effective engine of the Agent era | High-performance open-source base model | Open source ecosystem flagship model |
Application scenarios of Spark X2-Flash
-
Complex Agent Workflow: In-depth research report generation, multi-step task breakdown and execution, and multi-round context reading and correction.
-
Skill/Tool DevelopmentAutomatically generate and manage complex skills (such as AI video generation skills), including structure definition, core functions, and use cases.
-
Code generation and system controlScenarios requiring coding skills, such as script writing, system command execution, and automated operation and maintenance.
-
Long document analysisIt can process extremely long documents, papers, and reports based on a 256K context, and perform abstracting, extraction, and question answering.
-
Multimodal task orchestrationAs the brain of the Agent, it coordinates multiple platform toolchains such as text-based video and image-based video (e.g., Keling, Runway, Pika).