AB
AiBoss
project

Spark X2-Flash - A large language model based on the MoE architecture launched by iFlytek.

Spark X2-Flash is a large language model based on the MoE architecture released by iFlytek. It has a total of 30B parameters, supports 256K ultra-long context, and is trained on a domestic computing power cluster based on Huawei Ascend 910B.

What is Spark X2-Flash?

Xinghuo X2-Flash is a large language model based on the MoE architecture released by iFlytek. It has a total of 30 bytes of parameters, supports 256KB of ultra-long context, and is trained on a domestic computing cluster powered by Huawei Ascend 910B. Designed specifically for the Agent era, the model performs close to trillion-parameter models in scenarios such as agent task execution, code generation, and deep research, while its token consumption cost is less than one-third of mainstream large-scale models. The model achieves efficient training and inference through techniques such as DSA sparse attention and MTP multi-token prediction. Its API is open and integrated with platforms such as AstronClaw and Loomy.

Main functions of Xinghuo X2-Flash

  • Intelligent agent task executionIt supports complex agent workflows such as in-depth research report generation, skill management and invocation, and system control and execution, with results approaching those of a trillion-parameter model.
  • Code generationIt can quickly generate complex skills (such as AI video-generated skills), including complete descriptions of skill structure, core functions, and use cases.
  • Long context processingIt supports a maximum context window of 256K, which can meet the consumption needs of hundreds of thousands or even millions of tokens in long-chain Agent tasks.
  • Multi-platform accessIt has been integrated with products such as AstronClaw and Loomy, and is compatible with mainstream agent frameworks such as OpenClaw and Claude Code.
  • API serviceThe Xingchen Coding Plan fully supports this model by providing API calls through the iFlytek Open Platform and Xingchen MaaS Platform.

Technical Principles of Spark X2-Flash

  • MoE architectureThe model employs a hybrid expert architecture with a total of 30 parameters, achieving higher efficiency while maintaining performance.
  • Domestic computing power trainingTraining was completed based on the Huawei Ascend 910B cluster, and deep optimization was achieved through operators that are compatible with domestic chips and distributed training strategies.
  • Intelligent agent data closed loop: Construct a verifiable large-scale intelligent agent data automatic synthesis platform, in which the agent autonomously builds the environment and tests the accuracy of the results to achieve efficient data synthesis and closed-loop.
  • Efficient Training of Long TextsIt is the first domestic computing power to combine DSA (sparse attention) and MTP (multi-token prediction), with the context extended to 256K, and the training efficiency increased from 20% to 90% compared with the A800 cluster of the same size.
  • Sampling decoding efficiency optimizationIn reinforcement learning training scenarios, through algorithmic and engineering innovations, sampling and decoding efficiency can be improved by more than 2 times, alleviating the computational barrier for RL training in long-interaction scenarios.

Key information and usage requirements of Spark X2-Flash

  • Model NameSpark X2-Flash
  • PublisheriFlytek / iFlytek Open Platform
  • Model ArchitectureMoE (Hybrid Expert), Total Parameters 30B
  • Context windowMaximum support 256K
  • Training computing powerHuawei Ascend 910B domestic cluster
  • Already connected to the platformAstronClaw, Loomy
  • API EntryiFlytek Open Platform, Xingchen MaaS Platform
  • Compatible frameworkMainstream agent frameworks such as OpenClaw and Claude Code
  • Usage requirements:
    • Developers can call the API through the iFlytek Open Platform or the StarMaaS Platform.
    • The Starry Sky Coding Plan fully supports this model, and both new and existing users can switch to it independently.

The core advantages of Spark X2-Flash

  • Extremely high cost performanceThe performance of complex agent tasks is close to that of models with trillions of parameters, while token consumption is less than one-third of that of mainstream large-scale models.
  • Domestic computing power is independent and controllableTraining is based on Huawei Ascend 910B clusters, and it runs efficiently on local computing power architecture.
  • Extremely long contextA 256K context window meets the long-link requirements of complex agent workflows.
  • Breakthrough in training efficiencyThrough DSA+MTP technology, the training efficiency of domestically produced computing power has been increased from 20% to 90%.
  • Fast reasoning speedSampling and decoding efficiency is improved by more than 2 times, and the training time for reinforcement learning is significantly reduced.
  • Agent native optimizationDeeply compatible with mainstream agent frameworks such as OpenClaw, supporting automatic closed-loop data synthesis for intelligent agents.
  • Rapid Ecosystem AccessIt has been integrated with applications such as AstronClaw and Loomy, and developers can use it immediately.

Comparison of Spark X2-Flash with similar competing products

Comparison Dimensions Spark X2-Flash DeepSeek-V3 Qwen2.5-72B
Parameter size 30B (MoE) 671B MoE (37B per activation) 72B (Dense)
Context window 256K 128K 128K
Model Architecture MoE MoE Dense architecture
Training computing power Huawei Ascend 910B (domestic) NVIDIA H800 Cluster GPUs from NVIDIA, AMD, and other companies
Open source situation Closed-source (API service) Open source (can be deployed locally) Open source (can be deployed locally)
Agent adaptation Natively optimized, deeply compatible with OpenClaw and Claude Code Strong versatility, Agent ecosystem relies on community/third parties Strong versatility, Agent ecosystem relies on community/third parties
Task effect Model with nearly a trillion parameters Approaching GPT-40 level, with outstanding math/code skills. Excellent overall capabilities and strong multilingual support.
Token Cost Less than 1/3 the size of mainstream large-size models API pricing is lower (about 1/10 of GPT-4o). API pricing is lower (about 1/20th the price of GPT-4o).
Core positioning The cost-effective engine of the Agent era High-performance open-source base model Open source ecosystem flagship model

Application scenarios of Spark X2-Flash

  • Complex Agent Workflow: In-depth research report generation, multi-step task breakdown and execution, and multi-round context reading and correction.
  • Skill/Tool DevelopmentAutomatically generate and manage complex skills (such as AI video generation skills), including structure definition, core functions, and use cases.
  • Code generation and system controlScenarios requiring coding skills, such as script writing, system command execution, and automated operation and maintenance.
  • Long document analysisIt can process extremely long documents, papers, and reports based on a 256K context, and perform abstracting, extraction, and question answering.
  • Multimodal task orchestrationAs the brain of the Agent, it coordinates multiple platform toolchains such as text-based video and image-based video (e.g., Keling, Runway, Pika).