News
Alibaba's Tongyi platform launches the multimodal MoE model Qwen3.8-Flash.
Alibaba officially launched the multimodal MoE model Qwen3.8-Flash, with a total of 125 bytes of parameters. Only 6 bytes are activated per token, and it natively supports 260,000 token contexts, expandable to 1 million. The model features four major upgrades: hybrid attention using GDN and QSA, gated residuals, N-gram embeddings, and a Muon optimizer. Training costs are reduced to 1/9 of the previous generation, resulting in enhanced coding and office capabilities. The model is offered as an API service with a pricing of 1 yuan per input and 3 yuan per million tokens per output.