News
Alibaba open-sources Qwen3.8-Flash-Next, previewing the Qwen4 architecture.
Alibaba's Tongyi Qianwen open-sourced Qwen3.8-Flash-Next, which uses a 125B main model and 6B single-token activation parameters, and showcases the new architecture planned for Qwen4 in advance.
Alibaba's Tongyi Qianwen has released and opened up the Qwen3.8-Flash-Next weight. This multimodal MoE model contains 125B of main model parameters and 51B of N-gram embedding parameters, with approximately 6B of parameters activated for each token.
Qwen states that the model uses a hybrid architecture of GDN and Qwen Sparse Attention, and is a preview of the experimental architecture planned for Qwen4.
refer to:Qwen Official Announcement