News
Zhipu has released and open-sourced GLM-5.3-Flash.
Zhipu has released GLM-5.3-Flash, the first native multimodal model in the GLM-5 series. It adopts a MoE architecture with 320 total parameters and 18 activation parameters, and the model weights have been made public.
Zhipu officially released GLM-5.3-Flash, the first native multimodal model in the GLM-5 series. The official description states that it adopts a MoE architecture with 320 total parameters and 18 activation parameters, and combines sparse attention and linear attention to reduce the cost of long-context inference.
The model weights have been publicly available on Hugging Face and can be deployed using frameworks such as SGLang, vLLM, and TokenSpeed.
refer to:Z.ai Official Announcement