AB
AiBoss
News

ZCube, the next-generation large-model inference network architecture, is launched by Zhipu.

Zhipu, in collaboration with Yuhun Networks and Tsinghua University, launched the ZCube networking architecture. Addressing the congestion challenges of PD-separated inference, it eliminates the Spine layer and adopts a flat topology with hybrid single/multi-track access. Real-world testing with GLM-5.1 coding shows that ZCube reduces switch and optical module costs by 33%, increases GPU inference throughput by 15%, and reduces first-to-token latency (TTFT P99) by 40.6%, providing a high-efficiency foundation for next-generation hyperscale inference clusters.