The anonymous model codenamed Elephant Alpha has been officially unveiled: Ling-2.6-flash
Ant Financial's large model team has released Ling-2.6-flash, with a total of 104B parameters and 7.4B activation parameters. It adopts a hybrid attention architecture of MLA+Lightning Linear and sparse MoE. The model achieves an inference speed of 340 tokens/s in a 4-card H2O environment, with token consumption only about 1/10 of its competitors. It achieves state-of-the-art performance on Agent benchmarks such as BFCL-V4 and SWE-bench Verified.