News
Alibaba's Tongyi launches PawBench, a general benchmark for evaluating intelligent agents.
Tongyi Labs has launched PawBench, a general-purpose intelligent agent evaluation benchmark, which for the first time incorporates the base model and runtime framework (Harness) into a joint evaluation. PawBench v1.0 includes 150 real-world tasks and 4050 test units, covering a cross-matrix of 9 models and 3 Harnesses. The evaluation found that the performance difference of Harnesses can be as high as 6.4 points, and the difference can reach 11.5 points when changing Harnesses for the same model.