AB
AiBoss
News

Meituan launches next-generation search agent benchmark LoHoSearch

Meituan's LongCat team has launched a new benchmark for evaluating search agents, LoHoSearch, which automatically generates 544 highly challenging questions using a Wikipedia entity knowledge graph containing 7.62 million entries. The benchmark increases difficulty by controlling both the search space and structural complexity. Evaluations show that the most powerful model, GPT-5.5, only achieves an accuracy of 34.74%, significantly lower than its performance on BrowseComp, revealing a significant bottleneck for current search agents in long-chain, complex reasoning tasks.