HyperBrowseComp leaves 58% of multilingual search questions unsolved Alham Fikri Aji's new HyperBrowseComp benchmark found that five evaluated AI systems failed to answer 244 of its 423 multilingual search questions using their built-in web search, leaving 58% unsolved. The 423-question benchmark spans 13 languages and tests whether agents can connect evidence across web pages, video, maps and documents, according to the team's paper. The results show strong search-agent performance can depend on a provider's own search stack, while alternative retrieval setups can spend hundreds of millions of tokens while answering fewer questions. HyperBrowseComp leaves 58% of multilingual search questions unsolved Alham Fikri Aji's 423-question benchmark spans 13 languages and tests whether agents can connect evidence across web pages, video, maps and documents. By Ryan Merket https://runtimewire.com/author/ryan-merket · Published Primary source: X https://x.com/AlhamFikri/status/2107446184340271478 Why it matters The benchmark exposes two operational constraints behind search-agent claims: strong results can depend on a provider's own search stack, and alternative retrieval setups can spend hundreds of millions of tokens while answering fewer questions. It also shows why multilingual scores need difficulty controls before they can support clean comparisons. A new benchmark from Alham Fikri Aji @AlhamFikri https://x.com/AlhamFikri found that five evaluated AI systems failed to answer 244 of its 423 questions using their built-in web search. The result, reported in the team's paper https://arxiv.org/abs/2610.03574 , puts a number on a stubborn weakness in search agents: finding and checking evidence when clues cross languages and media formats.…