Alham Fikri Aji's 423-question benchmark spans 13 languages and tests whether agents can connect evidence across web pages, video, maps and documents.
By [Ryan Merket](https://runtimewire.com/author/ryan-merket)
· Published
Primary source: [X](https://x.com/AlhamFikri/status/2107446184340271478)
Why it matters #
The benchmark exposes two operational constraints behind search-agent claims: strong results can depend on a provider's own search stack, and alternative retrieval setups can spend hundreds of millions of tokens while answering fewer questions. It also shows why multilingual scores need difficulty controls before they can support clean comparisons.
A new benchmark from Alham Fikri Aji (@AlhamFikri) found that five evaluated AI systems failed to answer 244 of its 423 questions using their built-in web search. The result, reported in the team's paper, puts a number on a stubborn weakness in search agents: finding and checking evidence when clues cross languages and media formats.…