cd /news/ai-agents/the-hard-part-comes-after-search-ben… · home › topics › ai-agents › article
[ARTICLE · art-140776] src=machinebrief.com ↗ pub= topic=ai-agents verified=true sentiment=· neutral

The Hard Part Comes After Search: Benchmarking Web Agents on Synthesizing, Organizing, and Displaying Knowledge

A new benchmark called KNOWS, introduced in arXiv paper 2609.30604v1, evaluates computer-use agents on open-ended, complex, browser-based tasks that end in a produced artifact such as a document, presentation, or spreadsheet. The best frontier computer-use agent fully succeeded on fewer than 3% of the long-horizon tasks, and failures on visual steps rendered artifacts unusable even when agents completed more than 50% of other evaluation steps. The authors conclude that current agents fall short as end-to-end assistants and call for progress on tool use, visual understanding, and long-horizon reasoning.

by read1 min views1 publishedSep 28, 2026

arXiv:2609.30604v1 Announce Type: new Abstract: Existing computer-use agent benchmarks do not fully evaluate agents acting as assistants. A useful assistant retrieves information across complex, multi-step workflows, synthesizes it into artifacts (documents, presentations, spreadsheets), and navigates program interfaces to produce a coherent final product. Such workflows demand reasoning and synthesis, decomposition of complex tasks, as well as visual and spatial understanding. To study agents on workflows like these, we introduce KNOWS, a benchmark of open-ended, complex, browser-based tasks that jointly evaluate these capabilities, with each task culminating in a produced artifact. To write tasks, we develop a task design rubric and a protocol for ensuring that tasks meet the requirements. Each task is paired with an evaluator, a program that combines deterministic checks with LLM judgments to balance the richness, reliability, and automation tradeoff inherent to agent evaluation. We evaluate and analyze frontier computer-use agents and browser-based harnesses. They achieve moderate scores on partial-success metrics, but the best performer fully succeeds in fewer than 3% of our complex, long-horizon tasks. Failures on visual steps render the resulting artifacts unusable, even when agents complete more than 50% of other evaluation steps. Our results expose limitations of current agents acting as end-to-end assistants, and call for progress on tool use, visual understanding, and long-horizon reasoning.

── more in #ai-agents 4 stories · sorted by recency
── more on @knows 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/the-hard-part-comes-…] indexed:0 read:1min 2026-09-28 · —