22:10
2026-07-26
arxiv.org
large-language-models
Frontier LLMs drop from 83% to 43% once reasoning has to chain across domains
A new benchmark called Relay-Bench shows that frontier large language models (LLMs) drop from an average of 83% on single-domain tasks to 43% when required to chain reasoning across multiple domains iā¦