04:00
2026-07-22
machinebrief.com
large-language-models
Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains
A new benchmark called Relay-Bench, introduced in arXiv:2607.18438v1, measures large language models' ability to complete multi-domain reasoning chains in a single prompt, with the leading model GPT-5โฆ