04:00
2026-08-12
arxiv.org
large-language-models
From Reasoning Depth to Reasoning Breadth: Evaluating Multi-Point Associative Reasoning in Large Language Models
Researchers introduced MPAR-Bench, a bilingual English-Chinese benchmark with 1,000 items that tests large language models' reasoning breadth through multi-point associative reasoning, finding that peโฆ