ADAGE: A Language-Agnostic Pipeline for Analogical Reasoning Evaluation Researchers introduce ADAGE, a language-agnostic pipeline for analogical reasoning evaluation that uses native-speaker curation and LLM-assisted generation to create translation-free benchmarks. Testing on Arabic, Amharic, and Japanese, they find that 14 open-weight models show a consistent cultural reasoning gap, with accuracy dropping 12–52 percentage points relative to English. arXiv:2607.23058v1 Announce Type: new Abstract: Multilingual reasoning evaluation overwhelmingly relies on translating English benchmarks, a practice that introduces linguistic artifacts and fails to test culturally-grounded reasoning. We introduce ADAGE Analogical Difficulty-by-design Assessment for Grounded Evaluation , a language-agnostic pipeline that combines native-speaker curation with LLM-assisted generation to construct challenging, translation-free benchmarks for abstract analogical reasoning. We validate ADAGE by constructing benchmarks for Arabic, Amharic, and Japanese. Evaluating 14 open-weight models, we find a consistent cultural reasoning gap: models that perform well on English proverb reasoning struggle substantially on all three native benchmarks, with accuracy dropping by 12--52 percentage points relative to English. We release the pipeline, all three benchmarks, and the full evaluation suite.