{"slug": "cogarena-a-multimethod-evaluation-of-cognitive-ability-structure-in-large-models", "title": "CogArena: A Multimethod Evaluation of Cognitive Ability Structure in Large Language Models", "summary": "A new study introducing CogArena, a 13-paradigm benchmark for evaluating cognitive ability structure in large language models, finds that across 55 open-weight models, paradigm correlations are positive and a common axis explains about half the variance, but the evidence does not establish stable five-dimensional cognitive profiles. The researchers report that targeted scaffolds show a small matched-grouping advantage, but no scaffold-specific contrast survives multiplicity correction and selectivity does not improve held-out-family prediction, leading to a boundary conclusion that theory-aligned prompting produces only a small in-battery diagonal tendency.", "body_md": "arXiv:2607.24999v1 Announce Type: new\nAbstract: LLM cognitive scores are increasingly summarized as per-ability profiles whose dimensions should converge across tasks, respond selectively to matched interventions, and generalize beyond the models used to define them. We introduce CogArena, a procedurally generated 13-paradigm benchmark built around a multimethod framework for determining when cognitive-task scores warrant dimensional labels across five theory-motivated groupings. Across 55 open-weight models, nearly all paradigm correlations are positive and a common axis explains about half the variance. The within-grouping advantage is small, scoring-sensitive, and uncertain across model families. In a separately frozen, fully crossed study across 12 models from six families, targeted scaffolds show a small matched-grouping advantage, but no scaffold-specific contrast survives multiplicity correction and selectivity does not improve held-out-family prediction. The frozen confirmation criterion fails. A post-hoc alternate-wording replication produces a smaller positive estimate and again fails. Together, these results support a boundary conclusion. Theory-aligned prompting produces a small in-battery diagonal tendency, but the present evidence does not establish stable five-dimensional profiles. CogArena provides a workflow joining behavioral signatures, covariance, matched interventions, and out-of-family prediction before cognitive labels are attached to model scores.", "url": "https://wpnews.pro/news/cogarena-a-multimethod-evaluation-of-cognitive-ability-structure-in-large-models", "canonical_source": "https://arxiv.org/abs/2607.24999", "published_at": "2026-07-29 04:00:00+00:00", "updated_at": "2026-07-29 04:26:22.253978+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-research"], "entities": ["CogArena", "arXiv"], "alternates": {"html": "https://wpnews.pro/news/cogarena-a-multimethod-evaluation-of-cognitive-ability-structure-in-large-models", "markdown": "https://wpnews.pro/news/cogarena-a-multimethod-evaluation-of-cognitive-ability-structure-in-large-models.md", "text": "https://wpnews.pro/news/cogarena-a-multimethod-evaluation-of-cognitive-ability-structure-in-large-models.txt", "jsonld": "https://wpnews.pro/news/cogarena-a-multimethod-evaluation-of-cognitive-ability-structure-in-large-models.jsonld"}}