cd /news/artificial-intelligence/cogarena-a-multimethod-evaluation-of… · home topics artificial-intelligence article
[ARTICLE · art-78053] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

CogArena: A Multimethod Evaluation of Cognitive Ability Structure in Large Language Models

A new study introducing CogArena, a 13-paradigm benchmark for evaluating cognitive ability structure in large language models, finds that across 55 open-weight models, paradigm correlations are positive and a common axis explains about half the variance, but the evidence does not establish stable five-dimensional cognitive profiles. The researchers report that targeted scaffolds show a small matched-grouping advantage, but no scaffold-specific contrast survives multiplicity correction and selectivity does not improve held-out-family prediction, leading to a boundary conclusion that theory-aligned prompting produces only a small in-battery diagonal tendency.

read1 min views1 publishedJul 29, 2026

arXiv:2607.24999v1 Announce Type: new Abstract: LLM cognitive scores are increasingly summarized as per-ability profiles whose dimensions should converge across tasks, respond selectively to matched interventions, and generalize beyond the models used to define them. We introduce CogArena, a procedurally generated 13-paradigm benchmark built around a multimethod framework for determining when cognitive-task scores warrant dimensional labels across five theory-motivated groupings. Across 55 open-weight models, nearly all paradigm correlations are positive and a common axis explains about half the variance. The within-grouping advantage is small, scoring-sensitive, and uncertain across model families. In a separately frozen, fully crossed study across 12 models from six families, targeted scaffolds show a small matched-grouping advantage, but no scaffold-specific contrast survives multiplicity correction and selectivity does not improve held-out-family prediction. The frozen confirmation criterion fails. A post-hoc alternate-wording replication produces a smaller positive estimate and again fails. Together, these results support a boundary conclusion. Theory-aligned prompting produces a small in-battery diagonal tendency, but the present evidence does not establish stable five-dimensional profiles. CogArena provides a workflow joining behavioral signatures, covariance, matched interventions, and out-of-family prediction before cognitive labels are attached to model scores.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @cogarena 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/cogarena-a-multimeth…] indexed:0 read:1min 2026-07-29 ·