CogGym: Towards Large-Scale Comparative Evaluation of Human and Machine Cognition Researchers introduced CogGym, a scalable framework grounded in cognitive science that standardizes 258 cognitive experiments from 100 papers into a task-agnostic Experiment Markup Language (EML) to compare human and machine cognition, and used it to evaluate 50 large language models against human responses. The team reports a clear scaling trend in which larger and more recent AI models better reproduce human judgments, but model-human fit remains well below human split-half reliability of R² = 0.93 on text, 0.95 on image, and 0.92 on video, with the best models reaching only R² = 0.59 on text, 0.58 on image, and 0.43 on video experiments. The authors intend CogGym to serve as a living evaluation framework that continually incorporates new cognitive science experiments. arXiv:2609.21259v1 Announce Type: new Abstract: Understanding and modeling human intelligence are parallel goals shared by artificial intelligence AI and cognitive science. As AI systems grow increasingly capable, in what ways do model responses resemble human responses, and where do they systematically diverge? The sheer breadth and diversity of the tasks humans can perform and think about pose a challenge for scalable and rigorous comparison between humans and models. We introduce CogGym, a scalable, unified framework grounded in cognitive science for systematically comparing model and human behavior on matched experimental trials. CogGym uses a semi-automated, human-in-the-loop pipeline to standardize diverse experimental paradigms into a task-agnostic Experiment Markup Language EML , enabling reproducible and faithful comparison at scale. For initial release, we curate and standardize 258 cognitive experiments from 100 papers that focuses on human commonsense reasoning, and evaluate 50 large language models against human responses. We find a clear scaling trend where larger and more recent AI models better reproduce human judgments. Yet AI models' improvement on such common reasoning tasks is considerably slower than the gains observed on formal-reasoning benchmarks like math and coding, and model--human fit remains well below human splithalf reliability $R^2 = 0.93$ on text, $0.95$ on image, and $0.92$ on video with the best models achieving $R^2 = 0.59$ on text, $0.58$ on image, and $0.43$ on video experiments. We intend for CogGym to provide a living evaluation framework that continually incorporates new cognitive science experiments to characterize where model behavior resembles human behavior, where it systematically diverges, and how those patterns change as models and experiments evolve.