04:00
2026-08-03
arxiv.org
artificial-intelligence
Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review
A benchmarking study of autonomous AI research systems found that FARS benchmark papers significantly outperform four leading AI Scientist frameworks, achieving mean scores of 2.14β2.47 on a 1β5 scaleβ¦