Epoch AI's InnovationEval benchmark shows current models lag human algorithmic problem-solving Epoch AI's InnovationEval benchmark shows that current AI models lag human researchers at replicating algorithmic breakthroughs, indicating a measurable capability gap in novel problem-solving. The benchmark scored 89 on Hacker News, and Epoch AI says the result suggests limits to agent capability for independent research and discovery tasks. Epoch AI's InnovationEval benchmark shows current models lag human algorithmic problem-solving According to Epoch AI, the InnovationEval benchmark demonstrates that recent AI models struggle to replicate human algorithmic breakthroughs, suggesting a measurable capability gap in novel problem-solving compared to human researchers. The benchmark score of 89 on Hacker News indicates current models have not yet matched human-level innovation in algorithmic domains. The finding suggests limits to agent capability for independent research and discovery tasks. Topics Sources - Press Read article https://epoch.ai/publications/innovationeval Go deeper This intelligence is sourced automatically from public sources across the web and synthesised by the Prefactor AI pipeline. Stories are reviewed before publication.