InnoEval and new benchmarks show AI models struggle with original research A cluster of 2025-2026 studies led by the InnoEval evaluation framework found that frontier AI models recover the central idea of a paper from its pre-publication reference list only 3-15% of the time across 643 papers, and that AI-generated hypotheses scored an average novelty of 3.406 versus 3.968 for human hypotheses in an October 2026 Science study. InnoEval, introduced in a February 2026 arXiv paper and presented at ICML 2026, beat traditional LLM-as-judge methods by up to 16.18% in F1 scores, while a separate October 2026 study of more than 121,000 preprints found LLMs generate narrowly focused, low-diversity ideas. The benchmarks InnoGym and InnovatorBench reported low success rates for complete LLM pipelines producing genuinely new methods, indicating current models work better as research assistants than as sources of original hypotheses. Photo: Tima Miroshnichenko / Pexels InnoEval and new benchmarks show AI models struggle with original research A wave of 2025-2026 studies finds frontier AI models are far better at remixing known science than inventing new methods Frontier AI models can summarize a thousand papers before your coffee cools. Ask them to invent a genuinely new research technique from scratch, and the results get a lot less impressive. That is the shared conclusion of a cluster of studies from 2025 and 2026, led by a new evaluation framework called InnoEval. Together, they suggest AI models underperform when asked to develop new research methods without prior information to lean on. What the research actually found InnoEval first appeared in an arXiv paper in February 2026 and was later presented at ICML 2026. Its goal is to measure innovation in a structured way, using what its authors describe as knowledge-grounded, multi-perspective evaluations. That approach paid off on its own terms. InnoEval beat traditional LLM-as-judge methods by up to 16.18% in F1 scores on specific tasks. Better measurement, however, produced some unflattering readings. The more carefully researchers looked, the less originality they found. Take the “Reconstruction” benchmark. It tested whether AI models could recover the core idea of a paper using only the reference list that existed before publication. Across 643 papers, as of August 2026, models recovered the central idea only 3–15% of the time. AI, tech, and the markets they move—in one daily briefing. Daily. Free. Join 34,000+ readers across crypto, finance, and policy. A study published in Science in October 2026 compared AI-generated hypotheses with human ones directly. The AI hypotheses scored an average novelty of 3.406. Human hypotheses averaged 3.968. The diversity problem A separate large-scale study, also from October 2026, analyzed more than 121,000 preprints. It found that LLMs often generate ideas that are narrowly focused and show little diversity. Two more benchmarks, InnoGym and InnovatorBench, were introduced across 2025 and 2026 to probe the same question from a workflow perspective. Rather than grading a single answer, they test complete LLM pipelines on a range of real-world tasks. The results reported low success rates in producing genuinely new methods. Why this keeps coming up Until recently, a common way to judge AI-generated ideas was to ask another AI model whether they were novel. Frameworks like InnoEval exist because that approach was not reliable enough. The 16.18% F1 improvement suggests grounding evaluations in real knowledge catches things that a lone LLM judge misses. What this means for AI labs, researchers and investors Current models appear well suited as research assistants that surface literature and test hypotheses. They appear much less suited to serving as the source of the hypotheses themselves, at least when working without strong prior context. The 121,000-preprint study hints at the risk that a tool which reliably proposes the same narrow cluster of ideas could quietly reduce the diversity of research directions a lab explores. None of this means AI cannot contribute to original research. What the studies do establish is a clearer baseline, with numbers like a 3–15% recovery rate and a novelty score of 3.406 against 3.968 that future systems will have to beat. Disclosure: This article was edited by Diego Almada Lopez. For more information on how we create and review content, see our Editorial Policy https://cryptobriefing.com/editorial-policy/ .