AI isn't ready to research itself A study posted on arXiv in late July found that AI systems are not yet capable of fully automating open-ended research, with an AI system built by Princeton University researchers scoring 2/6 and 1/6 on two research tasks. The system, which used Claude Opus 4.8 and the agentic system OpenClaw, succeeded at engineering work but failed to explore hypotheses effectively and often settled too early. Sayash Kapoor, a computer scientist at Princeton and co-author of the preprint, said, "I don’t think full automation of open-ended research is on the horizon right now. Thank you for visiting nature.com. You are using a browser version with limited support for CSS. To obtain the best experience, we recommend you use a more up to date browser or turn off compatibility mode in Internet Explorer . In the meantime, to ensure continued support, we are displaying the site without styles and JavaScript. Artificial intelligence has made leaps in AI research itself, finding ways to make existing algorithms smarter — or writing new ones. But computers are not yet ready to replace their makers, according to a study posted on the preprint server arXiv1 in late July. “I don’t think full automation of open-ended research is on the horizon right now,” says Sayash Kapoor, a computer scientist at Princeton University in New Jersey and a co-author of the preprint. The effort to fully automate the scientific process, from idea generation to the writing and self-evaluation of a paper, was pioneered by a team mostly from the firm Sakana AI in Tokyo. The team unveiled their system, called The AI Scientist, in 2024, and later published results from an improved version in March in Nature2. The AI Scientist was tasked with studying pitfalls in machine learning. Three of the papers it produced were submitted for peer review at a conference workshop, and one achieved a score high enough for acceptance. But, according to Kapoor, peer review is an unreliable way of assessing the quality of a paper, especially in AI research. More generally, some researchers have questioned whether automating AI research from idea creation to publication can produce true breakthroughs yet, suggesting it is better at optimizing existing techniques. To hold AI to a higher standard than peer review, Kapoor and his colleagues created a new challenge, called shadow evaluation. First, they picked two papers that had been submitted to this year’s Neural Information Processing Systems conference. Then they asked an AI tool to do research and write papers based on a research question from each paper, and asked the original authors to scrutinize its output. The idea was that these authors would have more of the expertise and dedication to dig into the output than a harried peer reviewer would. The team built its AI system by ‘harnessing’ the large language model Claude Opus 4.8 in a modified version of the agentic system OpenClaw, and then wrapped that in a broader ‘scaffold’ that contained general instructions and ways to check progress. The harness gave the model an array of tools, including the ability to create sub-agents and to access the Internet, software libraries, computer processors for running experiments and software that simulates peer review. For each paper, the AI system had six days and US$3,000 in computing credits. Following the general research direction of the first paper, Kapoor’s team asked the system to design a method for the precise control of chatbot personality. Mirroring the second paper, it had to design a failure detector for a certain kind of neural network. The authors were surprised by the system’s success at the engineering work of its research. It could run hundreds of experiments over several days, without getting stuck in a loop of unresolvable errors. And it performed solid literature reviews and made some minor findings. It also caught its own false claims sometimes called hallucinations and did not try to cut corners or ‘reward hack’ , despite the authors’ predictions. Poor marks But the AI system mostly failed at its two assigned tasks, earning overall scores of 2/6 and 1/6 from the original papers’ authors. A typical way in which it would fail was to select a few hypotheses to explore, but settle too early on one and not backtrack sufficiently when its approach wasn’t working. Subsequent self-review wasn’t sufficiently negative, so the system persisted on its initial choices, whittling down its claims until it said little of interest. Enjoying our latest content? Log in or create an account to continue Access the most recent journalism from Nature's award-winning team Explore the latest features & opinion covering groundbreaking research