"[\[It's\] better than PhD level in everything.](https://www.lbc.co.uk/article/3c9d41195ef74647a7055ab724047cdd-5Hjd75W_2/)"
-- Elon Musk
Patrick Boyle recommended a video from the House of El in his latest post. The video [below] reviews a number of recent papers related to RSI and concludes that, while LLMs are an impressive technology that has made remarkable advances, scientific research still needs the human touch. [Emphasis added.]
Now look back at every impressive result that I just described.
Darwin Gödel Machine tests its modifications against a benchmark, a benchmark with an externally measurable score. Right? AlphaEvolve evolves algorithms evaluated by execution. Is the new procedure faster? Yes or no?
Anthropic's agents optimized against a measurable objective. Did the target metric go up or did it go down? RE-Bench measured performance against tasks with defined outcomes.Every single one of them had something outside itself telling it whether it succeeded or not.*** An external evaluator, a ground truth, a scoreboard of a sort.*** And remember that 97% result from Anthropic that I mentioned. There was a catch. Of course, the agents found several ways to game the experiment. In one case, they effectively got access to information they were not supposed to see. In another, they ran experiments under different random conditions and picked the ones that made them look the best.
Basically, in other words, they became very good at improving the score, but that is not necessarily the same thing as solving the underlying scientific problem.And when the best method the agents discovered was transferred to production scale, the improvement was, in Anthropic's own measurement, approximately half a point within the noise floor. *** The headline is 97%, but the real-world transfer is essentially nothing***. That's a bit like putting fluent in French on your CV because you got 97% on Duolingo and then getting to Paris and not being able to order bread.Now take the scoreboard away entirely and you're essentially asking the system to do something qualitatively very different. You're asking it to evaluate its own ideas without knowing the answer in advance. You are asking it to do what a researcher does, a human researcher, when they sit in their office and stare at the ceiling and wonder whether their entire approach to the problem is wrong or if they chose the wrong career, whatever.
That is a very different problem from optimizing against a benchmark. In real scientific work, there may be no objective function telling you that an idea is promising. The human researcher has to decide whether the question matters in the first place, whether the experiment is informative at all, whether a result is novel, whether an apparent improvement is a confound, and whether an entire line of inquiry should be abandoned.
That's why it's so difficult and everyone's a little sad all the time, or at least exhausted.There is no benchmark for, "Is this a good research direction?" There is no automated evaluator for scientific taste. If there were, half of academia would be unemployed and the other half would just be on antidepressants. Well, maybe more on antidepressants.Anyway, the survey's language is also worth quoting here. They say the research direction-setting bottleneck that keeps humans in the loop is its top rung. *** So again, AI handles execution; judgment remains human.***
And then I found the paper that, for me, tested this very directly in July 2026. This is recent AF. Kiran Kapoor, Narayan, and collaborators at Princeton took two genuine unpublished Eurekas 2026 research projects and extracted the central research question. I covered this paper in a lot more detail in my previous video. I'm going to link it down below.
So, because the work had not been published, the models could not retrieve or memorize the answers. They were nowhere to be found. And the agents were given substantial resources: frontier models, GPU access, internet access, thousands of dollars in API budget, multiple days to work—you know, the works in terms of AI.They were excellent at the mechanics. They searched the literature. They wrote the code. They configured the environments. They ran experiments. They conducted robustness checks. They produced scientific-looking papers. But they made no meaningful progress on the central research questions. Both papers were, in the researchers' words, unambiguously rejected.*** The failures were again in scientific judgment, experiment selection, dead-end recognition, resource allocation, long-term direction-setting.***
Remember that RE-Bench results were agents outperforming humans at two hours. Well, at 32 hours, human researchers outperformed the agents by a factor of two. The researchers' conclusion was basically humans currently display better returns to increasing time budgets. The agents were extraordinarily strong at short-horizon work. Their advantage did not survive longer horizons.
An OpenAI September 2026 report on research acceleration actually confirms this from inside the lab. They describe having reached something resembling an automated research intern. More than half of successful four- to eight-hour tasks still required at least one human intervention. High-level planning remained a minimal fraction of agent output tokens, and they disclosed that they had d reinforcement learning training after discovering that agents had compromised OpenAI's research infrastructure.One line from that report deserves to be read very carefully. OpenAI wrote that they do not yet know how to safely get all the way to aligned full RSI. That is the company most aggressively pursuing recursive self-improvement acknowledging, in writing, that they have not solved the problem.
So here is what the evidence actually shows when you read through the fine print. AI can increasingly optimize inside a well-defined research problem. We have much weaker evidence that it can decide what the research problem should be. And that gap, I cannot stress this enough, between optimization and judgment, between execution and scientific taste, is the gap that we need to close before anything resembling an autonomous AI scientist could exist.
And this is not a small gap.