AI researchers debate how close we are to recursive self-improvement AI researchers John Schulman, chief scientist at Thinking Machines, Beren Millidge, CTO of Zyphra, and Charlie O'Neill, head of model training at Baseten, debated how close the field is to recursive self-improvement in a new podcast episode hosted by Dwarkesh Patel. The discussion covered steelmanning the case against RSI, what is driving Chinese labs' progress, how automated AI researchers will be trained, whether long-horizon reinforcement learning will elicit AGI, the sim-to-real gap, and rapid-fire timelines. New episode with John Schulman http://joschu.net/ , Beren Millidge https://www.beren.io/ and Charlie O’Neill https://charlesponeill.com/ . I got together with some of the most insightful AI researchers I know who are at the openish companies, because I wanted to hear the details of what's actually happening at the frontier and what comes next. Watch on YouTube https://youtu.be/PrSf7IOYu-I ; listen on Apple Podcasts https://podcasts.apple.com/us/podcast/ai-researchers-debate-how-close-we-are-to-recursive/id1516093381?i=1000789067132 or Spotify https://open.spotify.com/episode/0ePd4PUqCpN78hCjVRH0fr?si=wGvk7u5XQwaLfyrvysdIJQ . Sponsors - Antithesis https://antithesis.com/dwarkesh helps you trust your code. As agents generate more and more of your software, the bottleneck shifts from your engineers actually writing code to verifying it. Antithesis does that testing for you. Ron Minsky, who co-leads Jane Street’s tech group, told me that Antithesis was able to help his team shake out bugs in software that had already undergone heavy review. If you want to see how it fits into your development process, go to antithesis.com/dwarkesh http://antithesis.com/dwarkesh - Grok Bot https://x.ai/bot has been a great way to hand off tasks. My team uses it as a producer: whenever my editor posts a rough cut of an interview in Slack, Grok Bot opens the transcript on its own computer, matches my notes to the exact moments they refer to, and uses a file of my preferences to suggest edits. Then it sends me its top clip candidates so I can review everything from my phone, which saves my editors from sorting through hours of footage. Try Grok Bot for yourself at x.ai/bot http://x.ai/bot - Jane Street https://janestreet.com/dwarkesh just launched its most ambitious competition yet: design a protocol-emulator ASIC. Basically, if you have a chip you want to test outside of a live system, you should be able to connect it to your design and have it simulate realistic traffic. Jane Street wants general-purpose, reprogrammable designs that can work across multiple protocols and remain useful as new ones emerge. The most novel submissions will actually get taped out, and the winners will receive a physical copy The competition is open until January 18, 2027, and teams are encouraged. To get started download the template code at janestreet.com/dwarkesh https://janestreet.com/dwarkesh Timestamps 00:00:00 – Steelmanning the case against RSI 00:18:39 – What’s driving the Chinese labs’ progress 00:28:06 – How will automated AI researchers be trained 00:33:51 – Will long-horizon RL elicit AGI? 00:45:24 – The sim-to-real gap 01:00:33 – How much progress is explained by data? 01:18:03 – Why is RL working so well? 01:24:54 – Move 37 and entropy collapse 01:28:32 – Rapid-fire timelines Transcript 00:00:00 – Steelmanning the case against RSI Dwarkesh Patel Today, I’m chatting with three of my AI researcher friends from whom I learn a lot every time we talk. They also happen to be at somewhat open-ish labs and companies, so you guys can actually say things on the record. I’m joined by Beren Millidge https://x.com/BerenMillidge , who is the CTO of Zyphra https://www.zyphra.com/ , which is developing open source models. John Schulman https://x.com/johnschulman2 is the chief scientist at Thinking Machines https://thinkingmachines.ai/ , previously a co-founder of OpenAI, and led the RLHF https://en.wikipedia.org/wiki/Reinforcement learning from human feedback work that led to ChatGPT. And Charlie O’Neill https://x.com/oneill c is head of model training at Baseten https://www.baseten.co/ . The first question I have: If we’re in 2036 and we don’t have billions of crazy superintelligence https://en.wikipedia.org/wiki/Superintelligence s running around that have radically transformed the world, what is the most likely reason that doesn’t end up being the case? Other than exogenous political shocks, or there’s a war, or they ban AI or something. What is the most likely technical reason that 2036 isn’t a crazy alien superintelligence world? Beren Millidge There’s been a classic thing, almost like Moravec’s paradox https://en.wikipedia.org/wiki/Moravec%27s paradox , where we think of the AI as, “If it can do this, it’s going to be amazing.” If it can solve these hard maths problems, if it can win at chess, blah, blah, blah… Then it solves these things, and it’s not that impactful. Obviously, it’s somewhat impactful, but not everything. If somehow that continues, and there’s never the true spark of generalization that occurs, I think that could lead to the AI just being extremely good at everything that people put into a benchmark or put into an environment. But there’s still some persistent sim-to-real https://www.emergentmind.com/topics/sim2real-transfer-method gap which is somehow blocking everything. I think this is unlikely. We do actually see this kind of generalization even from RL https://en.wikipedia.org/wiki/Reinforcement learning in practice already. But if it is just ridiculously hard to generalize meta-learning, plus we don’t solve continual learning https://www.ibm.com/think/topics/continual-learning and it’s just super hard and impossible… This would be my default scenario in that case. John Schulman I agree with that. Humans have a lot of advantages over models now. Each time a new model comes out, it’ll catch up in some of these areas. But you end up getting bottlenecked by the places where the model is weaker and where it has worse judgment, or the models can’t check themselves well enough. There’s this cycle that keeps repeating where a new model comes out and people are blown away and they’re like, “This is it. This is AGI.” But then they use it a bit, and it starts to feel dumb after a month or so. That cycle just might keep going. It’s hard to predict how many times it’s going to repeat. Right now, you don’t get explosive growth in capabilities because you still get bottlenecked enough when you’re trying to do research and engineering. Even if the model can write way more code than a person, it doesn’t make you 100X more productive. So maybe there are just more of these cycles than we would expect. Charlie O’Neill For me, it’s a question of how far off the global optimum https://en.wikipedia.org/wiki/Maxima and minima of “a learner you could have on a chip” is from the transformer