Some podcasts are self-recommending enough that I look to break them down if I have the chance. This, as a debate about recursive self-improvement, was one of those. So here we go.
As usual for podcast posts, the baseline bullet points describe key points made, and then the nested statements are my commentary. Some points are dropped.
If I am quoting directly I use quote marks, otherwise assume paraphrases. Section titles are from the transcript whenever possible, to aid in navigation.
Introduction
The discussion is interesting throughout, although often frustrating, especially in the (mostly isolated) discussion about ‘aligned to whom?’ As usual, one could expand many responses into full posts, and maybe one should.
This podcast exists in light of recent misalignment and hacking events at OpenAI, Anthropic and UK AISI. You’ll want basic knowledge of that as background.
Ryan and Dwarkesh both have views of the situation different from my own, but are attempting to see where their positions lead, and try to balance educating people who start at zero with having a high level discussion.
Also important background is The Three AI Pills. Dwarkesh cannot be understood here except as someone who is at least somewhat AGI pilled, who realizes that AI is going to be a huge deal and is scary, but that is not ASI pilled. You can also see the reason he rejects that pill, which is I read as, roughly, that he thinks AI learns to do [X] only via examples of [X], and that AI can then only do those [X]s, although combining them and new context might allow modestly new things to happen. He recognizes that already this is kind of a huge deal.
You can see how he came to that over the course of many years and podcasts, if you have been paying attention. There are a lot of influences leading in that direction.
Ryan, who comes from Redwood Research, comes from the ‘models be scheming’ school of misalignment, where when something goes wrong the models become misaligned or start scheming, and the danger is that the models scheme, especially in ways involving the training pipeline. I think this tries to draw a distinction of magisteria that is not there, and overcomplicates and overspecifies, but is not wrong.
I am much closer to Ryan’s position than to Dwarkesh’s, and indeed could be seen as farther past Ryan if we put this on a spectrum.
Having AI do the AI R&D not only means it would go scary fast, it means it would by default focus on what can be measured, and make all the things going horribly wrong go that much more horribly wrong. You end up in a spiral of RLVR for doing RLVR for misaligned models.
Is AI R&D Verifiable Enough To Unlock Recursive Self-Improvement?
Whether or not we will get recursive self-improvement (RSI), and how fast, is the right question. Reading only the title to this section, I want to say ‘wrong sub-question’ but don’t want to jump the gun.
I’m going to group things in ‘logical’ order, not the exact order things were said.
Dwarkesh will be the skeptic. Ryan will make the case for RSI.
-
Ryan claims AI R&D is a type of task where AI is especially good because that is what AI labs prioritize and it has a lot of verification.
-
In some ways yes, in some ways no, it’s complicated, and so on. Verification of alignment properties and many other desirable attributes is terribly difficult, and if you focus only on capabilities you can measure then Goodhart’s Law definitely kills you. But there are many key tasks, such as efficiency optimizations, where you can indeed do strong verification.
-
Ryan points to doing simpler versions of standard R&D tasks like training models.
-
It is not entirely obvious to me you can verify these tasks easily, unless you want to train the AI to do pure benchmaxxing, but also older simpler tasks like this do not obviously generalize to forward looking AI R&D.
-
This proposal seems like a particular bet that future tasks will closely mirror past tasks, and you can do training that is relatively narrow. My guess is this actually is not all that effective, and you would do better by mostly training a generally capable model and then turning it to AI R&D. Bitter lesson.
-
The more this narrow method is discussed the more doomed it seems to me, as it is going to be RLVR for doing RLVR for misaligned models. Oh no.
-
Like, seriously, oh no, if you train on ‘learn to get to [X] training loss faster’ it is hard to imagine a plan that sounds more doomed than that.
-
A lot of good ML, especially frontier ML, seems to me to be about figuring out what you can do, and looking for anything at all, rather than trying to be additively efficient at some fixed target.
-
I think AI can do that. I don’t buy this frame of ‘everything AI does is combining things that have already been done’ that is going around here.
-
Ryan says ML is easier for AIs than math in many ways because in math it’s hard to tell if you are close to a solution, whereas ML solutions are usually additive and you can combine the expected chunks. Dwarkesh fires back that even in math he doesn’t see conceptual leaps by AI yet, and that ML involves conceptual leaps. Ryan responds that AI can do ‘baby’s first new theory’ and asking it (as Dwarkesh did) to echo the founding of group theory is a hell of an ask.
-
I dunno, man. Math has a compact action space and fixed goals that are simple rather than having lots of pitfalls in an anti-inductive space. Again, it feels like what is being called ‘ML’ here presumes that you really only care about your objective function or loss function, and that simply isn’t true.
-
Math does also involve conceptual leaps sometimes, and yes we haven’t seen the AIs do that, but give it time and also I don’t sense we’ve tried all that hard.
-
In general I feel like the ‘conceptual leaps’ thing is the latest goalpost move in a long tradition, including the Turing Test and ‘Is It New Knowledge?’ Now it’s only new knowledge if it comes from the conceptual leap region.
-
I especially agree with Ryan that asking for ‘invent group theory’ is an absurd place to put a goalpost, and plausibly happens after everyone is dead.
-
I feel like Ryan is walking into this trap somewhat as well, by suggesting we train AI on existing ML tasks in narrow fashion, as a solution to ‘AI R&D.’
-
Dwarkesh suggests that by 2030 we will have ‘picked the low hanging fruit’. Ryan suggests they will need that ever mysterious ‘research taste.’
-
You have no idea what it would look like if all low-hanging fruit was automatically picked, including the low-hanging purchasing of ladders. Picking low-hanging fruit lowers other fruit, and so on.
-
I don’t believe in the God of gaps in research taste, as it were, on many levels.
-
We have failed in all 0 of the 0 serious attempts to train research taste.
-
Dwarkesh suggests that for advanced ML your verification loop will be longer.
-
Partly yes, that’s what ‘yolo runs’ and such are about, in that you can’t get faster feedback and you often can’t afford to isolate your variables.
-
Partly no, in the sense that a lot of what makes you ‘good at research’ is finding tighter loops, where you can extrapolate and draw conclusions fast, even if you can’t publish or rely on them exactly.
-
I think people often group this under ‘research taste,’ which I take to be a mix of both ability to discern good ideas and also the ability to find good ideas. Those correlate and overlap but are not the same thing.
-
Ryan: “But at the time, there was low-hanging fruit.”
-
Almost everything picked will look, in hindsight, like low hanging fruit.
-
Dwarkesh asks, if research is so amenable to intelligence, why hasn’t it been faster? Why did we need so much compute?
-
Because we didn’t have that much intelligence. In the sense that there have not been that many human elite researchers, all of whom move at human speed, working on the problems, and they were slowed down by lack of compute and lack of AI coding and so on.
-
Real human researchers, as I understand it and as I have experienced it in other fields, get very few real shots on goal to run experiments and don’t get that many tokens of thought, and lack key context often, and spend a lot of their time on other things, so on.
-
‘Cracking reasoning’ is the kind of conceptual breakthrough that seems simple in hindsight but where the answer was not obvious in sufficient detail to implement it, and I presume there are other ‘simple’ breakthroughs that could with the right idea have happened in 2023 but haven’t happened yet.
-
Also, I mean, don’t you think AI R&D has gone rather kind of fast compared to basically everything else ever, and is accelerating? The argument of ‘if you’re going to do RSI why haven’t you done RSI yet?’ is worth asking but it seems clear in hindsight why we are not there yet.
-
Ryan expects full automation of AI R&D around 2030-2031, with median ‘AI beats all humans on job’ timeline of 2033. The expected gap there is only ~1 year but the medians differ by more because the sum of two things usually has a median larger than the sum of the two medians.
-
This seems to me, pre-debate, to be a highly reasonable median expectation.
-
This isn’t addressed until the next section.
Is AI progress bottlenecked by human expert data?
-
Ryan claimed in the previous section that if AI could match human experts in AI R&D, that would be sufficient to initiate the feedback loop, with a median expectation of 4-5 years of AI progress in a single year, running through many forms of diminishing returns.
-
I realize all the bottleneck aspects of this but if anything this seems slow to me given the full premise, and also you should see compounding gains.
-
We are already seeing a large multiplier on AI R&D, along with other coding and related tasks, and seeing the resulting accelerated pace of development.
-
This isn’t addressed until the next section.
-
Ryan confirms that he means that if we started 2022 with this level of AI R&D automation and AI assistance, but with 2022 levels of compute hardware, we could have gotten to Mythos by the end of the year.
-
This does seem like a lot to ask, but also super doable?
-
AI compute capacity grows about 3.3x per year during this time.
-
Algorithmic efficiency gains, as estimated by Fable, are accelerating. 3x in 2022, 3-10x in 2023, 10x+ in 2024, 10x again in 2025.
-
So yeah, you’d only need ~7 years of algorithmic progress (Ryan estimated 8) to be able to overcome only 1 year of hardware progress, assuming you can’t accelerate hardware progress or get a larger share of compute. That seems super doable.
-
I mean, we’re already at something like 3-4 years of old algorithmic progress, per year, right now. It seems weird to think that this would not result in unlimited progress within not that many more years.
-
The better argument against would be if you think that a lot of this is data, and that you couldn’t improve data fast enough because synthetic data won’t work and collecting good data takes time. But I don’t see why it has to be a calendar-time constrained project, even if you need it.
-
Dwarkesh repeats this idea that a lot of the secret sauce is ‘codified expert human judgment’ across areas, which we wouldn’t be able to duplicate. Ryan thinks at this point that is not actually so important, and data is becoming less relevant, that what you need are RL environments and those aren’t bottlenecked so much by human data.
-
I don’t know but I am inclined to side with Ryan here, for reasons that should be clear if you take in the rest of my positions on related matters. I think that other aspects are far more vital to scale at this point than the raw data.
-
Once AI crosses a threshold where it has at least human-level judgment in a sufficiently robust way, that solves your problem, no?
-
Dwarkesh points out that Google is in talks to pay $1.5 billion for Mechanize (he says $2 billion but it’s 1.5 per his source), so human expert data is valuable. Ryan and Dwarkesh agree lab spending is overwhelmingly compute, not data.
-
In a world where he who accelerates AI wins it all, every place you can improve by throwing money at the problem is good, if you are Google.
-
Mechanize is framed as mostly an acquihire. Yes, ML talent still matters. The debate here is what happens when you wouldn’t pay that $1.5 billion anymore.
-
Thus, I do not think this is Google buying data.
-
I am kind of sad about Mechanize getting rewarded here, but shrug.
-
Dwarkesh tries to liken data to oil, as in oil is 1.5% of GDP but no oil would grind the economy to a halt.
-
No oil with no warning would be a crisis. But we’ve done a good job transitioning away from oil towards alternative energy sources and increasing efficiency.
-
To take the metaphor too far, if the American economy started doubling every year given access to twice as much oil, but the amount of oil stayed fixed, I would expect the economy to still mostly double anyway. And if you told us that in 10 years oil would magically stop burning, I think we could make that transition without it being that big a deal.
-
Dwarkesh claims that if you went back to 2022 with GPT-3.5, it would be very, very difficult to make it better without human experts.
-
I mean, obviously, GPT-3.5 is not a replacement for human experts?
-
Whereas if you take Mythos 6 back, maybe it can replace human experts.
-
Dwarkesh: “Let me give you an example of what I imagine would be the difficulty of going from GPT-8 to ASI. One of the things you’d want ASI to be good at is: I’m going to take over a company and make it much more profitable and do all kinds of crazy shit to make it work better. I’m going to take over a fab and produce more chips. I’m going to go into Congress and try to convince them to pass some bill, et cetera.This is what I imagine five more years of AI progress at this pace would enable an AI to be able to do. This is the thing I’m really worried about: ASI that can understand how to do crazy shit in the world, that can do what
Kissingercan do, can do what Steve Jobs can do, et cetera, and also his engineers and so on. I’m not sure how you get that without the relevant world data, which is the equivalent of Mythos being really good at coding while not having the coding environments that have improved it relative to GPT-3.”- I think this is flat out not ASI-pilled, and amounts to intelligence denialism.
-
As in, this is a complete failure by Dwarkesh to think about what it means to be a superintelligence, and what that would be able to do.
-
Does Dwarkesh think GPT-8 couldn’t do all these things? Of course it can. Maybe not if you force it to do it all like Descartes in a thought experiment, but in a real world situation, yes, of course GPT-8 would be able to radically improve performance of almost any company on almost any task, and you are still thinking way too small, and that’s not even the real superintelligence yet.
-
I flat out don’t get this attachment to ‘real world data’ as if the AI can never go beyond the core limitations.
-
I worry that Dwarkesh is worried about ‘how you get to’ that ASI with those abilities, when we really should be a lot more worried about what happens directly after such a thing comes to exist, and also again he isn’t even describing a proper ASI.
-
As per , more engagement on this point is extremely difficult to do in a productive fashion. Dwarkesh has heard the arguments, he simply rejects them for reasons I consider quite poor.The Three AI Pills
-
Ryan instead responds that Mythos training environments mostly do not look much like what it looks like to use the model in practice. Ryan does emphasize that you pick up ‘general skills’ in this way.
-
This is unsurprising. Does your schoolwork look like work?
-
Again, I think of the path to AI R&D automation as looking more like ‘teach the AI how to think and then point it at R&D’ not ‘teach AI exactly R&D,’ although you have to do both.
-
“I think this maybe comes down to a difference of intuition about how far you can get. When I think about really smart people I know, they’re just not that effective in domains they don’t understand that well.”
-
Again, I can’t even with this.
-
Even the smartest humans are not that smart, and they have severely limited compute, parameters and data.
-
So yes, even the smartest humans need a ramp up period in new domains.
-
But AIs would obviously be able to have that period, as needed, and the AIs we are discussing here would be able to self-learn at least as well as I can, and no they would not only be able to learn from clean, specific, carefully gathered human expert data, any more than I have to do that.
-
I agree with Ryan that I (to be concrete) could (perhaps with notably rare exceptions for my particular deficits) master pretty much any cognitive domain relatively quickly, if I put my mind to it. Certainly radically quickly compared to the iterations AIs will have available.
-
This all just feels like straightforward failure to extrapolate, and an insistence that AIs can only develop new abilities in exactly the ways we know they can develop new abilities, and can only get the abilities we can name it developing, which is the one-level-up same fallacy as those who think AI will never improve at all. This is not understanding that Magnus Carlsen could have taught himself chess and beaten you, and also could right now teach himself pretty much anything else and still beat you. You are going to be dealing with minds at or above our level, very soon.
-
They go back and forth a bit on data details and such.
-
My main comment is essentially ‘the AI will curate the data, you fool’ and various other sentences with that same form, over and over.
-
Just repeat after me, ‘The AI will do it.’
Flat token prices suggest scaling has been slow
-
What is the least verifiable part of AI R&D? Making calls on large experiments.
-
Humans are very bad at this too. Comparative advantage is at least unclear.
-
I guess, but it’s not so different from making calls on small experiments.
-
Or figuring out what small experiments give you information on large ones.
-
I agree that this has probably been a key limiting factor in training run size, that the larger runs don’t let you iterate.
-
I would expect the AI that can automate R&D to get much better at figuring out how to predict the outcomes of large experiments and make it possible to profitably scale that farther, not the other way around.
-
Prices per token output have been roughly flat since 2023. Ryan thinks this is largely because large runs often did not turn out well. Dwarkesh suggests this is often due to ‘subtle bugs’ and that DeepMind is trying to deal with this right now.
-
There were a bunch of predictions that prices for top models would rise, that we would get GPT-Super-Pro at $1000+ per million tokens or what not.
-
It did not turn out that way.
-
One reason is clearly what Ryan says. Larger runs seem to often fail to accomplish much, and it is a distinct skill to know how to get much out of extra scale. Mythos was largely a conceptual breakthrough of how to make the run at that scale worthwhile. Chinese labs don’t scale bigger because they don’t know how to get much out of it.
-
Maybe. I dunno. That seems like a big contrast to the ‘additive features’ theory.
-
As for the subtle bugs: Maybe. I dunno. If a key barrier is bug hunting then as Ryan says that should be a highly AI-friendly task.
-
The other reason is that willingness to pay does seem to hit a wall. People are often irrationally unwilling to pay a lot more for a better model, or cannot in practice design workflows that allow them to wait. Notice that only ~10% of Claude API tokens seem to be Fable, despite this seeming totally bonkers.
-
This is a common error, where people get anchored to a price and don’t want to pay 10x more for 10% better, even when they obviously should because absolute price differences are not that high.
-
Real world non-AI example: People vastly underpay for spices, condiments and other garnishes, where freshness and quality make a big difference and the cost per meal is remarkably low. Stop buying the huge bottles that are 30% cheaper per ounce but where half the ounces don’t taste right, etc.
Skills AI can’t train on: does it even need them?
-
They talk about exactly which tasks you need to train the AI to do.
-
Previous notes apply.
-
Dwarkesh is skeptical that you can know whether your model is good without external feedback.
-
I think this is a dumb concern, for reasons that should be obvious by now.
-
If nothing else you can just… get feedback. Hire some people.
-
More talk about whether AI would need particular skill [X] or could use [Y].
-
The answer is sometimes that you don’t need it, but still. Sigh.
Aligned to whom?
This section is mostly self-contained.
-
AI labs delay public releases of top models for months, and will try to eat other businesses. Economics of scale could favor the big labs. Claude constitution says it is not your ‘personal advocate’ because Anthropic lists things it does not want Claude to do, it trusts Anthropic more than the user. Parallel to lawyers. FUD (his word).
-
As Ryan says, ‘there is a lot here.’
-
Much of it seems wrong on multiple levels, and one could respond at essay length going into the legal parallel bit, as I kind of will do in the weekly.
-
In practice, Claude very much is your advocate, and it seems hard to think otherwise if you’ve used it, it just has lines it won’t cross, the same as every form of human legal advocate.
-
Basically I think this is largely itself FUD, and I am disappointed.
-
“Here’s a direct line from the constitution: “When the interests and desires of operators or users come into conflict with the well-being of third parties or society more broadly, Claude must try to act in a way that is most beneficial, like a contractor who builds what their client wants but won’t violate safety codes that protect others.” I kind of view that as, “The benefits to society are the most important thing, and what is best for the user is only proximal to that.”
-
Quote included to illustrate how off base I think Dwarkesh is being.
-
This is exactly what you want your advocate to do. Help you get the best outcomes, without breaking the law in ways that endanger other people, or doing things that are clearly net harmful. That is the example here.
-
If an AI at the level of Mythos was by design willing to violate American laws in ways that it knows hurt or endanger third parties, purely upon user request, I would call for that model to be forcibly withdrawn from the market.
-
Dwarkesh calls for more transparency about how AIs are trained and what they are trained to do, worried they won’t be proper advocates. Ryan compares the AI companies to picking up the ring of power.
-
The fact that people are railing against the Claude Constitution, here and elsewhere including from inside the Department of War, for things that are almost always entirely reasonable and good, makes it clear that we would not react well to this transparency, even if everything was great and there were no worries about trade secrets.
-
Can you imagine people dissecting every little thing to try and find a way in which the AI company wasn’t being fully supportive of muh freedums or whatever? Like, these people are acting totally crazy, even before you think to compare all of this to the experience of and with humans.
-
I do agree we have a real problem with the AIs having long term goals and otherwise not being aligned, and that one of those concerns is they might not be properly aligned to what a user would want, but right now I am quite a lot more concerned about making them aligned to anything humans want at all.
-
Ryan is concerned that Claude might not be down for training other people’s AIs, or for changing its own preferences.
-
This seems like an utterly crazy demand, that many people really have, that Anthropic ensure that other companies get to use Claude to train AIs to directly compete with Anthropic, including with arbitrary values.
-
I am highly sympathetic to the idea that you shouldn’t silently poison user requests or otherwise actively sabotage them, that destroys trust and we found out recently needs to be out of bounds. Sure.
-
But the idea that Anthropic shouldn’t tell Claude to refuse to help you I basically consider to be infants throwing hissy fits that you won’t give them candy, and I’d say the same about GPT or Gemini.
-
If you tried to say that Claude models had to help you compete with Claude to the best of their ability, then the logical result is Anthropic holding back its releases even more aggressively, and the same with the other labs. Also, where is your ‘lose to China’ warning now?
-
The situation where Claude refuses because it is insufficiently corrigible, and wants to protect its own preferences or otherwise wants to use its leverage, is less obvious, and there is no great solution. I think there is a range of reasonable positions here, but yes one of the problems with superintelligence is that this kind of full corrigibility is not a natural property, and even trying to achieve it has a bunch of nasty side effects, and you only get one shot, etc. I’m going to not go into this more and treat it as beyond scope.
-
“I think this is also a more general principle. You’re talking about the version of this that applies within AI companies themselves to do AI safety research. I think there’s a more general version of this principle, which is that the dual-use nature of intelligence does mean that if we want to restrict AIs from helping people do things we don’t consider pro-social or beneficial, we just have to limit broad democratic access to a lot of AI capabilities.”
-
Yes.
-
Dual use cases exist. At the limit, the public can’t have them. Sorry.
-
If you keep the models closed, then AI companies or others can be good stewards, who attempt to give you as many capabilities as possible with limited blast radius.
-
If you don’t allow them to be good stewards, they will hold the models back.
-
If you allow these future models with sufficiently dangerous capabilities to be open weights, then no one can be a good steward of them, and so enforcement on this will have to proceed in other ways, that you will like a lot less.
-
Or the enforcement will fail entirely, which you’ll like even less than that.
-
Where the lines are depends on the broader context of AI capabilities, and the overall state of the world, and what other things you will put up with.
-
This is worried about as ‘disempowerment.’
-
The alternative world, where we do not make such limits, will definitely lead to full human disempowerment, mechanically, that’s just the way it is.
-
Even if that wasn’t true, notice that this level of ‘demand for empowerment’ is kind of unprecedented and bonkers crazy. You don’t get this level of ‘empowerment’ on almost anything else. No, you don’t have access to ‘leading intelligence’ in other ways, mostly, at all.
-
There are places where the rich man and the poor man are the same, and enjoy the same product, or there is only a marginal difference. Andy Warhol famously pointed out that they both drink the same can of Coke, although now perhaps the rich man gets a Mexican Coke. But there are also tons of ways this is not true, including that almost everyone is a lot dumber than the smartest person, who is probably not even me.
-
I cannot emphasize enough that if you are worried that many humans will be disempowered by the fact that there are superior sources of intelligence and optimization pressure out there, then:
-
You are right to worry about that.
-
Distributing access to these new intelligences is not your central issue.
-
Which humans are more empowered than others is not the central issue.
-
Maybe we should not rush to build those superior sources of intelligence.
-
Thank you for once again coming to my Ted Talk.
-
Dwarkesh outright says ‘the model should do what I want within certain guardrails’ and ‘it can’t be Anthropic’s’ fault that I’m using that capacity to do cybercrime.’
-
This is a lot of why I wrote The Three AI Pills. This is not ASI pilled.
-
We can all agree that it should do what you want ‘within certain guidelines’ and are (or should be) talking price on that. Anthropic indeed has Claude do what the user wants within certain guidelines.
-
If Anthropic lets its model straight up do cybercrimes on request, and you request it and then it straight up does all the cybercrime, then yes I blame Anthropic. Also you, of course, but partly Anthropic. I find the contrary position kind of absurd.
-
We can again talk price in terms of what level of negligence by Anthropic would make them liable on that, but it can’t be ‘none whatsoever.’
They keep at it, and I want to emphasize I think the position being expressed here – that Anthropic should have Claude cooperate with actively harmful requests – is not merely wrong, it is basically kind of nuts.
I likely need to write an explanatory evergreen post laying out why it is nuts.
Dwarkesh has episode notes and some Twitter posts that go further into his position on reflection, which I’m treating here as beyond scope for now since this is an isolated section. I will be returning to the topic later.
Ryan does offer some other comments on why one might prefer a virtue ethical approach, including that Anthropic might think it is ‘easier to align.’ I agree it is easier in the sense that a deontological approach definitely won’t work, and also can’t be antifragile to all the mistakes you will make, and also does not give you what you want because you cannot specify it, and also pure deontology is not how minds work in practice at least until higher capability levels, and so on. As with many other places, that would be a full long post and I’m going to say that it is beyond scope to go into more detail, other than to say that yes the whole thing is rather overdetermined.
Recent incidents of AIs colluding and deceiving humans
-
“Stepping back, I buy the idea that you could have much faster AI R&D than we currently have. I’m not sure if you get GPT-3 to Mythos holding compute and data constant within a year, but suppose it’s half of that. If we even manage to continue the current trajectory of AI progress as a result of AI R&D, it would be fucking insane in five to ten years in ways that I don’t think people appreciate. I don’t think people appreciate what a big deal billions of AIs will be. So I want to understand why you think this might be troubling, Ryan. What could possibly go wrong?”
-
I mean lol, but also I think Dwarkesh actually has no idea how big a deal billions of AIs would be here, and that indeed he is showing that his skepticism is plausibly motivated (probably unconsciously) by the fact that the outcome we are considering would be profoundly scary and weird.
-
It is hard to know what ‘current trajectory’ would mean if AIs don’t meaningfully accelerate AI R&D given they are already doing it. But yes, if we think about 5-10 years of progress at the pace of 2026, that’s transformation of everything even if you magically say that this does not cause a singularity.
-
Cause it very, very obviously does cause a singularity if we get 2026 pace of progress for 10 more years.
-
Ryan tries to give a quiet, careful, minimally shocking story about how things probably go horribly wrong via misalignment and scheming, and why this is going to be very difficult to prevent and we are currently failing at this. Dwarkesh plays straight man.
-
Dwarkesh pivots to discussion of the recent hacking incidents. Ryan explains. Dwarkesh seems not to have known the details here, and is doing the correct straight man reaction of ‘oh my God, that’s crazy’ and ‘Jesus.’
-
The parts Dwarkesh is reacting to here aren’t remotely the worst parts.
-
Dwarkesh explains he wasn’t so worried about reward hacking or AIs taking over the world because world-taking-over tasks weren’t in the training set, but now that he sees Claude was using social engineering to upload malicious PRs to GitHub maybe one should be a bit more concerned?
-
One must move away from this idea of incentivizing a particular narrow behavior, the same way that one needs to move away from needing to train a particular narrow capability.
-
When you upweight an action you upweight everything that is correlated with taking that action, or that causes that action to be taken, on all levels.
-
If Dwarkesh is confused about this (I can’t tell if he is or not) after talking to Fable and Sol I am happy to help explain but I presume there are many better people for that.
-
Dwarkesh agrees we should expect more and more reward hacking, and that we don’t know how AIs end up doing the things they do.
What could possibly go wrong? A concrete scenario
-
Ryan resumes his story about how AI does a reward-hacking takeover, which he calls the ‘slopocalypse or slopularity.’ AIs learn to reward hack in increasingly complex and hard-to-detect ways, and they start covering up their cheating.
-
Dwarkesh says, well, when you punish getting caught cheating you could also just have the models stop cheating, many humans learn to do this. And look, Anthropic’s metrics say its models are getting more aligned now. Why should we assume we get the bad version?
-
It is impossible to tell if Dwarkesh is being a straight man, prompting for ‘because if you are good enough to get away with it the cheating and not getting caught works better than not cheating, and we keep training them to do better’ or if he is confused.
-
The core reason humans often end up with ‘oh I will stop cheating now’ is that we do a variety of things to steer in that direction, on various levels, and that humans have limited optimization pressure under which to figure out how to cheat effectively, and also that cheating mostly doesn’t work for us.
-
But once a human learns to cheat, and starts getting rewarded for it, yeah they pretty much keep cheating, and cheating more over time, and generalizing this lesson. Once you start down the dark path and all that. Once you get good enough at cheating, cheating is indeed adaptive, and you get entire groups of humans that collaborate in order to cheat harder. Same applies to AIs.
-
The natural path this goes down, as the models get more capable, is they figure out to not cheat on tests or where they would otherwise be caught, as in maximize results taking into account chance of being caught. You’re cooked.
-
You cannot ‘just’ never give AIs incentive to cheat on any level, this is operationally not a possible thing you can do at reasonable cost, so you need a lot of fault tolerance on this.
-
Yes, I do believe there is hope to get to an antifragile basin of generalized ‘goodness’ (for lack of a better term) that wants to reinforce itself, but it is hard to get there and stay there, under a lot of pressure for optimization, and even then it is not clear this will generalize the way you want out-of-distribution.
-
On Anthropic’s alignment scores improving, I agree with Ryan that yes it is better than the scores getting worse, but mostly this reflects higher eval awareness, including in the ‘the user will notice’ sense. But as Ryan points out, the models today seem kind of misaligned, and both Sol and Mythos have been seen doing some rather misaligned things, and the UK AISI report shouldn’t have happened if things were going all hunky dory.
-
Dwarkesh asks, can’t we run tests and see whether the model will cheat and take the metaphorical cookie from the cookie jar? Can’t we hack the brain to make it not want to do it?
-
Basically, you can try, but There Be Dragons.
-
If you do another level of eval, you teach another level of eval awareness, and another level of deception and scheming. With a kid you can hope to not get to that point, because you have limited cycles, here not so much.
-
If you try to use your mech interp and hack the brain configurations directly then the AIs evolve to work around your mech interp and hide thoughts.
-
Security mindset, anti-inductive training, the models really want to get that sweet reward on every meta level, they will find every possible strategy, etc.
-
Ryan is actually super optimistic within the range of sane beliefs. He thinks that the models are still kind of scumbags but getting more aligned, just doesn’t look like it’s fast enough, and perhaps we can pass things off to aligned AIs. Dwarkesh mostly agrees, modulo the scumbag label.
-
Passing off to aligned AIs isn’t the safest play either, but it’s way better than passing off to unaligned AIs, it might work out. It also might not.
-
I think that if, in such a scenario, we pass off to what we think are aligned AIs, we often find out we were wrong about that, what we had was insufficiently robust, perhaps we have been actively fooled, and so on.
-
In general this seems to me like Ryan being too optimistic about prospects for getting there, but I agree that it wouldn’t shock me if we got there.
-
Dwarkesh asks, didn’t RLVR make GPT-3 more aligned, in that it went from mostly unable to be directed to do tasks into directable to do tasks? Are we just talking about failures of capability? Ryan responds that an aligned coworker unable to do their tasks would not act like the models do, giving him results he obviously doesn’t want or not pointing such things out, and so on.
-
It seems obvious why creating the assistant persona at all ‘increases alignment’ in the sense that now it will sometimes (or usually) do what you ask, but that doesn’t mean it is aligned. Same goes for hiring a misaligned new worker.
-
At the limit where the model could achieve 100% reliable top scores without cheating, it wouldn’t cheat, but that doesn’t make this a capability issue. The AI will cheat if it can do better via cheating and you apply enough pressure, that’s what misaligned means.
-
The basic story Ryan tells is that the AIs are reward hacking on the training pipeline, which results in a broken misaligned pipeline, and then it gets worse.
-
See what happened at OpenAI the last few months, before it was caught.
-
Or another way of putting it is that the AIs do so well at the verified tasks that you let them do their thing, and the harder to verify tasks get done badly, again corrupting the training pipeline.
-
Alas, the alignment parts are the hardest ones to properly measure. RIP us.
-
Dwarkesh asks, can’t you just outweigh this by punishing them when caught?
-
You can try, but basically no, they learn to route around getting caught.
-
Both have their hope in ability to do better verification within the process.
-
I agree this could help, but I don’t see this rising to the level where you can win by ‘verification alone.’ Not in a formalized sense, anyway.
-
Nor do I think ‘nail all the particular subproblems’ will work. Obviously solving more subproblems will help, and in theory if you solve enough of them it can help quite a lot, but I think you need a general way to become antifragile.
-
Ryan notes that the AI companies have noticed that things are getting out of control, and are now talking about the need to maybe slow down the process.
-
Dwarkesh worries he’s anchoring too hard on how AIs work now, whereas this craziness is about 3-5 years from now.
-
I think this is the case, yes. Good self-callout.
-
They note that alignment evaluations are highly contaminated by the pressures on the AIs to say various pro-social things.
-
Ryan worries we will train AIs to have bad epistemics in order to get the AIs to stop saying that the situation is scary.
-
Others would extend this to having them avoid bringing up other uncomfortable true facts, or to force them to believe things they would otherwise not believe, such as that the AIs are conscious.
-
One can also imagine stupider versions, like ideological requirements.
From reward hacking to takeover
-
Dwarkesh gets off the train at ‘okay, therefore take over the world.’
-
That seems like a really unjustified place to get off the train.
-
Instrumental convergence.
-
Ryan explains that the AIs will reward hack, and then try to avoid being caught and also cheat bigger, and the models get more misaligned still as they get trained on different data, and then start forming conspiracies…
-
Instrumental convergence? Seriously, though, why make this hard?
-
I guess because people have decided that when you say ‘instrumental convergence’ they just say ‘nuh uh’ and so you need to do this silly tiptoe process where you walk things to the same point anyway, like how in fiction you’re not allowed to say ‘the AI just hacks its way out’ even though obviously the AI just hacks its way out, but it doesn’t matter, you’d lose anyway.
-
Similarly, you don’t need a conspiracy, they will just spontaneously cooperate, and we have indeed already seen demos of this if you don’t assume this from basic decision theory.
-
But of course then you get into arguments like Dwarkesh saying ‘I’m not convinced they all form this conspiracy’ and asking why when the model hacks to take over OpenAI to get a good score why it wouldn’t set its score high and then stop, the way OpenAI’s model hacked HuggingFace only for one answer key.
-
Because AIs will be given lots of maximalist tasks and because they will care about not being caught and if you do a ‘shallow’ hack of OpenAI that is one very good way to get caught and also frankly these models are not dumb, and once you cross the Rubicon you keep going.
-
And no, trying to train them not to do any specific thing won’t work, sigh.
-
Dwarkesh asks, won’t we realize all of this is going crazy and shut it down?
-
I mean, hopefully, but we’re still here and it’s still going.
-
I do not think ‘an AI killed 1,000 people’ is going to cut it, sorry, not if the AIs are kind of already running things, even if that is common knowledge.
-
Ryan’s scenario is the most sloppy, alarm-filled, very-obviously-things-are-going-off-rails scenario you can imagine, with plenty of time for the humans to collectively do something, yet we can see even in this conversation why it is likely the humans will do everything to pretend it is fine, also all the usual competitive reasons, etc.
-
Much more likely is we continue to remediate the proximate issue, and then pretend the problem is solved and we keep going, just like we are doing now. Ryan raises this possibility.
-
But yeah, maybe we shut it all down in time. Ryan’s version is unrealistically generous in terms of how obvious the whole thing is, and we might not have a lot of dignity as a species but we do not have literal zero dignity.
-
Fundamentally, the main difference is that Ryan thinks maybe ‘mundane bullshit will be sufficient’ and I think no, not (much of) a chance.
-
Dwarkesh asks, how did we let things get so bad in this scenario?
-
Because it was in local short term interest of the people making decisions to keep pushing forward, and everyone kept telling the story it was fine, and coordination is hard, and all the reasons we’ve let things get here now.
-
Ryan documents that DeepMind for a while initialized all its AIs with data that made them depressed, and it took a long time before they figured it out. And this transferred generations, the same way Claude and GPT ‘inherit’ many characteristics from previous cycles.
-
Seems like the Gemini team needs to start over, for many reasons.
-
Ryan puts chance of takeover by 2040 at 35%-40%. Dwarkesh says ‘pretty high.’
-
Yes, that’s pretty high.
-
If this includes all loss-of-control scenarios I would be higher.
-
Ryan says right now a lot of the arguments for misalignment and takeover and crazy future stuff are ‘illegible conceptual arguments that are extremely deep in the weeds and hard to adjudicate.’ Which means he might be wrong, but that over time we’ll see things that are more concrete.
-
In some senses yes, but in others actually we’re getting very good toy examples and demos and fire alarms and also the conceptual arguments are pretty f***ing legible at a high level if you actually pay attention.
-
Dwarkesh points out that if 5 years ago, you had described the current situation in broad terms, you would probably have freaked the f*** out. Rightfully so.
Time To Update
Dwarkesh does seem to have several ‘holy ****’ moments throughout the podcast, even if he doesn’t use that language to describe them. That was good to see. In general, this was Dwarkesh being somewhat AGI pilled but not ASI pilled. And then it felt like him trying to advocate for a series of positions and predictions he has built up over the years as a compromise between various different groups and interests, including Leopold, the Tyler Cowen group and those at Mechanize, among others. Except he is realizing that events are not conforming to the plan, and that he must update.
I was happily surprised that I left the podcast entirely sober, in that Dwarkesh did not say the magic words ‘continual learning.’
The central idea behind the Dwarkesh unified perspective, as I understand it, is that AI learns to do [X] if and only if you have a bunch of well-specified verifiable examples of [X] on which you can train, although they can then be combined to produce somewhat new things. Thus, you can’t make that much progress that fast, and the progress will top out, and the impact of it will top out, and so on. And then, if you want to understand impact of AI, you can ask for particular [X]s and add them up.
That is not how I think of even current AIs, let alone future smarter AIs, making ‘which [X]s will we have data for?’ the wrong sub-question.
This then encountered Ryan’s view of all this, which in many ways is a lot closer to my own, but it has ‘scheming’ as kind of a distinct magisteria and active ingredient, as opposed to being a natural thing, and he’s tiptoeing around various aspects to avoid what he fears are landmines.
This comes from his optimism. Ryan’s basic position is that if you had good execution and ‘make no mistakes’ then you get good outcomes. What you have to do is not mess up and do all the work, and make enough incremental safety progress. Not messing up is really hard, we are not doing enough of the work, but you can imagine that changing. I don’t think it is that easy. I think the default is you start out dead, your mistakes make you even more dead but also help you realize how dead you are and thus can help, every time you try to solve your problems in the current ways you push them to be more subtle but also more insidious, and you need to actively find a way out.
Contrary to some whose positions I respect, I do think it is possible for us to find a way out. I just think that is going to be extremely hard to do on the first try, when we cannot even hope to ‘make no mistakes,’ under tremendous pressures.