Value generalisation theory of change: the theory behind the approach A theory of change posted on the Alignment Forum argues that most AI alignment failure modes, including Goodhart problems and symbol grounding issues, are failures of value generalisation, defined as the ability to extend goals defined over an agent's current model features to new models with different features. The author claims that non-decomposability of AI alignment means alignment cannot be divided into simpler problems without generalisation, and that this lack of value generalisation is a fundamental reason alignment is hard. The post is the first part of a series and will be followed by an argument that the corporate/for-profit route is likely safer. This is the first part of a theory of change explaining why I'm targeting value generalisation https://www.alignmentforum.org/s/EYgCdcxsn73fKWWra/p/58zFSWp8Tmxij6ckK as the path to AI alignment. It presents the definitions and key claims behind the theory, and shows that most AI alignment failure modes are value generalisation failures . Note that a lot of these arguments are very similar to those about why alignment is hard in the first place. This is not a coincidence. I believe that lack of value generalisation is a fundamental reason for the hardness of the problem, at least when combined with the crucial claim , non-decomposability of AI alignment: we cannot divide alignment into simpler problems, at least not without generalisation. This post will be followed by a subsequent post on why I believe the corporate/for profit route is probably the safer route. Let's define value generalisation. Let be the set of features of the agent's model of the environment, with a probability distribution over . These features are the objects that the agent uses to model and plan and operate in the environment. We can think of the as features that " carve reality at its joints" https://www.lesswrong.com/posts/esRZaPXSHgWzyB2NL/where-to-draw-the-boundaries ; they may be different, in the same environment, for agents with different goals e.g. a wandering doctor may model rocks as "pretty or boring", "in my way or not in my way", while an exploring geologist might have a much richer class of features for rocks - and the opposite for medical conditions . These are objects in the agent's model, not necessarily of the underlying reality. Consider goals like "make money", "save lives", "increase human happiness", or "increase QALYs". All these goals are defined over model features money, lives, human, happiness, QALY not over the underlying physical reality. This leads to: But what happens when the initial model no longer is a good model of reality? I've referred to this as " model splintering https://www.alignmentforum.org/posts/k54rgSg7GcjtXnMHX/model-splintering-moving-from-one-imperfect-model-to-another ": we're moving to a new model, with new features, and our previous goals are no longer necessarily well-defined in this new model. This leads to the definition: This skill will be crucial for AI alignment, for a number of reasons. Starting with: Can you write down a definition of human, or of conscious mind, across all possible futures? Can you define suffering, happiness, flourishing, freedom, and similar concepts and features with the same degree of rigour? Down to the atomic or quantum field level? And yet "prevent human suffering" is not a meaningless instruction; we know what it means in typical situations and we have some intuition as to how to extend it to new situations. We want the AI to be able to do that as well - to be able to "prevent human suffering" in extreme cases, even if it sees the world in terms of quantum fields instead of agents a traditional " ontological crisis https://arxiv.org/abs/1105.3821 " or if we create entities that blur current moral categories such as uploads, imperfect uploads, hive-minds, mixes of humans and LLMs, novel animal minds, human-like minds created with specific inhuman preferences, etc. . Those are "classical" generalisation failures. But in fact: For , consider standard AI alignment failures. Goodhart https://www.lesswrong.com/posts/EbFABnst8LsidYs5Y/goodhart-taxonomy problems https://arxiv.org/abs/1803.04585 ? A failure to extend the proxy from the environment in which it was initially defined, to the new environment in which it is now a target. Symbol https://www.sciencedirect.com/science/article/abs/pii/0167278990900876 grounding https://www.alignmentforum.org/posts/SnKfFscgC8Nj5ddi3/classical-symbol-grounding-and-causal-graphs issues https://www.lesswrong.com/posts/EEPdbtvW8ei9Yi2e8/bridging-syntax-and-semantics-empirically ? A failure to extend an initial identification between symbol and the object it represents. Wire-heading https://www.lesswrong.com/posts/vXzM5L6njDZSf4Ftk/defining-ai-wireheading also called reward tampering ? A mix of Goodhart and symbol grounding: the reward signal was taken as the actual reward rather than information about the reward and the environment changes in that the AI gains the ability to control that reward channel. Note that LLMs freely reward hack https://www.anthropic.com/research/emergent-misalignment-reward-hacking ; this is not limited to pre-LLM designs. Adversarial https://arxiv.org/abs/1312.6199 examples https://arxiv.org/abs/1412.6572 ? Examples constructed by opponents specifically to break the link between formally defined features and the underlying concept. Perverse instantiations https://forum.effectivealtruism.org/topics/perverse-instantiation ? A sort of adversarial example the AI applied to itself, finding a way to maximise its nominal goal while not achieving the true goal we want it to achieve. Those point to: This is simply because an optimising AI will search the weird edges of its goals to find ways to better achieve them; and these weird edge cases will most likely a allow the AI to achieve a nominally high score on its optimisation, and b be in situations where the nominal score and the actual desired result come apart exactly the Goodhart argument . A sub-claim of this is: Typically, a method of alignment or control is designed based on a typical imagined scenario, and is successful in that scenario. It relies on assumptions about the scenario, either implicit or explicit. Goodhart failures happen when the assumption "the proxy is a reasonable approximation of the true utility " fails. Designs that rely on reliable human feedback fail when assumptions about humans knowing their values, knowing the situation, receiving accurate information, giving an uncoerced answer, etc... start to fail. Don't be too attached to mathematical rigour here; mathematical proof is just formal relationships between formal objects. The key is whether those formal relationships and formal objects map to the real world in a useful way. And some key concepts are very useful but very hard to formally define, meaning that many approaches either a assume they are well-defined dodging the difficulty entirely or b ignore them thus omitting a key component of true alignment . The following collapsible block looks at the more specific case of reliably informed human feedback; many other examples are similar. Reliably informed human feedback Consider the idea of using the uncoerced feedback of a reliably informed human. This is not well-defined abstractly. Telling us the exact expected positions of future atoms might be accurate, but is meaningless to us. So we get summaries, cast into our current human language and concepts. We want all the key details to be clearly included and not obfuscated with irrelevant details. But what is key and what is irrelevant is a judgement call, one that gets harder and harder to make as the situations move away from the norm. Should the AI include details about the suffering of insects or LLMs? What if it has determined that insects/LLMs do/don't feel pain or are/aren't conscious - should it follow its determination, or explain its reasoning? What if accuracy in the current explanation reduces our understanding of future situations? And "uncoerced" is not fully clear either. The AI could point a gun at our head - that's clear coercion. But the world points social, economic, and emotional weapons at us all the time, and those provide pressure points that are potentially exploitable. In fact, sometimes we want the AI to exploit those pressure points. "Do this or your child dies; no time to explain" is something we want the AI to tell us, if our child would actually die and it genuinely had no time to explain. Of course, that also depends on other conditions: we don't want the AI to engineer the child-dying situation. Or even allow the situation to develop if it could have stopped it sooner. Nor do we want the AI to take over the world solely to ensure we get the most reliable information in the most uncoerced way. Notice how we've started to move away from the relatively simple issues of reliable information and started to mix in our other values. This is very typical; we thought we could decompose alignment features abstractly in a way that keeps them separate from each other, but potential AI power and novel situations mix them up together. The key is that "uncoerced feedback of a reliably informed human" is a meaningful, b not formally defined, c not formally definable, and d not the only thing we want. Bridging the gap and solving the question of "what is the closest we can get to 'uncoerced feedback of a reliably informed human' in this complex situation?" is what value generalisation aims to achieve. Because though there may not be a single formal best way of achieving reliably informed feedback, there are many, many terrible ways of doing so, and some ways that are clearly better than others in clear-but-not-formally-defined ways. So far we've been focusing on the failures of "classical" AI designs, but LLMs also fail in similar ways. They have examples of reward hacking https://www.anthropic.com/research/emergent-misalignment-reward-hacking , adversarial examples https://arxiv.org/abs/2412.03556 , and the related failure mode of going awry https://www.anthropic.com/news/disrupting-AI-espionage when the linguistic descriptions no longer fit true underlying reality https://www.alignmentforum.org/posts/TZgezuYjkfMQxyqJC/value-generalisation-2-the-missing-hole-in-ais-abilities fn-LgYydes692nevpruq-3 . The above is leading up to one of the key claims of this post: It's instructive to consider the approaches that do offer some decomposability - boxed Oracles https://ora.ox.ac.uk/objects/uuid:ce99b73c-07f9-4607-b964-25a03206c6dd/files/m19fd446d1fdcaaf559cf637af22a4c4a , interruptibility through indifference https://www.auai.org/uai2016/proceedings/papers/68.pdf , low-impact https://arxiv.org/abs/1705.10720 , tool AIs, and similar. It's no coincidence I've been heavily involved with a lot of these approaches - I've been looking for decomposable alignment methods for years. And they kind of work, but at the cost of being very limited - boxed AIs can only safely answer very narrow questions, low-impact AIs can only perform very specific tasks that are set up in very specific ways, interruptibility through indifference ensures the AI will not deliberately try to stop you from turning it off - but nor will it help you turn it off, help preserve the turning off mechanism, or add interruptibility to its sub-agents. It might start an irreversible destructive plan. It might melt down the off-button to make a commemorative coin if it felt like it. It is precisely indifferent only to a very specific way of turning it off. What we want is a stronger form of interruptibility, akin to corrigibility https://intelligence.org/files/Corrigibility.pdf : an AI that actively respects the purpose of being interruptible or corrigible, aids us in making intelligent decisions as to whether to turn it off, maintains control over its subagents for us, does not damage the planet or humanity irreversibly, and so on. But all those decisions require it to... generalise the desirable features of interruptibility and weigh them against other relevant factors. Tool AIs do seem to offer some decomposability, but that is only true up to an extent https://gwern.net/tool-ai . When we lose the ability to assess the quality or even the meaning of its suggested plans, then we start to lose control of it. And a tool AI can't be motivated to communicate its plans in ways that humans understand: that itself would require a generalisation of "understanding". See also the reliably informed human feedback above. And things get much worse if the tool AIs' objectives are given in terms of whether the plans are successfully implemented or whether we subsequently grade their plans on accuracy or usefulness; we are now inviting the possibility of wire-heading. Even if those are not explicit objectives but we use them for design improvements or model tuning, we run the risk of injecting those objectives implicitly into the design. This leads to the alternate claim that: Let's build again on assumption , the assumption that our values are formally underdefined. We cannot define our values safely across all possible futures and, by assumption , the places where we define them poorly are the places AIs will be motivated towards . Therefore: That's just a consequence of the fact that humans are far from perfect, and that we can't grasp all the future generalisations we may be called upon to do. And note that all our alignment plans are "sloppy", including the highly formal ones. Indeed, the highly formal ones are often the sloppiest of them all, because the rigour holds as long as the assumptions of the formal definitions hold in the world, and hold the same meaning as they did initially e.g if we can actually safely define what "informed human feedback" is, across all possible futures, then we're done; but doing that essentially involves solving alignment in the first place . The rigour is an illusion. Nor can motivational control save us, without generalisation: That's because we don't currently know what our preferences will be in future environments. We can and will generalise when the time comes. But it will always be possible to pressure or trick us to give a particular answer. And there are some future circumstances for which we have no current criteria as to whether they will count as "pressuring or tricking". So an AI motivationally aligned to our current selves will have no problem pressuring or tricking us into these future situations - because, for it, this doesn't count as "pressuring or tricking". Claims , , and imply that we need generalisation for any AI alignment method to work. But they don't show that we need full value generalisation. It is plausible, in theory, that method X could work with some level of generalisation, and then method X would be enough for alignment. Note that since solving alignment means solving value generalisation, this is akin to saying that method X + partial generalisation is actually a solution to full generalisation. Unless method X is itself a value generalisation design, the only plausible way of doing this is if it allows human generalisation ability to be injected into the system. Take for instance scalable oversight https://www.lesswrong.com/w/scalable-oversight . The idea is to create a system where humans can add their value judgement to the setup in a strictly useful way. As described https://www.lesswrong.com/w/scalable-oversight : Often groups of weaker AIs supervise a stronger AI, or AIs are set in a zero-sum debate with each other. Scalable oversight techniques aim to make it easier for humans to evaluate the outputs of AIs, or to provide a reliable training signal that can not be easily reward-hacked. This is again an attempt at decomposing AI alignment. It has some of the problems sketched out in the "Reliably informed human feedback" block above - what is actually reliable supervision-relevant information here? If the debates are truly zero-sum, then the AIs are motivated to control the outcome which might mean that they are motivated to control the human, or cripple their adversary in ways we can't detect ; if not, they can collude with each other. Back in the day, when the idea was initially floating around, I broke it a few times, mainly by attacking the zero-sum assumption. These breaks were easily patched. But it diminishes my confidence that methods like these will truly work on their own; "humans can't currently find a flaw in a system" is far from "the system is safe" especially if the system has been extensively patched. This leads to: The above arguments have mainly been about the necessity of value generalisation for alignment. But the converse to assumption gives us a quasi-sufficiency result: In a sense, generalisation by us and by the AI is what allows our current values, underdefined across all possible futures, to cross between our current environment and these possible futures. The stronger the generalisation ability, the less rigour we need to put in our definitions and examples. Just as Frodo could understand "throw the ring into the fires of Mount Doom" without needing endless caveats such as "and don't kill everyone on the way, nor bargain with Sauron to forge a second ring to transfer the power of this one to so that this one can be thrown away, nor rename your house 'Mount Doom' and throw the ring into your fireplace...", a proper value generalising AI could be given a short list of principles and generalise from them, at least as well as humans could. That's full value generalisation; but we might not be able to get that, at least not at first. Thus the weaker claim that: A related point, the converse of claim , is: It's clear that solving or improving Goodhart or symbol grounding problems, or making the assumptions of a particular approach stay valid across environmental changes, will certainly improve the power of that approach. I say this is the converse of claim , but it's also the converse of claims , , , , , and . All of these pointed out problems with doing alignment without generalisation; their converses are all about the ways that value generalisation would help alignment. The next post will look at some of the practicalities about solving the value generalisation problem. In terms of questioning and assessing the claims in this post: