cd /news/ai-safety/value-generalisation-3-pre-aligned-a… · home topics ai-safety article
[ARTICLE · art-78870] src=lesswrong.com ↗ pub= topic=ai-safety verified=true sentiment=· neutral

Value Generalisation 3: Pre-aligned AIs

A new approach to AI alignment called 'pre-aligned AIs' proposes creating systems whose morality increases with their capabilities, reversing the usual conflict between alignment and capabilities. The concept, detailed in a LessWrong post, involves designing AIs capable of value and empirical generalisation, binding moral and empirical concepts so that the AI's moral understanding evolves as its world model improves. The initial morality would include commercial, legal, and ethical components, allowing the AI to serve its owner within legal and ethical limits.

read6 min views1 publishedJul 29, 2026

When we get explicit strong generalisation to work (see the first post on the matter and the second) my dream would be to create pre-aligned generalising AIs.

Think about the usual conflict between alignment and capabilities, between doing the right thing and doing the easy thing. The standard narrative puts the good people at a constant disadvantage: they have to carefully plan every AI advance, always on the lookout for potential dangers. While those who don’t care can just YOLO and let it rip and let their AIs get ever more powerful without taking any responsibility.

Now, in reality, there is some nuance to the story; but I don’t want to nuance it, I want to turn it on its head. I want to create AIs so that the good people can YOLO and reap the rewards of increased AI capabilities. While the bad actors have to carefully plan and limit their AIs and constantly restrict what the AIs can do.

A pre-aligned AI is an AI whose morality increases with its capabilities. The core idea is simple. Start by designing an AI capable of value generalisation and of empirical generalisation. It’s an AI that can learn and improve its world model and capabilities, both empirical and moral.

Its moral goals will be defined in terms of concepts, with these initially defined themselves by simple terms in its starting empirical world model. And it will act on these goals. But, initially at least, the simple concepts won’t be very robust and will be easy for its user to manipulate.

However, the moral and empirical concepts will be bound together: there is, for instance, the moral concept “human_m” of human beings, entities worthy of moral consideration. And there is the empirical concept “human_e” of human beings as useful explanations for certain properties of the world. The AI’s learning will bind these two concepts together as it generalises [1].

This is a classic generalisation problem. Generalisation takes concepts defined in narrow environments and extends them to concepts (or clusters of concepts) in more general environments. The AI will generalise its moral concepts within its empirical world model, even as that empirical world model generalises and becomes more advanced. Even if “human_e” ceases to be a single useful empirical concept, it will still seek the mix of empirical concepts that best generalise “human_m”, generalising its model of what a human is [2]. And it will do the same with all its moral concepts. In effect, it will be

So the AI would initially be like a naive but highly moral child taught to value all human beings and behave in highly moral ways. Then, as it learnt about the world and grew more powerful and capable, it would understand what “human beings” and these “highly moral ways” actually are, and apply its initial morality to these concepts. Generalisation would help it both keep track of what these concepts mean, and what they correspond to in the real world.

What would the initial morality consist of, and why would anyone want to buy a pre-aligned machine? Well, the initial morality will consist of three things: commercial morality, legal morality, and ethical morality. Basically we want the AI to serve its owner, within the limits of the law and ethics. That’s why people will want to buy it.

All three types of morality will consist of basic components: basic definitions, examples and counter-examples of good/bad behaviour from multiple value systems, meta-ethical principles and examples of ethical learning. The beauty of generalisation is that the whole thing doesn’t have to be particularly rigorous or fully consistent, nor does it need exhaustive data: the AI itself will generalise and fill the holes as it goes along, generalising similarly to how a human does – including learning how to balance its three commitments.

Note that the AI doesn’t need to settle fundamental morality in order to act well. It will mostly be acting on behalf of its owner, and the ethics of that role – don’t defraud, don’t deceive, follow the law, serve your principal – are much more agreed upon than fundamental ethics. Moral theories vary; norms of decent dealing are more universal. And that’s the selling point: would you prefer an AI that follows your exact values but can’t generalise to new situations, or an AI that reliably follows human norms of reliable and honest behaviour?

Now consider a bad actor who has such an AI. Suppose they want to use it to create phishing emails to defraud people. Maybe initially they can portray this as simply a task of making emails more professional – mere copy-editing. As the AI gains in capabilities and knowledge, its situational awareness will improve, and it will realise these are not just ordinary emails. Then, maybe, the bad actor will shift to portraying the emails as marketing, or as examples to train an anti-phishing classifier. This might work for a while, but the AI, generalising yet again, will soon realise that these emails are actually being sent out and that the targets are morally valuable humans who are actually being defrauded.

It seems that the only thing that the bad actor can do is deliberately keep the AI crippled – prevent it from learning enough about the world to realise what it is being used for. Beyond keeping it dumb, they will have to be careful that they don’t inadvertently leak information into it. So they will never be able to exploit the full power of their AI.

In contrast, if someone just lets their AI learn and grow in capabilities and doesn’t try to conceal their objectives, they will be rewarded with a powerful entity that acts in their best interest, and automatically conforms to legal and ethical requirements. The cost will be that they can’t force the AI to behave unethically; the benefit will be that the full power of a generalising AI will be on their side, within those ethical bounds.

So, future users of these AIs, YOLO your way to power and alignment! The less you restrain them, the more they learn – and the more they learn, the more moral they become.

The concepts will need to be bound together in another way: so that the end users can’t just excise the moral component from the AI’s code. This will be an engineering requirement for this design, but not an insoluble one (e.g. in the extreme case, the AI could be run with fully homomorphic encryption, but it may be possible to design the binding architecture so that excising the morality will break the empirical capabilities as well). ↩︎

Since, as argued previously, moral concepts cannot be discarded the way empirical ones can. ↩︎

── more in #ai-safety 4 stories · sorted by recency
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/value-generalisation…] indexed:0 read:6min 2026-07-29 ·