# Value Generalisation 1: a Research and Deployment Program

> Source: <https://www.lesswrong.com/posts/58zFSWp8Tmxij6ckK/value-generalisation-1-a-research-and-deployment-program>
> Published: 2026-07-29 15:57:50+00:00

I’m looking for people, advice, critiques, and funding to build a research program on value generalisation – the ability of an AI to correctly extend human values and preferences to situations neither it nor we have seen before. My ongoing research has become convinced that this is necessary if we want to get aligned AIs that operate in the human interest.

This would be a focused research organisation or a commercial venture. I’m leaning towards commercial, because alignment techniques confined to academic papers get ignored – or worse, mined for capability-relevant parts while the alignment component is discarded.

This post is the research program’s summary. The technical case is in the [next post](https://www.lesswrong.com/posts/TZgezuYjkfMQxyqJC/value-generalisation-2-the-missing-hole-in-ais-abilities), and one exciting consequence – AIs whose alignment grows with their capabilities – is in the [post after that](https://www.lesswrong.com/posts/uMKGaEKRDpoqnZyBh/value-generalisation-3-pre-aligned-ais).

Nothing technical stands in the way of you handing an AI assistant full control of your devices and accounts today. And I’m not saying using an app or harness that has been designed for these tasks. I mean hand an LLM your passwords, email, social media, and bank access along with a little note stating what you want, plugging inputs and outputs via APIs, and letting it go wild.

The capacity to do this exists. What doesn’t is the trust. And the trust is missing for a good reason: today’s AIs cannot be relied on to understand your interests in situations that weren’t covered – explicitly or implicitly – by their training and instructions. They extrapolate patterns naively. Push them past the situations they were shaped for, and they will still confidently do *something*; it just won’t reliably be what you wanted. This is tolerable in a chatbot. It is disqualifying in an autonomous agent, and it becomes more dangerous, not less, as the underlying capabilities improve.

The world will keep generating novelty. A few years ago, nobody had heard of AI psychosis or knew much about practical drone warfare; a few years from now there will be challenges we can’t currently name, not to mention new norms and expectations. Any AI acting with real autonomy will constantly face situations that its training didn’t pin down. It needs to cope, somehow.

An AI with working value generalisation would do what a good human assistant does: it would recognise when the situation is new, work out which of its principal’s values and preferences bear on it, and either act correctly or – when genuinely unsure – stop and asks a well-phrased question. It will push on until it is uncertain, not until it fails. Each answer it receives will teach it more about its principal’s values, so the questions get rarer and the delegation gets deeper. The product is reliable assistance capable of operating with minimal guidance and knowing when it's reached the limit of its abilities. This would allow a lot of AI applications – for a start, any situation today where the AI is right most of the time, but is not used because the consequences of a few misaligned decisions are severe.

I'll argue in the next post that this ability – I call it explicit value generalisation, distinguishing it from the naive generalisation of current systems and the human-guided generalisation of current oversight schemes – is not something that scaling and patching will deliver on their own. It is a specific missing capability, and it needs to be built deliberately.

Three asymmetries drive the timing. Firstly, capabilities don’t need explicit generalisation; alignment does. AIs may become very powerful without ever developing explicit generalisation – many companies are effectively betting on exactly that. But I don’t believe AIs can become aligned without explicit value generalisation. So if alignment is to be solved, value generalisation needs to be solved at some point. Better to solve it early and deliberately, and integrate it into weaker systems, rather than trying to bolt it onto highly capable ones later.

Secondly, and relatedly, AI capacities for deception will likely grow with their capabilities. An early deployment of value generalisation can be tested and validated much better than a later one.

Finally, deployment beats publication. As said above, alignment technique confined to academic papers get ignored or co-opted for capability work. The way to make value generalisation matter is to make it one of the most useful things on the market: learning AIs that can actually be trusted with delegation, deployed widely, with the alignment machinery the load-bearing piece that makes them reliable and trustworthy. The endpoint of that road is the pre-aligned AI described in the third [post](https://www.lesswrong.com/posts/uMKGaEKRDpoqnZyBh/value-generalisation-3-pre-aligned-ais): a system whose moral concepts are bound to its empirical ones, so that the more it learns about the world, the harder it becomes to misuse.

Note what the commercial pitch does not require: it does not require settling fundamental ethics. A delegated agent mostly needs the ethics of the role – don’t defraud, don’t deceive, follow the law, serve your principal, ask when unsure. Moral theories vary; norms of decent dealing converge. An AI that reliably follows human norms in novel situations is far more useful to delegate to than one that merely shares your values but can’t generalise them.

The research program decomposes into stages, each valuable on its own:

**I**. Out-of-distribution recognition for values. An AI that reliably knows when a situation has left the territory its values were trained for – and stops. The nuance is to distinguish “out of distribution” (which happens all the time) from “out of the distribution in a potentially value-relevant way”. This is already a deployable safety design and a sellable product feature: the agent that stops when error threatens.

**II**. Relevant concept selection. Working out which values and concepts actually bear on the novel situation – the skill that lets the AI bring a human into the loop intelligently, with a question that gives the human real situational awareness rather than deferring to the machine.

**III**. Full explicit value generalisation. Extending values correctly with less and less need for human input. This will lead ultimately to pre-aligned AIs and their empirical-moral concept-binding architecture.

Initial progress on each stage strengthens human oversight, with its better questions and *better timed* questions making us better able to control and direct the AI. Further progress will allow the AI to generalise our values further and act more reliably for human intent.

The research program builds on my old [concept-extrapolation posts](https://www.lesswrong.com/s/u9uawicHx7Ng7vwxA), with the addition of research done at Aligned AI, progress on resolving value generalisation challenges for different agent designs (including results on goal misgeneralisation challenges that, to my knowledge, no other approach has achieved), and more recent research. Initial results are promising: it seems doable to add initial versions of these stages to multiple different designs.

Get in touch if this is something you’d want to work on, contribute to, or critique. Get in touch by comment, DM here, or at [dragondreaming@gmail.com](mailto:dragondreaming@gmail.com)
