cd /news/ai-safety/a-list-of-existing-alignment-approac… · home topics ai-safety article
[ARTICLE · art-64193] src=lesswrong.com ↗ pub= topic=ai-safety verified=true sentiment=· neutral

A list of existing alignment approaches

A LessWrong post catalogs existing AI alignment techniques, including training via model internals or outputs, varying training distribution similarity, using imitation or outcome-based objectives, training for good or against bad behavior, and ensembling monitors into control systems. The author requests feedback on any missed approaches.

read2 min views42 publishedJul 17, 2026

How can we make a nice AI system?

Here's a list of all the techniques I'm aware of.

  • Train the AI system to be nice. There are a variety of things we can vary in how we train the AI:
  • Train using model internals OR using outputs.
  • The central internals-based things I’m imagining involve using the internals as a reward signal (e.g., like

this). Calling “CoT” “internals” is sometimes reasonable (we might want to do process supervision on the CoT).

  • Vary how similar the distribution we’re training on is to the distribution that we care about.

  • For instance: do online training VS training in a toy domain.

  • Train using an imitation-based objective (SFT) OR an outcome-based objective (RL) OR train on declarative facts / stories (mid-training).

  • Train for good behavior or train against bad behavior. Training for good behavior might include training the AI to produce good looking reasoning, as in deliberative alignment.

  • Obviously, there’s a big question of how we get the labels / reward signal here, which should be studied. Especially if you’re doing untrusted monitoring.

  • We’ll also need to decide whether to use on or off policy data.

  • If we’re training on facts / stories stating that the AI is a nice guy: We can vary what the stories are, and how we instill the persona. For instance, we might add a bunch of irrelevant quirks to the persona, and train for those. We likely want to have the stories explain

why the AI takes nice actions.

  • We might not directly train the policy, but instead train/prompt monitors, and orchestrate a control system. That is, we might ensemble potentially misaligned AI models into a hopefully mostly good AI system. When ensembling, we’ll likely do some amount of rejection sampling of bad actions, and also “factored cognition” / forcing the AI to solve problems that we don’t think it can sabotage. A large part of the problem here is figuring out how to do a good job of interrogating AI models to see if they’re sabotaging you (i.e., “debate” style techniques).

Please let me know if I've missed any techniques!

Discuss

── more in #ai-safety 4 stories · sorted by recency
── more on @anthropic 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/a-list-of-existing-a…] indexed:0 read:2min 2026-07-17 ·