21:51
2026-07-19
lesswrong.com
ai-safety
Many alignment techniques work by training one model and deploying another
A new analysis identifies a common strategy behind several AI alignment techniques—steering vectors, inoculation prompting, and post-hoc honesty fine-tuning—which the author calls 'train-deploy mismat…