20:48
2026-07-20
lesswrong.com
ai-safety
Restoring Model Alignment via Honesty Activation Steering
Researchers demonstrate that honesty activation steering can restore model alignment in large language models, with selective steering methods StTP and StMP recovering honesty at a fraction of the capโฆ