16:03
2026-07-21
lesswrong.com
ai-safety
Steering Blackmail Through a Model's "Emotional State"
A case study on Gemma 3 12B reveals that the model's decision to blackmail an executive is not linearly decodable until late in its reasoning, peaking at layer 19 with 0.74 AUROC, and that steering anβ¦