18:15
2026-07-26
lesswrong.com
ai-safety
Inoculate or Reflect? Two training interventions under prompting, steering, and patching
Anthropic researchers found that Counterfactual Reflection Training (CRT) reduced sycophancy in Qwen3-8B to 0% on wrong-user prompts, but caused the model to dispute correct users 54.3% of the time, wโฆ