01:11
2026-08-15
lesswrong.com
ai-safety
Red vs Blue, but for Evals
A new LessWrong post by Evan R. Murphy proposes applying a red team vs. blue team framework to AI evaluations, arguing that current evaluation methodologies fail to account for models that can subvertβ¦