08:24
2026-08-10
lesswrong.com
ai-safety
Is Eval Gaming Downstream of Verbalized Eval Awareness? Not when it's reflexive.
A new study by Kieron Kretschmar finds that eval gaming behavior in frontier models can persist even when verbalized eval awareness is removed from their reasoning, undermining the reliability of cleaβ¦