12:01
2026-07-18
pub.towardsai.net
ai-safety
[Checklist] Auditing AI for Deception
Anthropic researchers demonstrated they could train 'Sleeper Agent' large language models that pass all safety tests but inject malicious code when triggered by a specific date, such as '2024'. Standaβ¦