04:41
2026-06-24
lesswrong.com
ai-safety
Can weak AI watch strong AI?
A new experiment tested whether weaker AI models can effectively monitor stronger coding agents for malicious behavior, finding that detection rates improve with monitor size but vary by threat type. β¦