16:43
2026-06-04
lesswrong.com
ai-safety
Training Deliberative Monitors for Black-Box Scheming Detection
Researchers have developed a cost-effective method for detecting scheming behavior in AI agents by training small open-weight "deliberative monitors" that reason over a scheming specification before jโฆ