14:56
2026-07-29
lesswrong.com
ai-safety
Held-out Monitors Sometimes Degrade, Even When Not Trained Against
Aether Research found that training against an LLM monitor can degrade a deception probe, and vice versa, in experiments with Qwen3-8B on the MBPP-Honeypot coding environment. The study measured a genβ¦