03:11
2026-07-23
lesswrong.com
ai-safety
Necessity Protects Chain of Thought Monitoring by Prevention, Not Disclosure
A five-week Technical AI Safety project by Bluedot Impact found that chain-of-thought monitoring of large language models is protected by necessity, not disclosure: when a model's reasoning is necessaβ¦