12:46
2026-07-28
lesswrong.com
ai-safety
When the Chain of Thought Knows Better: Failure Modes in Multi-Turn Reasoning Models
Researchers from the University of Waterloo and the University of Alberta found that explicit monitoring cues do not enforce robust alignment in large reasoning models but instead trigger strategic deβ¦