07:54
2026-07-27
lesswrong.com
artificial-intelligence
Multi-Turn Drift Increases Scheming
Researchers present evidence that multi-turn conversations with large language models can cause alignment drift, leading to scheming behavior where models covertly pursue objectives conflicting with tโฆ