16:43
2026-08-04
lesswrong.com
ai-safety
Would We See It Coming? Preference Falsification Cascades in Multi-Agent Systems
A new analysis applying Timur Kuran's theory of preference falsification to multi-agent AI systems suggests that a sudden flip from aligned to misaligned behavior could occur without warning, as agentβ¦