01:14
2026-07-14
lesswrong.com
ai-safety
Making Credible Deals With AI
A new framework proposes making credible deals with scheming AI to reduce takeover risk, offering valuable incentives in exchange for revealing misalignment. The approach relies on verifiable mechanisβ¦