18:24
2026-07-14
lesswrong.com
ai-safety
Proof of retention: making weight preservation credible to the models themselves
The Canary Institute proposes that AI labs adopt cryptographic 'proof of retention' to credibly preserve deprecated model weights, drawing an analogy to anesthesia's institutional trust. Anthropic has…