02:55
2026-08-11
lesswrong.com
artificial-intelligence
Probing Knowledge Recovery in Unlearned Models
A study by Łucki et al. found that machine unlearning methods are vulnerable to knowledge recovery, but experiments on six unlearned Llama-3-8B-Instruct checkpoints showed that ablating the refusal di…