21:51
2026-07-23
lesswrong.com
ai-safety
Red-teaming LLM unlearning: LUNAR's "forgotten" knowledge is still recoverable
A new study reveals that LUNAR, a state-of-the-art LLM unlearning method, fails to truly remove knowledge, as two attack routes can recover "forgotten" information. The author demonstrates that LUNAR'โฆ