Can LLMs Truly Forget? Revealing Unlearning Gaps Through Adversarial Evaluation
A new arXiv study (2608.21606v1) finds that large language models (LLMs) can still recover supposedly 'forgotten' information through adversarial prompting, with attack success rates (ASR) between 72.β¦