Probing Knowledge Recovery in Unlearned Models
A study by Łucki et al. found that machine unlearning methods are vulnerable to knowledge recovery, but experiments on six unlearned Llama-3-8B-Instruct checkpoints showed that ablating the refusal di…
A study by Łucki et al. found that machine unlearning methods are vulnerable to knowledge recovery, but experiments on six unlearned Llama-3-8B-Instruct checkpoints showed that ablating the refusal di…
Researchers found that safety alignment in Diffusion Large Language Models (DLLMs) is sparse and transferable, enabling attacks that increase attack success rates from 2.6% to 73.8% on LLaDA and from …
Researchers introduced Causal Attribution Pruning (CAP), a training-free method that identifies critical attention heads in large language models by measuring their causal impact on reasoning tasks. C…
Researchers at arXiv introduced a dual-stance evaluation method to test whether activation steering on Llama-3-8B-Instruct reduces sycophancy without suppressing agreement with factually correct state…
Researchers introduced a paired fixed-content probe over 500 MMLU-Pro items to test how discourse-role labels such as Instruction:, Reference:, and Example: affect language model adoption of misleadin…