Decoy Direction Optimization: A Post-Hoc Defense Against LLM Abliteration Researchers propose Decoy Direction Optimization, a post-hoc defense that counters Refusal Feature Ablation (RFA), an attack that projects out a linear refusal direction from a language model's residual stream to bypass safety guardrails at a high attack success rate (ASR) while preserving model capability. The defense targets open-weight language models whose safety guardrails can be readily bypassed by RFA. Safety guardrails in open-weight language models can be readily bypassed using Refusal Feature Ablation RFA , a technique that identifies and projects out a linear refusal direction from the residual stream, often achieving a high attack success rate ASR while preserving model capability. Defendi