{"slug": "decoy-direction-optimization-a-post-hoc-defense-against-llm-abliteration", "title": "Decoy Direction Optimization: A Post-Hoc Defense Against LLM Abliteration", "summary": "Researchers propose Decoy Direction Optimization, a post-hoc defense that counters Refusal Feature Ablation (RFA), an attack that projects out a linear refusal direction from a language model's residual stream to bypass safety guardrails at a high attack success rate (ASR) while preserving model capability. The defense targets open-weight language models whose safety guardrails can be readily bypassed by RFA.", "body_md": "Safety guardrails in open-weight language models can be readily bypassed using Refusal Feature Ablation (RFA), a technique that identifies and projects out a linear refusal direction from the residual stream, often achieving a high attack success rate (ASR) while preserving model capability. Defendi", "url": "https://wpnews.pro/news/decoy-direction-optimization-a-post-hoc-defense-against-llm-abliteration", "canonical_source": "https://aiflash.com/news/120521/", "published_at": "2026-09-16 05:30:05+00:00", "updated_at": "2026-09-16 05:36:34.177435+00:00", "lang": "en", "topics": ["ai-safety", "large-language-models", "ai-research"], "entities": ["Refusal Feature Ablation", "Decoy Direction Optimization"], "alternates": {"html": "https://wpnews.pro/news/decoy-direction-optimization-a-post-hoc-defense-against-llm-abliteration", "markdown": "https://wpnews.pro/news/decoy-direction-optimization-a-post-hoc-defense-against-llm-abliteration.md", "text": "https://wpnews.pro/news/decoy-direction-optimization-a-post-hoc-defense-against-llm-abliteration.txt", "jsonld": "https://wpnews.pro/news/decoy-direction-optimization-a-post-hoc-defense-against-llm-abliteration.jsonld"}}