# Decoy Direction Optimization: A Post-Hoc Defense Against LLM Abliteration

> Source: <https://aiflash.com/news/120521/>
> Published: 2026-09-16 05:30:05+00:00

Safety guardrails in open-weight language models can be readily bypassed using Refusal Feature Ablation (RFA), a technique that identifies and projects out a linear refusal direction from the residual stream, often achieving a high attack success rate (ASR) while preserving model capability. Defendi
