Safety guardrails in open-weight language models can be readily bypassed using Refusal Feature Ablation (RFA), a technique that identifies and projects out a linear refusal direction from the residual stream, often achieving a high attack success rate (ASR) while preserving model capability. Defendi
The Reasoning Heist: Stealing Encrypted LLM Thoughts from GPT-5, Claude & Gemini — Fix It Now