{"slug": "does-fine-tuning-undo-activation-steering-behavioural-recovery-without-weight", "title": "Does Fine-Tuning Undo Activation Steering? Behavioural Recovery Without Weight-Edit Reversal", "summary": "A new arXiv preprint (arXiv:2608.24988v1) finds that activation steering embedded in language model weights survives fine-tuning mechanistically but not behaviorally: across five instruction-tuned models (3B-14B), refusal suppression lost 64% of its effect on average under supervised fine-tuning (SFT), yet the weight edit remained nearly intact (mean vector recovery ρ = 0.004, mean cosine similarity 0.074). The authors conclude that embedded steering is mechanistically durable but functionally vulnerable, requiring behavioral re-validation after downstream training.", "body_md": "arXiv:2608.24988v1 Announce Type: new\nAbstract: Activation steering can be embedded directly into a language model's weights, shaping behaviour without inference-time intervention and offering a way to encode alignment prior to release. However, models are routinely fine-tuned after deployment, and it is unknown whether embedded interventions survive this. We study the stability of embedded steering for refusal suppression and brevity induction across five instruction-tuned models (3B-14B) under non-adversarial SFT and RLHF. Behaviourally, preservation tracks the training data: steering degrades when optimisation pressure contradicts the targeted behaviour and persists otherwise, with refusal ablation losing 64% of its effect on average under SFT. Mechanistically, however, the weight edit survives almost untouched even where behaviour reverts: mean vector recovery is $\\rho = 0.004$, and the fine-tuning update along the steering direction is near-orthogonal to its pre-edit weight pattern (mean $\\cos\\theta = 0.074$). When steered behaviour degrades, fine-tuning does not achieve it by dismantling or reversing the steering mechanism itself. Embedded steering is therefore mechanistically durable but functionally vulnerable, and requires behavioural re-validation after downstream training.", "url": "https://wpnews.pro/news/does-fine-tuning-undo-activation-steering-behavioural-recovery-without-weight", "canonical_source": "https://arxiv.org/abs/2608.24988", "published_at": "2026-08-27 04:00:00+00:00", "updated_at": "2026-08-27 04:19:53.858633+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-safety", "ai-research"], "entities": ["arXiv"], "alternates": {"html": "https://wpnews.pro/news/does-fine-tuning-undo-activation-steering-behavioural-recovery-without-weight", "markdown": "https://wpnews.pro/news/does-fine-tuning-undo-activation-steering-behavioural-recovery-without-weight.md", "text": "https://wpnews.pro/news/does-fine-tuning-undo-activation-steering-behavioural-recovery-without-weight.txt", "jsonld": "https://wpnews.pro/news/does-fine-tuning-undo-activation-steering-behavioural-recovery-without-weight.jsonld"}}