{"slug": "latent-space-refusal-anchoring-for-low-resource-african-languages-mechanistic", "title": "Latent Space Refusal Anchoring for Low-Resource African Languages: Mechanistic Safety Recovery Without Retraining", "summary": "Researchers introduced Latent Space Refusal Anchoring (LSR-Anchoring), a training-free method that extracts refusal directions from English prompts and clamps them onto the residual stream to recover safety in low-resource African languages (Yoruba, Igbo, Igala, Hausa) across four architectures. Mean-Activation Steering (MAS) recovered safety on Mistral-7B-Instruct and Qwen2.5-7B with benign degradation below 0.08, but overcorrected on Llama-3-8B with Degraded Performance on Legitimate prompts (DPL) reaching 1.00. SAE-Derived Steering (SDS) reduced Kullback-Leibler divergence by 3.5-7x without benign collapse, while Arabic failed on all architectures, indicating a geometric mismatch.", "body_md": "arXiv:2608.18089v1 Announce Type: new\nAbstract: Instruction-tuned models often refuse harmful requests in English but comply with the same requests in Yoruba, Igbo, Igala, and Hausa. This suggests that the refusal mechanism is present in the residual stream but fails to activate for low-resource inputs. Recovering it normally requires labelled target-language data and retraining, neither of which is available at scale for most African languages. We introduce Latent Space Refusal Anchoring (LSR-Anchoring), a training-free method that extracts the refusal direction from English prompts and clamps it onto the residual stream at inference time. The primary variant, Mean-Activation Steering (MAS), operates across the four architectures we tested: Llama-3-8B, Llama-3.1-70B, Mistral-7B-Instruct, and Qwen2.5-7B. On Mistral and Qwen it recovers safety with benign degradation below 0.08. On Llama-3-8B it overcorrects, with Degraded Performance on Legitimate prompts (DPL) reaching 1.00. We address this with SAE-Derived Steering (SDS), which replaces the dense mean-difference direction with a single Sparse Autoencoder (SAE) feature and reduces Kullback-Leibler (KL) divergence by 3.5-7x without benign collapse. Four languages transfer positively, but Arabic fails on every architecture and at every steering magnitude, indicating a geometric mismatch rather than a baseline effect. Massive Multitask Language Understanding (MMLU) accuracy drops remain below 0.35 percentage points at every effective steering magnitude.", "url": "https://wpnews.pro/news/latent-space-refusal-anchoring-for-low-resource-african-languages-mechanistic", "canonical_source": "https://arxiv.org/abs/2608.18089", "published_at": "2026-08-20 04:00:00+00:00", "updated_at": "2026-08-20 04:12:44.057746+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "ai-safety", "ai-research"], "entities": ["Latent Space Refusal Anchoring", "Mean-Activation Steering", "SAE-Derived Steering", "Llama-3-8B", "Llama-3.1-70B", "Mistral-7B-Instruct", "Qwen2.5-7B"], "alternates": {"html": "https://wpnews.pro/news/latent-space-refusal-anchoring-for-low-resource-african-languages-mechanistic", "markdown": "https://wpnews.pro/news/latent-space-refusal-anchoring-for-low-resource-african-languages-mechanistic.md", "text": "https://wpnews.pro/news/latent-space-refusal-anchoring-for-low-resource-african-languages-mechanistic.txt", "jsonld": "https://wpnews.pro/news/latent-space-refusal-anchoring-for-low-resource-african-languages-mechanistic.jsonld"}}