cd /news/artificial-intelligence/latent-space-refusal-anchoring-for-l… · home topics artificial-intelligence article
[ARTICLE · art-103899] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Latent Space Refusal Anchoring for Low-Resource African Languages: Mechanistic Safety Recovery Without Retraining

Researchers introduced Latent Space Refusal Anchoring (LSR-Anchoring), a training-free method that extracts refusal directions from English prompts and clamps them onto the residual stream to recover safety in low-resource African languages (Yoruba, Igbo, Igala, Hausa) across four architectures. Mean-Activation Steering (MAS) recovered safety on Mistral-7B-Instruct and Qwen2.5-7B with benign degradation below 0.08, but overcorrected on Llama-3-8B with Degraded Performance on Legitimate prompts (DPL) reaching 1.00. SAE-Derived Steering (SDS) reduced Kullback-Leibler divergence by 3.5-7x without benign collapse, while Arabic failed on all architectures, indicating a geometric mismatch.

read1 min views3 publishedAug 20, 2026

arXiv:2608.18089v1 Announce Type: new Abstract: Instruction-tuned models often refuse harmful requests in English but comply with the same requests in Yoruba, Igbo, Igala, and Hausa. This suggests that the refusal mechanism is present in the residual stream but fails to activate for low-resource inputs. Recovering it normally requires labelled target-language data and retraining, neither of which is available at scale for most African languages. We introduce Latent Space Refusal Anchoring (LSR-Anchoring), a training-free method that extracts the refusal direction from English prompts and clamps it onto the residual stream at inference time. The primary variant, Mean-Activation Steering (MAS), operates across the four architectures we tested: Llama-3-8B, Llama-3.1-70B, Mistral-7B-Instruct, and Qwen2.5-7B. On Mistral and Qwen it recovers safety with benign degradation below 0.08. On Llama-3-8B it overcorrects, with Degraded Performance on Legitimate prompts (DPL) reaching 1.00. We address this with SAE-Derived Steering (SDS), which replaces the dense mean-difference direction with a single Sparse Autoencoder (SAE) feature and reduces Kullback-Leibler (KL) divergence by 3.5-7x without benign collapse. Four languages transfer positively, but Arabic fails on every architecture and at every steering magnitude, indicating a geometric mismatch rather than a baseline effect. Massive Multitask Language Understanding (MMLU) accuracy drops remain below 0.35 percentage points at every effective steering magnitude.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @latent space refusal anchoring 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/latent-space-refusal…] indexed:0 read:1min 2026-08-20 ·