cd /news/artificial-intelligence/component-and-dimension-sparsity-in-… · home › topics › artificial-intelligence › article
[ARTICLE · art-146567] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Component and Dimension Sparsity in Transformer Refusal Mechanisms

A study of four open-weight large language models found that refusal behavior concentrates in sparse component mechanisms comprising 28–48% of upstream attention and MLP components, which retain 88–101% of full steering effectiveness, according to arXiv:2610.06903v1. Within those mechanisms, effective steering further concentrates in roughly 50% of residual stream dimensions, retaining 85–98% of the component-mechanism baseline. The authors released all code and raw experimental results at https://github.com/wang-research-lab/Refusal_Mechanisms.

by read1 min views1 publishedOct 7, 2026

arXiv:2610.06903v1 Announce Type: new Abstract: Activation steering manipulates large language model behavior by intervening on internal activations, but the mechanistic basis of these interventions remains poorly understood. We decompose refusal steering into component-level interventions across four open-weight models, identifying the sparse subsets of attention and MLP components whose steering suffices to reproduce the full behavioral effect. We find that refusal directions concentrate in sparse component mechanisms comprising 28--48% of upstream components, retaining 88--101% of steering effectiveness. Within these mechanisms, effective steering further concentrates in approximately 50% of residual stream dimensions, retaining 85--98% of the component-mechanism baseline, consistent with a privileged basis structure. Sparsity thus operates at two levels: which components are steered, and which dimensions within those components carry the signal. Together these findings show that refusal is not diffusely encoded across a transformer but assembled by a structured, identifiable mechanism, providing a foundation for mechanistic understanding of how refusal behaviors are represented and steered. To facilitate reproducibility, we release all code and raw experimental results in https://github.com/wang-research-lab/Refusal_Mechanisms.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @arxiv:2610.06903v1 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/component-and-dimens…] indexed:0 read:1min 2026-10-07 · —