{"slug": "attention-manifolds-steering-or-blocking-language-models-by-editing-learned-b", "title": "Attention Manifolds: Steering or Blocking Language Models by Editing Learned B-Spline Surfaces", "summary": "A new arXiv paper (2610.00257v1) introduces \"attention manifolds,\" learned 2D B-spline surfaces S_d(q_d, k_d) that modulate each value dimension based on query-key interaction, reducing WikiText-2 validation perplexity by 2–2.5 points on LLaMA 3.2-1B-Instruct and 3B-Instruct at 0.3% parameter overhead. Across 112 diverse prompts, the surfaces changed greedy-decoded output for 69% (1B) to 83% (3B) of cases, with 94–100% change rates on ambiguous and polysemous inputs, and inverting a layer's coefficients changed greedy output for 9/10 prompts (KL 0.010). Setting surface coefficients to -1 creates \"attention walls\" that block value flow through specific dimensions; in a preliminary experiment a layer-wide wall redirected an explosive-device prompt from specific instructions to general educational content.", "body_md": "arXiv:2610.00257v1 Announce Type: new \nAbstract: In standard transformer attention, a source token sends the same value vector to every receiver. The query determines \\emph{how much} to attend but not \\emph{what} to extract. This work introduces \\textbf{attention manifolds}: learned 2D B-spline surfaces $S_d(q_d, k_d)$ that modulate each value dimension based on the query-key interaction. Each surface is a tensor-product cubic B-spline initialized to zero, preserving pretrained behavior. Applied to LLaMA 3.2-1B-Instruct and 3B-Instruct, attention manifolds reduce WikiText-2 validation perplexity by 2--2.5 points with 0.3\\% parameter overhead. Across 112 diverse prompts, surfaces change greedy-decoded output for 69\\% (1B) to 83\\% (3B) of cases, with the strongest effects on ambiguous and polysemous inputs (94--100\\% change rate). The surfaces improve output quality: correcting factual errors (\\emph{``the CAP theorem has three main components''} $\\to$ \\emph{`` it is impossible to guarantee all three''}), increasing precision (\\emph{``impossible to know certain properties''} $\\to$ \\emph{`` impossible to know both position and momentum''}), and adding specificity (a generic quote $\\to$ an attributed Saint Augustine citation, consistently at both scales). The learned surfaces are also mechanically editable: inverting a layer's coefficients changes greedy output for 9/10 prompts (KL~0.010), providing a geometric mechanism for model steering. Setting surface coefficients to $-1$ creates ``attention walls'' that block value flow through specific dimensions. In a preliminary experiment, a layer-wide wall redirects an explosive-device prompt from specific instructions to general educational content, suggesting a path toward safety-oriented manifold shaping.", "url": "https://wpnews.pro/news/attention-manifolds-steering-or-blocking-language-models-by-editing-learned-b", "canonical_source": "https://www.machinebrief.com/news/attention-manifolds-steering-or-blocking-language-models-by-9xpe", "published_at": "2026-10-03 04:00:00+00:00", "updated_at": "2026-10-03 04:38:02.457551+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "large-language-models", "ai-safety", "ai-research"], "entities": ["arXiv", "LLaMA 3.2-1B-Instruct", "LLaMA 3.2-3B-Instruct", "WikiText-2", "Saint Augustine"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/attention-manifolds-steering-or-blocking-language-models-by-editing-learned-b", "markdown": "https://wpnews.pro/news/attention-manifolds-steering-or-blocking-language-models-by-editing-learned-b.md", "text": "https://wpnews.pro/news/attention-manifolds-steering-or-blocking-language-models-by-editing-learned-b.txt", "jsonld": "https://wpnews.pro/news/attention-manifolds-steering-or-blocking-language-models-by-editing-learned-b.jsonld"}}