{"slug": "steering-instruction-hierarchies-at-inference-time", "title": "Steering Instruction Hierarchies at Inference Time", "summary": "A new training-free inference method called V-Steer restores instruction hierarchy compliance in large language models by editing cached value vectors, raising primary constraint accuracy from under 18% to 92% on controlled role conflict benchmarks across models from 7B to 70B, according to a preprint on arXiv.", "body_md": "arXiv:2607.26228v1 Announce Type: new\nAbstract: Instruction hierarchies are a core safety assumption of language model deployment: higher priority inputs, such as system prompts, should override conflicting lower priority inputs from users or tools. Yet frontier LLMs often violate this hierarchy. We introduce V-Steer, a training-free inference time method that restores privileged influence by editing cached value vectors at prompt positions. Using direct logit attribution on the first next token prediction, V-Steer identifies heads where lower priority spans dominate privileged ones, then boosts privileged spans and suppresses conflicting lower priority spans through in-place multiplicative edits to cached V tensors. Since the method acts only on cached values, it remains compatible with fused attention backends and adds only a one time prefill overhead. Across models from 7B to 70B, this attribution guided intervention raises primary constraint accuracy from under 18% up to 92% on controlled role conflict benchmarks, and on broader instruction hierarchy evaluations substantially outperforms prompt only baselines while matching or exceeding SoTA training based methods on 3 of 4 scales of LLMs, with negligible decoding-speed overhead. The code is available at https://github.com/cindy2000sh/v-steer.", "url": "https://wpnews.pro/news/steering-instruction-hierarchies-at-inference-time", "canonical_source": "https://www.machinebrief.com/news/steering-instruction-hierarchies-at-inference-time-ni8m", "published_at": "2026-07-30 04:00:00+00:00", "updated_at": "2026-07-30 04:32:12.142101+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-safety"], "entities": ["arXiv", "V-Steer"], "alternates": {"html": "https://wpnews.pro/news/steering-instruction-hierarchies-at-inference-time", "markdown": "https://wpnews.pro/news/steering-instruction-hierarchies-at-inference-time.md", "text": "https://wpnews.pro/news/steering-instruction-hierarchies-at-inference-time.txt", "jsonld": "https://wpnews.pro/news/steering-instruction-hierarchies-at-inference-time.jsonld"}}