{"slug": "instruction-duplication-as-an-inference-time-control-primitive", "title": "Instruction Duplication as an Inference-Time Control Primitive", "summary": "A new arXiv paper (2609.04024v1) introduces instruction duplication, a black-box inference-time control that repeats procedural instructions without retraining, and reports that across seven instruction-tuned models and 16,800 scheduled generations on 300 medical multiple-choice questions, moving from one to two copies raises the deterministic All-8 diagnostic from 90.22% to 93.17%, eliminating 30.2% of remaining failures. The authors note that while a blinded challenge audit did not meet its prespecified 28/30 confirmation criterion (10/30 directional confirmations, 20/30 ties), the control can matter operationally when downstream systems act on generated trajectories, citing an Answer Engineering example where a trailing duplicate raised an endpoint from 84.2% to 97.1%.", "body_md": "arXiv:2609.04024v1 Announce Type: new\nAbstract: Procedural instruction following is a basic requirement for controllable language-model systems, especially when generated trajectories are inspected or repaired downstream. We introduce instruction duplication, a minimal black-box inference-time control that repeats only the procedural instruction, without retraining or decoding changes. Across seven instruction-tuned models, 300 medical multiple-choice questions, eight placement conditions, and 16,800 scheduled generations, moving from one to two copies raises the deterministic All-8 diagnostic--responses passing all eight observable tests--from 90.22% to 93.17% (+2.95 percentage points), eliminating 30.2% of the failures remaining after one copy. Pre-provisional TF-IDF recall rises from 73.44% to 74.81% (+1.38 points; Holm-adjusted p < .001), while final-answer accuracy remains exactly 60.21%. Premature commitment increases from 1.52% to 2.30% (p_Holm = .00536). A blinded challenge audit yields 10/30 directional confirmations, 20/30 perceptual ties, and no reversals; its prespecified 28/30 confirmation criterion is not met. Yet this distinction can matter operationally when a downstream system acts on the generated trajectory. In Answer Engineering (AE), where explicit trajectory state determines local repair, the published reason-first no-editing SSNHL endpoint was 25.1%; system-only AE was later reproduced at 84.2%, and the same trailing duplicate raised it to 97.1%. For conductive diagnostic branch preservation, the corresponding values are 58.9% published without editing, 78.6% with reproduced AE, and 73.8% with AE plus duplication--a within-AE decrease, but still 14.9 points above the no-editing baseline. Instruction duplication is therefore a low-complexity, placement-sensitive control whose practical value can emerge through the downstream system that consumes the exposed trajectory.", "url": "https://wpnews.pro/news/instruction-duplication-as-an-inference-time-control-primitive", "canonical_source": "https://www.machinebrief.com/news/instruction-duplication-as-an-inference-time-control-primiti-im7i", "published_at": "2026-09-04 04:00:00+00:00", "updated_at": "2026-09-04 04:52:24.573551+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-research", "ai-safety"], "entities": ["arXiv"], "alternates": {"html": "https://wpnews.pro/news/instruction-duplication-as-an-inference-time-control-primitive", "markdown": "https://wpnews.pro/news/instruction-duplication-as-an-inference-time-control-primitive.md", "text": "https://wpnews.pro/news/instruction-duplication-as-an-inference-time-control-primitive.txt", "jsonld": "https://wpnews.pro/news/instruction-duplication-as-an-inference-time-control-primitive.jsonld"}}