{"slug": "characterizing-a-configuration-where-inference-time-prm-pruned-fragment-grafting", "title": "Characterizing a Configuration Where Inference-Time PRM-Pruned Fragment Grafting Is Inert: Evidence from Three Reasoning LMs", "summary": "A study of PRM-Pruned Fragment Grafting (PPFG) on Qwen2.5-7B-Instruct with Math-Shepherd on full MATH500 (n=500, three seeds) found the inference-time intervention statistically indistinguishable from an independent parallel chain-of-thought baseline on every measured axis, with a hindsight oracle bounding any per-problem gain at +0.13 pp. A four-bucket classification of 322 stagnation-rule injection events showed only 14% targeted a genuinely struggling chain, and a random control matched the same parity at 2.4x the firing rate, so the inertness is not heuristic-specific. The finding replicated across three base LMs, six benchmarks, a second PRM, and a compatibility-gate sweep, with two-one-sided-tests analysis promoting the parity to positive equivalence on all twelve Qwen/LLaMA cells.", "body_md": "arXiv:2610.00047v1 Announce Type: new \nAbstract: Diversity collapse in parallel chain-of-thought has motivated inference-time interventions built on a natural design: when a process reward model (PRM) prunes a chain, its high-PRM prefix is extracted and grafted verbatim as an in-context demonstration into a still-decoding sibling. We isolate this mechanism, PRM-Pruned Fragment Grafting (PPFG), as the most cost-minimal operationalization of cross-trajectory step-level transfer, and test it at the operating point where prior fragment-grafting work reports gains only under additional compensating ingredients. On Qwen2.5-7B-Instruct with Math-Shepherd on full MATH500 (n=500, three seeds), PPFG in both stagnation- and random-targeting variants is statistically indistinguishable from an independent parallel-CoT baseline on every measured axis. We characterize why: a four-bucket classification of 322 stagnation-rule injection events shows only 14% targeted a genuinely struggling chain; the rest landed on chains that had already succeeded, were near completion, or sat on a flat PRM plateau, states a rescue graft cannot change. No compound-gate refinement jointly achieves well-targeted firing and adequate density, and a random control matches the same parity at 2.4x the firing rate, so the inertness is not heuristic-specific. The finding replicates across three base LMs, six benchmarks, a second PRM, and a compatibility-gate sweep; two-one-sided-tests analysis promotes the parity to positive equivalence on all twelve Qwen/LLaMA cells. A per-event spot-check finds injected chains prune at 2.75x the matched-step rate, but a surviving-sibling counterfactual finds no population-level compensation. A hindsight oracle bounds any per-problem gain from choosing PPFG over independent at +0.13 pp. We contribute an equivalence-testing template for establishing inference-time mechanism nulls, with every claim scoped to its tested operating point.", "url": "https://wpnews.pro/news/characterizing-a-configuration-where-inference-time-prm-pruned-fragment-grafting", "canonical_source": "https://arxiv.org/abs/2610.00047", "published_at": "2026-10-02 04:00:00+00:00", "updated_at": "2026-10-02 04:15:10.483007+00:00", "lang": "en", "topics": ["ai-research", "large-language-models", "machine-learning", "artificial-intelligence"], "entities": ["Qwen2.5-7B-Instruct", "Math-Shepherd", "MATH500", "LLaMA", "PRM-Pruned Fragment Grafting"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/characterizing-a-configuration-where-inference-time-prm-pruned-fragment-grafting", "markdown": "https://wpnews.pro/news/characterizing-a-configuration-where-inference-time-prm-pruned-fragment-grafting.md", "text": "https://wpnews.pro/news/characterizing-a-configuration-where-inference-time-prm-pruned-fragment-grafting.txt", "jsonld": "https://wpnews.pro/news/characterizing-a-configuration-where-inference-time-prm-pruned-fragment-grafting.jsonld"}}