{"slug": "sampling-luck-masquerades-as-allocation-gain-auditing-test-time-budget-for", "title": "Sampling Luck Masquerades as Allocation Gain: Auditing Test-Time Budget Allocation for Neural Combinatorial Optimization", "summary": "A new arXiv preprint (arXiv:2608.13087v1) reports that apparent gains from non-uniform test-time budget allocation in neural combinatorial optimization (NCO) are largely an artifact of in-sample evaluation. Across three pretrained solvers (POMO, AM, SymNCO) on uniform TSP-100, an oracle allocation computed and evaluated on the same stored samples reports a 2.2-2.6% gain, but out-of-sample the gain is indistinguishable from zero (0.457%, 0.015%, -0.512%). Under distribution shift, a pre-registered confirmatory experiment found that allocation guided by held-out sample statistics improves best-of-k by 11.5% for AM (95% CI [7.4, 19.7]) and 12.0% for SymNCO at equal evaluation budget, while a negative control (POMO) showed -0.3% [-0.7, 0.24]. The authors provide a correction procedure, a reporting checklist, and release all data, code, and pre-registration records.", "body_md": "arXiv:2608.13087v1 Announce Type: new\nAbstract: Neural combinatorial optimization (NCO) solvers report the best of many sampled solutions per instance, and the sample count is, by convention, identical for every instance. Whether a non-uniform allocation of a fixed total budget would buy anything has not been measured. We measure it, and we audit the measurement itself.\nFirst, on in-distribution workloads the allocation headroom is not detectable. Across three pretrained solvers (POMO, AM, SymNCO) on uniform TSP-100, an oracle allocation computed and evaluated on the same stored samples reports a 2.2-2.6% gain with intervals excluding zero; measured out of sample the same gain is indistinguishable from zero (0.457, 0.015, -0.512 percent). Following the customary in-sample procedure, all three solvers would have supported a published 2%-level gain that does not exist. We calibrate this bias against an instance-wise null in which the true gain is zero by construction; over the ranges we test it does not shrink with more samples or more instances.\nSecond, the same correction that removes the phantom gains preserves a real one. Under distribution shift (a workload mixing uniform and clustered instances), a pre-registered confirmatory experiment finds that allocation guided by held-out sample statistics improves best-of-k by 11.5% (AM, primary endpoint; 95% CI [7.4, 19.7]) and 12.0% (SymNCO, replication) at equal evaluation budget, with the signal-acquisition cost not charged; a pre-registered negative control (POMO, an order of magnitude more robust to shift) shows -0.3% [-0.7, 0.24]. The gain exceeds a frozen distribution-label baseline by 4.2 points [1.9, 7.7]. An exploratory policy charging a 20-sample probe against the same budget retains 3.4% (AM) and 4.6% (SymNCO).\nWe give a correction procedure and a reporting checklist, and release all data, code, and the pre-registration record.", "url": "https://wpnews.pro/news/sampling-luck-masquerades-as-allocation-gain-auditing-test-time-budget-for", "canonical_source": "https://www.machinebrief.com/news/sampling-luck-masquerades-as-allocation-gain-auditing-test-t-6zpq", "published_at": "2026-08-14 04:00:00+00:00", "updated_at": "2026-08-14 05:11:54.631757+00:00", "lang": "en", "topics": ["machine-learning", "artificial-intelligence"], "entities": ["arXiv", "POMO", "AM", "SymNCO", "TSP-100"], "alternates": {"html": "https://wpnews.pro/news/sampling-luck-masquerades-as-allocation-gain-auditing-test-time-budget-for", "markdown": "https://wpnews.pro/news/sampling-luck-masquerades-as-allocation-gain-auditing-test-time-budget-for.md", "text": "https://wpnews.pro/news/sampling-luck-masquerades-as-allocation-gain-auditing-test-time-budget-for.txt", "jsonld": "https://wpnews.pro/news/sampling-luck-masquerades-as-allocation-gain-auditing-test-time-budget-for.jsonld"}}