{"slug": "more-programs-or-more-rolls-separating-coverage-from-specialization-in-llm", "title": "More Programs or More Rolls? Separating Coverage from Specialization in LLM Harnesses", "summary": "A controlled evaluation of eight generated LLM harnesses plus a baseline and nine byte-identical baseline copies on 386 MATH-500 tasks found that identical programs alone yield 2.16 percentage points of repeat-averaged oracle headroom, meaning coverage gains cannot be attributed to task specialization. Generated programs showed more repeatable score patterns, but losses versus the baseline persisted across all three repeats on 100 tasks while persistent wins occurred on only one task and were sensitive to answer extraction; the frozen selector gained 0.00 percentage points and both populations reached 98.70% oracle coverage at 27 harness executions. The authors conclude that coverage and repeatability alone cannot justify claims of useful specialization and propose an evaluation standard requiring task advantages to persist across executions, guide usable decisions, and beat additional fixed-program executions under matched inference budgets.", "body_md": "arXiv:2609.35873v1 Announce Type: new \nAbstract: Automated generation of LLM harnesses promises to improve inference through task specialization. Yet additional answer coverage can arise from repeated execution of the same program, making specialization difficult to identify. We introduce a controlled evaluation that separates answer coverage, repeatable task advantages, and gains from pre-execution selection. On 386 MATH-500 tasks, we compare eight generated harnesses plus a baseline with nine byte-identical baseline copies, using three executions per member. Identical programs yield 2.16 percentage points of repeat-averaged oracle headroom. Generated programs exhibit substantially more repeatable score patterns, but these chiefly reveal persistent weaknesses: losses relative to the baseline persist across all three repeats on 100 tasks, while persistent wins occur on only one task and are sensitive to answer extraction. The frozen selector gains 0.00 percentage points, and both populations reach 98.70% oracle coverage at 27 harness executions. Stable complementarity remains unresolved at three repeats. Supporting BIRD traces locate failures in mechanism implementation, activation, and output validity. Together, these findings establish why coverage and repeatability alone cannot justify claims of useful specialization. They motivate an evaluation standard for harness diversity: task advantages should persist across executions, guide usable decisions, and improve on additional fixed-program executions under matched inference budgets.", "url": "https://wpnews.pro/news/more-programs-or-more-rolls-separating-coverage-from-specialization-in-llm", "canonical_source": "https://arxiv.org/abs/2609.35873", "published_at": "2026-09-30 04:00:00+00:00", "updated_at": "2026-09-30 04:17:48.138033+00:00", "lang": "en", "topics": ["large-language-models", "ai-research", "machine-learning", "artificial-intelligence"], "entities": ["MATH-500", "BIRD"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/more-programs-or-more-rolls-separating-coverage-from-specialization-in-llm", "markdown": "https://wpnews.pro/news/more-programs-or-more-rolls-separating-coverage-from-specialization-in-llm.md", "text": "https://wpnews.pro/news/more-programs-or-more-rolls-separating-coverage-from-specialization-in-llm.txt", "jsonld": "https://wpnews.pro/news/more-programs-or-more-rolls-separating-coverage-from-specialization-in-llm.jsonld"}}