cd /news/ai-agents/an-exact-generate-transform-decompos… · home › topics › ai-agents › article
[ARTICLE · art-142222] src=arxiv.org ↗ pub= topic=ai-agents verified=true sentiment=· neutral

An Exact Generate - Transform Decomposition of Small-LLM Team Scaling Across Orchestration Architectures

A sweep of eight agent orchestration architectures across five instruction-tuned 7-9B models and six benchmarks found that scaling a team from three to thirty calls raises accuracy by up to 17 points on the arithmetic word-problem benchmarks GSM8K and GSMHard but by at most four points on ARC, GPQA, and MMLU, according to an arXiv paper (2609.36104v1). The Proposer-Critic architecture captured the arithmetic gains, scaling steepest and surpassing every other architecture at the largest budget with item-clustered intervals excluding zero, though it ranked among the weakest elsewhere and no architecture won across tasks. The authors attribute the split to an exact generate-transform decomposition separating an extensive coverage dividend from an intensive transformation change, and note token cost still varies 2.1x at equal call budgets.

by read1 min views3 publishedSep 30, 2026

arXiv:2609.36104v1 Announce Type: new Abstract: Replacing one LLM agent with a collaborating team can raise accuracy, but whether scaling the team helps, and which architecture to scale, is unclear. Sweeping eight agent orchestration architectures across five instruction-tuned 7-9B models, five short-answer benchmarks, and an executable-code benchmark up to 30 calls, we find that the returns to team scaling are sharply task-dependent: from three to thirty calls accuracy rises by up to 17 points on the two arithmetic word-problem benchmarks (GSM8K, GSMHard) but by at most four on ARC, GPQA, and MMLU, for every architecture, a split the usual task-averaged number conceals. Proposer-Critic captures the arithmetic gains, scaling steepest and, in aggregate, surpassing every other architecture at the largest budget (item-clustered intervals exclude zero), though it ranks among the weakest elsewhere, and no architecture wins across tasks. We explain these trajectories with an exact generate-transform decomposition. Partitioning any workflow into proposal coverage and a downstream transform, any accuracy change splits exactly into an extensive coverage dividend and an intensive transformation change. The decomposition diagnoses each task: arithmetic offers coverage headroom that a critic-guided transform converts, whereas the multiple-choice benchmarks either saturate in coverage or fail to convert it, and on open-ended code generative recovery nearly vanishes so accuracy tracks coverage. At equal call budgets token cost still varies 2.1x. Extra calls therefore create candidate opportunity that only some architectures, on some tasks, convert. Team scaling is a task- and architecture-specific bet, not a uniform lever.

── more in #ai-agents 4 stories · sorted by recency
── more on @gsm8k 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/an-exact-generate-tr…] indexed:0 read:1min 2026-09-30 · —