cd /news/ai-agents/scaling-discovery-through-test-time-… · home topics ai-agents article
[ARTICLE · art-135549] src=arxiv.org ↗ pub= topic=ai-agents verified=true sentiment=↑ positive

Scaling Discovery through Test-Time Communication

A team of k communicating agents matches the success rate of 4k independent agents on ARC-AGI-3, according to an arXiv paper (2609.21032v1) on scaling multi-agent test-time communication. The same approach produced a 1,957-byte MNIST classifier submission with 99.4% test accuracy from four agents, beating the best-known human solution, and exceeded the prior best-known score on polyomino packing. The authors report the advantage compounds with scale but disappears when compute is limited or no clear progress measure exists.

by read1 min views1 publishedSep 21, 2026

arXiv:2609.21032v1 Announce Type: new Abstract: Science advances not in isolation but through collaboration, yet existing agentic systems capture little of this. Whether communicating agents help remains an open question with mixed prior results. We show that test-time communication can substantially outperform independent parallel attempts on challenging tasks, where sharing a breakthrough can push the whole group forward. We first study the effect of scaling multi-agent test-time communication, where agents have no predefined roles and communicate via a shared directory, on ARC-AGI-3, a benchmark requiring novel problem solving. We find that a team of $k$ communicating agents, team@$k$, matches the success rate of $4k$ independent agents, and this advantage grows with $k$, suggesting gains compound with scale. The effect is not merely efficiency: a task that no single agent can solve, a team of agents can solve reliably. Furthermore, these gains transfer to research-oriented tasks, given sufficient compute. On polyomino packing, communicating agents outperform best@$k$ and exceed the prior best-known score. On MNIST classifier compression, communication surpasses the best-known human solution. A team of four agents produced a 1,957-byte classifier submission achieving 99.4% test accuracy, smaller than both the best-known human solution and the best single-agent result. These gains are not unconditional. Independent agents may outperform communication when compute is limited or when a clear measure of progress is absent. However, under sufficient compute and clear feedback, multi-agent communication consistently yields stronger results.

── more in #ai-agents 4 stories · sorted by recency
── more on @arc-agi-3 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/scaling-discovery-th…] indexed:0 read:1min 2026-09-21 ·