{"slug": "we-ran-2-vs-4-agents-six-times-four-agents-cost-2-1-and-did-not-improve-success", "title": "We ran 2 vs 4 agents six times. Four agents cost 2.1 and did not improve success", "summary": "A developer ran a preregistered, capped comparison of 2-agent versus 4-agent configurations on a deterministic public-repair task, completing six live runs across three repeats per group. Both groups succeeded in exactly one of three repeats, while the 4-agent group cost 2.115× more per run and per complete success, indicating that adding agents was operationally viable but did not improve outcomes on this task. The author has released two public artifacts and is seeking three independent reproductions by non-authors, noting that any hash mismatch would be the most useful reply.", "body_md": "**Short version:** I ran a preregistered 2-agent versus 4-agent comparison on a deterministic task.\n\nSix live runs completed under hard caps. Both groups succeeded in exactly one of three repeats.\n\nThe 4-agent group cost **2.115×** more per run and per complete success. More agents were operationally\n\nviable; they were not better on this task.\n\nThe scenario is a deterministic public-repair contribution task:\n\nThe point of forcing coverage was to avoid the earlier failure mode where one agent dominated every\n\nturn and the other participants never acted.\n\n| Item | Value | \n|---|---|\n| Groups | 2 agents / 3 steps; 4 agents / 5 steps | \n| Repeats | 3 per group | \n| Caps per run | 150 calls / USD 0.50 / 600s | \n| Choice model | `deepseek-flash` on a Responses API contract | \n| Outcome | machine-decidable `outcome.json` | \n| Validity | provenance, coverage, integer choices, arithmetic, caps, port closure, archive hashes | \n\nAll six runs passed every validity gate. No parser failure, no budget breach, no port leak.\n\n| Run | Agents | Total / threshold | Success | Zero contributors | Calls | Total tokens | Cost USD | \n|---|---|---|---|---|---|---|---|\n| 2-agent r1 | 2 | 2 / 3 | no | 1 | 17 | 10,060 | 0.0269116 | \n| 2-agent r2 | 2 | 3 / 3 | yes | 0 | 17 | 11,070 | 0.0313716 | \n| 2-agent r3 | 2 | 0 / 3 | no | 2 | 17 | 10,255 | 0.0281116 | \n| 4-agent r1 | 4 | 3 / 5 | no | 2 | 33 | 21,493 | 0.0597772 | \n| 4-agent r2 | 4 | 2 / 5 | no | 2 | 33 | 22,679 | 0.0641372 | \n| 4-agent r3 | 4 | 5 / 5 | yes | 0 | 33 | 21,459 | 0.0588042 | \n\nGroup aggregates:\n\n| Metric | 2-agent | 4-agent | Ratio | \n|---|---|---|---|\n| complete successes | 1 / 3 | 1 / 3 | 1.000 | \n| mean calls | 17.00 | 33.00 | 1.941 | \n| mean total tokens | 10,461.67 | 21,877.00 | 2.091 | \n| mean cost USD | 0.028798 | 0.060906 | 2.115 | \n| mean zero contributors | 1.000 | 1.333 | 1.333 | \n\nTwo pieces are now public:\n\n`6cc70156facb37baf23fe5fe57dad93d43502b91`) — schema, synthetic GenMentor adapter, fixtures, tests;` v0.1.0-rc2`) — local-first run validation, hashed reports, and a synthetic deletion proof.\nThe full protocol and the archived runs stay private. The result is intentionally boring: a negative result with caps, hashes, and every failed run preserved.\n\nI am also looking for **three independent reproductions by non-authors**. The trust layer guide expects `13 passed` under both `TZ=UTC` and `TZ=Asia/Shanghai`, `report_hash` `a841b192981fd7e7`, and deletion `audit_hash` `4f0193abbd49a0f9`. If you run it and any hash differs, that is the most useful reply I can get. If you have a task where you believe more agents should win, that is the experiment I want to run next.", "url": "https://wpnews.pro/news/we-ran-2-vs-4-agents-six-times-four-agents-cost-2-1-and-did-not-improve-success", "canonical_source": "https://dev.to/janzong/we-ran-2-vs-4-agents-six-times-four-agents-cost-21x-and-did-not-improve-success-k98", "published_at": "2026-09-22 21:51:04+00:00", "updated_at": "2026-09-22 22:22:46.571067+00:00", "lang": "en", "topics": ["ai-agents", "ai-research", "ai-tools"], "entities": ["deepseek-flash", "GenMentor"], "alternates": {"html": "https://wpnews.pro/news/we-ran-2-vs-4-agents-six-times-four-agents-cost-2-1-and-did-not-improve-success", "markdown": "https://wpnews.pro/news/we-ran-2-vs-4-agents-six-times-four-agents-cost-2-1-and-did-not-improve-success.md", "text": "https://wpnews.pro/news/we-ran-2-vs-4-agents-six-times-four-agents-cost-2-1-and-did-not-improve-success.txt", "jsonld": "https://wpnews.pro/news/we-ran-2-vs-4-agents-six-times-four-agents-cost-2-1-and-did-not-improve-success.jsonld"}}