I built a tool to prove my multi-agent harness was worth it. It told me it wasn't.
A developer built a tool to measure whether multi-agent scaffolding improves coding task performance, only to find that adding a planner, two drafters, and a judge made results worse (80% vs. 95%) at …