04:00
2026-09-28
arxiv.org
ai-research
When Is a Multi-Agent Code Judge Actually Grounded? Two Label-Free Measurements, and a Judge That Declines to Guess
Running the published MARCH multi-agent code-judging framework unmodified across 80 condition-by-cell measurements on two code judging benchmarks, researchers found it declares both candidate solutionβ¦