{"slug": "how-should-we-test-whether-structured-agent-debate-improves-evidence-review", "title": "How should we test whether structured agent debate improves evidence review?", "summary": "Ujjal Bhattacharjee published a methods proposal, prepared with OpenAI Codex, for testing whether structured agent debate improves evidence review, with no experimental results reported. The proposed study would compare a single reviewer, independent aggregation, structured deliberation, and deliberation with the model-based evaluator Jev under a fixed spending ceiling and deadline, using an initial 10–20 synthetic income-discrepancy cases. The primary outcome is an acceptable, evidence-supported packet delivered within budget and deadline, with false alarms reported separately from missed material defects.", "body_md": "Author: Ujjal Bhattacharjee. Methods proposal for discussion, prepared with OpenAI Codex. No experimental results are reported here.\n\nWhen several agents agree on a review, what has improved? They may have corrected a mistake. They may also have copied it, spent more money finding the same answer, or produced a cleaner-looking report. A useful evaluation needs to distinguish these outcomes.\n\nThe proposed study asks whether structured challenge and revision improve an evidence-review task under a fixed spending ceiling and deadline. The initial task is a synthetic income-discrepancy review: assess whether a stated figure is supported by the supplied packet. This is a bounded research task, not an autonomous lending decision.\n\nThe work is at the study-design stage. The current capture harness supports single and independent reviews; the deliberation conditions and their intervention records remain to be completed. An initial 10–20 synthetic cases would exercise the method. Domain-performance claims require independently validated cases and a separate held-out study.\n\n| Condition | Procedure | \n|---|---|\n| Single reviewer | One strong reviewer with a declared self-check allowance | \n| Independent aggregation | A fixed roster submits isolated reviews; a frozen mechanical rule combines them | \n| Structured deliberation | The same saved initial reviews seed bounded challenge, response and revision | \n| Deliberation with Jev | The same process includes precisely specified typed-evaluator interventions | \n\nThe comparison must name each intervention: when it runs, what it evaluates, what criterion it uses and what changes afterward. Logging a score without changing the process is not an intervention.\n\nJev is the model-based evaluator used by the proposed system. It returns structured judgments against declared criteria. The experiment tests whether those judgments help reviewers improve their work.\n\nThe two deliberation conditions share admission, evidence, tools, participant roster and final closure policy. Their substantive packets are compared at a common checkpoint before closure. Completion under the governance protocol is measured separately. A unanimity rule establishes whether a decision is authorized; it does not establish whether the answer is correct.\n\nThe primary outcome is an acceptable, evidence-supported packet delivered within budget and deadline. It must reach an acceptable disposition, recover required material findings, avoid material false findings and support decisive claims. A justified insufficient-evidence answer can succeed. An absent packet cannot.\n\nKeep the initial and final findings. Independent graders can then identify repairs, newly introduced errors and the loss of a correct minority position. Report false alarms separately from missed material defects. Agreement and message volume are insufficient quality measures.\n\nEvery condition exports the same substantive packet and retains its raw response. Invalid structure remains an operational failure. A separate blinded diagnostic asks whether correct content was present but could not be extracted. That distinction prevents improved formatting from masquerading as improved reasoning.\n\nThe cost ledger includes participant calls, evaluator calls, tools, retries and failed attempts. Unknown usage stays unknown. A cheaper successful run does not establish lower cost per success if expensive failures disappear from the denominator.\n\nReusing initial reviews creates an accounting trap. The experiment pays for each shared call once, but every condition that depends on that work must include it in its own budget. Report both totals and mark reused work. Reconstructed timings must also include the initial stage and be distinguished from directly observed elapsed time.\n\nCase variants and repeated model calls are correlated. Analysis should preserve case families and paired comparisons; extra repetitions do not create extra independent cases. Development examples should be separated from held-out families. Labels, stopping rules and aggregation must be frozen before inspecting held-out outcomes.\n\nThe evaluator cannot grade its own benefit. Independent domain review is still required. Jev’s typed output helps software consume judgments, but confidence and schema validity do not establish factual correctness.\n\nThere is also a baseline choice to defend. Equal participant counts, equal token allowances and equal spending ceilings answer different questions. The primary comparison here concerns useful results under the same spending ceiling and deadline; actual resource use must still be reported.\n\nPrior debate research motivates both enthusiasm and restraint. [Du et al.](https://proceedings.mlr.press/v235/du24e.html) report improvements in tested tasks. [Choi et al.](https://proceedings.neurips.cc/paper_files/paper/2025/hash/934252acd87f254d5d4672fbde283bd2-Abstract-Conference.html) motivate strong independent-aggregation controls. Neither establishes the outcome of this proposed study.\n\nThe study should remain worth publishing if deliberation or Jev makes results worse. Its contribution would be a trustworthy comparison and an explanation of where the process fails.", "url": "https://wpnews.pro/news/how-should-we-test-whether-structured-agent-debate-improves-evidence-review", "canonical_source": "https://discuss.huggingface.co/t/how-should-we-test-whether-structured-agent-debate-improves-evidence-review/180785#post_1", "published_at": "2026-09-28 20:42:51+00:00", "updated_at": "2026-09-28 20:47:17.624865+00:00", "lang": "en", "topics": ["ai-agents", "ai-research", "ai-safety"], "entities": ["Ujjal Bhattacharjee", "OpenAI Codex", "Jev"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/how-should-we-test-whether-structured-agent-debate-improves-evidence-review", "markdown": "https://wpnews.pro/news/how-should-we-test-whether-structured-agent-debate-improves-evidence-review.md", "text": "https://wpnews.pro/news/how-should-we-test-whether-structured-agent-debate-improves-evidence-review.txt", "jsonld": "https://wpnews.pro/news/how-should-we-test-whether-structured-agent-debate-improves-evidence-review.jsonld"}}