Author: Ujjal Bhattacharjee. Methods proposal for discussion, prepared with OpenAI Codex. No experimental results are reported here.
When several agents agree on a review, what has improved? They may have corrected a mistake. They may also have copied it, spent more money finding the same answer, or produced a cleaner-looking report. A useful evaluation needs to distinguish these outcomes.
The proposed study asks whether structured challenge and revision improve an evidence-review task under a fixed spending ceiling and deadline. The initial task is a synthetic income-discrepancy review: assess whether a stated figure is supported by the supplied packet. This is a bounded research task, not an autonomous lending decision.
The work is at the study-design stage. The current capture harness supports single and independent reviews; the deliberation conditions and their intervention records remain to be completed. An initial 10–20 synthetic cases would exercise the method. Domain-performance claims require independently validated cases and a separate held-out study.
| Condition | Procedure |
|---|---|
| Single reviewer | One strong reviewer with a declared self-check allowance |
| Independent aggregation | A fixed roster submits isolated reviews; a frozen mechanical rule combines them |
| Structured deliberation | The same saved initial reviews seed bounded challenge, response and revision |
| Deliberation with Jev | The same process includes precisely specified typed-evaluator interventions |
The comparison must name each intervention: when it runs, what it evaluates, what criterion it uses and what changes afterward. Logging a score without changing the process is not an intervention.
Jev is the model-based evaluator used by the proposed system. It returns structured judgments against declared criteria. The experiment tests whether those judgments help reviewers improve their work.
The two deliberation conditions share admission, evidence, tools, participant roster and final closure policy. Their substantive packets are compared at a common checkpoint before closure. Completion under the governance protocol is measured separately. A unanimity rule establishes whether a decision is authorized; it does not establish whether the answer is correct.
The primary outcome is an acceptable, evidence-supported packet delivered within budget and deadline. It must reach an acceptable disposition, recover required material findings, avoid material false findings and support decisive claims. A justified insufficient-evidence answer can succeed. An absent packet cannot.
Keep the initial and final findings. Independent graders can then identify repairs, newly introduced errors and the loss of a correct minority position. Report false alarms separately from missed material defects. Agreement and message volume are insufficient quality measures.
Every condition exports the same substantive packet and retains its raw response. Invalid structure remains an operational failure. A separate blinded diagnostic asks whether correct content was present but could not be extracted. That distinction prevents improved formatting from masquerading as improved reasoning.
The cost ledger includes participant calls, evaluator calls, tools, retries and failed attempts. Unknown usage stays unknown. A cheaper successful run does not establish lower cost per success if expensive failures disappear from the denominator.
Reusing initial reviews creates an accounting trap. The experiment pays for each shared call once, but every condition that depends on that work must include it in its own budget. Report both totals and mark reused work. Reconstructed timings must also include the initial stage and be distinguished from directly observed elapsed time.
Case variants and repeated model calls are correlated. Analysis should preserve case families and paired comparisons; extra repetitions do not create extra independent cases. Development examples should be separated from held-out families. Labels, stopping rules and aggregation must be frozen before inspecting held-out outcomes. The evaluator cannot grade its own benefit. Independent domain review is still required. Jev’s typed output helps software consume judgments, but confidence and schema validity do not establish factual correctness.
There is also a baseline choice to defend. Equal participant counts, equal token allowances and equal spending ceilings answer different questions. The primary comparison here concerns useful results under the same spending ceiling and deadline; actual resource use must still be reported.
Prior debate research motivates both enthusiasm and restraint. Du et al. report improvements in tested tasks. Choi et al. motivate strong independent-aggregation controls. Neither establishes the outcome of this proposed study.
The study should remain worth publishing if deliberation or Jev makes results worse. Its contribution would be a trustworthy comparison and an explanation of where the process fails.