{"slug": "reasoning-jury-multi-model-consensus-for-evaluating-reasoning-traces", "title": "Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces", "summary": "Researchers introduced Reasoning Jury, a multi-model consensus system that uses a jury of open-weight LLMs and moderated deliberation to identify defects in long reasoning traces, outperforming frontier models such as Opus 4.6, Sonnet 4.6, and Gemini 3.1 Pro while costing only 8–15% as much. The system, detailed in arXiv:2608.12585v1, enables more accurate reasoning data curation and training signals for reasoning LLMs, and provides deeper insights into model failure modes on benchmarks.", "body_md": "arXiv:2608.12585v1 Announce Type: new\nAbstract: Improving reasoning LLMs requires the ability to judge the quality of long reasoning traces for effective reasoning data curation, strong training signals during reinforcement learning, and an in-depth understanding of reasoning behaviors during model performance evaluation. Additionally, surfacing reasoning mistakes that the model makes would enable improving the model's performance at runtime through providing feedback. Due to the difficulty of this complex task on long reasoning traces, single-model judges (even frontier models) do not do well at identifying reasoning defects. Additionally, leveraging frontier models during online training of reasoning LLMs is generally prohibited due to guardrails in terms of use. In this work, we introduce Reasoning Jury, a system that replaces the single judge with a jury of LLMs and a moderated consensus mechanism, to improve the fidelity of judgments for identifying reasoning defects. In reasoning jury, defects of a reasoning trace and their severity are surfaced through a deliberation where a moderator conducts a discussion amongst the jury where the jurors critique each other's judgments and get to modify their initial votes. The moderator derives a consensus through deliberation amongst jurors or consolidation of judgements. We show that Reasoning Jury with a jury of open-weight models (e.g., gpt-oss-120b) is able to significantly outperform frontier models (opus-4.6, sonnet-4.6, and gemini-3.1-pro) at correctly identifying reasoning defects. Besides accuracy performance improvements, the aggregated cost of the jury (initial verdicts, deliberations, consolidation, etc.) is a fraction (8 to 15%) of the cost of running frontier models in LLM-as-a-judge setup. We also show how these judgements can be leveraged to understand failure modes of reasoning LLMs on benchmarks, which allows much deeper understanding of a model's performance.", "url": "https://wpnews.pro/news/reasoning-jury-multi-model-consensus-for-evaluating-reasoning-traces", "canonical_source": "https://arxiv.org/abs/2608.12585", "published_at": "2026-08-14 04:00:00+00:00", "updated_at": "2026-08-14 04:10:15.713422+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-research", "ai-ethics"], "entities": ["Reasoning Jury", "arXiv", "gpt-oss-120b", "Opus 4.6", "Sonnet 4.6", "Gemini 3.1 Pro"], "alternates": {"html": "https://wpnews.pro/news/reasoning-jury-multi-model-consensus-for-evaluating-reasoning-traces", "markdown": "https://wpnews.pro/news/reasoning-jury-multi-model-consensus-for-evaluating-reasoning-traces.md", "text": "https://wpnews.pro/news/reasoning-jury-multi-model-consensus-for-evaluating-reasoning-traces.txt", "jsonld": "https://wpnews.pro/news/reasoning-jury-multi-model-consensus-for-evaluating-reasoning-traces.jsonld"}}