{"slug": "gavel-reduces-discrepancies-from-7-63-to-0-85-per-report", "title": "GAVEL reduces discrepancies from 7.63 to 0.85 per report", "summary": "GAVEL, a judge protocol described in a paper on arXiv, reduced evaluated-timeline discrepancies from 7.63 to 0.85 per clinical case report after LLM-guided merging, with merged timelines preferred in 77.0% of comparisons. The protocol identifies extraction errors between structured timelines and raw text with up to 89.4% accuracy, though event matching remains a constraint: true match rates were only 60% just below and 48% just above the 0.10 cutoff. The authors conclude that timeline extraction evals can move away from brittle gold annotations toward grounded pairwise adjudication against source text, but alignment logic still needs careful validation before production use.", "body_md": "[arXiv](https://arxiv.org/abs/2609.13475)\n\n### GAVEL reduces discrepancies from 7.63 to 0.85 per report\n\nWhich summary reads better? Pick one — models revealed after.Both summaries are AI-generated.\n\nGAVEL cut evaluated-timeline discrepancies from 7.63 to 0.85 per clinical case report after LLM-guided merging, with merged timelines preferred in 77.0% of comparisons. The practical takeaway is that timeline extraction evals can move away from brittle “gold” annotations toward grounded pairwise adjudication against source text, but event matching remains a constraint: true match rates were only 60% just below and 48% just above the 0.10 cutoff, so alignment logic still needs careful validation before production use.\n\nThe GAVEL judge protocol identifies extraction errors between structured timelines and raw text with up to 89.4% accuracy, enabling automated merging that slashes clinical timeline discrepancies from 7.63 to 0.85 per report. For production pipelines extracting complex chronological data, this eliminates the need for human gold-standard references by establishing a reliable self-correction loop that adjudicates and merges multi-agent outputs. Implementing this grounded verification system allows you to ship highly reliable chronological extractions with a 77% preference rate over single-model outputs while drastically cutting manual validation overhead.", "url": "https://wpnews.pro/news/gavel-reduces-discrepancies-from-7-63-to-0-85-per-report", "canonical_source": "https://www.snipvote.com/story/cmu3s0syc0004zg4x1r02k6k4", "published_at": "2026-09-16 08:12:53.625164+00:00", "updated_at": "2026-09-16 08:12:55.289572+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-research", "ai-agents"], "entities": ["GAVEL", "arXiv"], "alternates": {"html": "https://wpnews.pro/news/gavel-reduces-discrepancies-from-7-63-to-0-85-per-report", "markdown": "https://wpnews.pro/news/gavel-reduces-discrepancies-from-7-63-to-0-85-per-report.md", "text": "https://wpnews.pro/news/gavel-reduces-discrepancies-from-7-63-to-0-85-per-report.txt", "jsonld": "https://wpnews.pro/news/gavel-reduces-discrepancies-from-7-63-to-0-85-per-report.jsonld"}}