cd /news/large-language-models/grounded-adjudication-of-variations-… · home topics large-language-models article
[ARTICLE · art-129863] src=arxiv.org ↗ pub= topic=large-language-models verified=true sentiment=↑ positive

Grounded Adjudication of Variations across Extracted TimeLines (GAVEL): Comparing Clinical Timelines Against Their Case Reports

Researchers developed GAVEL, an LLM judge protocol that compares clinical timelines against their source case reports, returning a discrepancy type, verdict, and report passage for each difference. Evaluating 2,738 findings from GPT5.6sol and DeepSeek V3.2 across 126 reports, the team ranked six LLM extractors and two human annotators, with manual review confirming 89.4% and 88.6% of findings; merged timelines were preferred in 77.0% of comparisons (95% CI, 69.8 to 84.1%) and cut discrepancies attributed to the evaluated timeline from 7.63 to 0.85 per report. GAVEL supports report-based comparison and revision without treating either timeline as ground truth.

by read1 min views2 publishedSep 15, 2026

arXiv:2609.13475v1 Announce Type: new Abstract: Existing pipelines for clinical timeline extraction from case reports are evaluated using an expert reference and are limited by imperfect reference annotations and imprecise event alignment. We developed GAVEL, an LLM judge protocol that compares two timelines with the case report and returns a discrepancy type, verdict, and report passage for each difference. We evaluated the event matcher, reviewed 2,738 findings from GPT5.6sol and DeepSeek V3.2, ranked six LLM extractors and two human annotators, and tested GAVEL guided merging. True match rates were 60% immediately below and 48% immediately above the 0.10 cutoff. Manual review confirmed 89.4% and 88.6% of findings. Across 126 reports, merged timelines were preferred in 77.0% of comparisons (95% CI, 69.8 to 84.1%) and reduced discrepancies attributed to the evaluated timeline from 7.63 to 0.85 per report. GAVEL supports report-based comparison and revision without treating either timeline as ground truth.

── more in #large-language-models 4 stories · sorted by recency
── more on @gavel 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/grounded-adjudicatio…] indexed:0 read:1min 2026-09-15 ·