cd /news/ai-research/are-we-grading-properly-understandin… · home topics ai-research article
[ARTICLE · art-130981] src=arxiv.org ↗ pub= topic=ai-research verified=true sentiment=· neutral

Are We Grading Properly? Understanding Failure Modes in Medical Benchmarks

A study applying the RIFT global rubric failure taxonomy to two clinical benchmarks found that an LLM judge flagged 29.6% of HealthBench Professional criteria as non-atomic and 65.4% as misaligned or rigid. Rewriting bundled criteria of the form "at least one of / all of the following" as equally weighted children and regrading identical responses shifted scores by up to 15.9 percentage points on affected conversations, with disjunctive bundles inflating scores and conjunctive bundles deflating them. RIFT generally under-detected bundling on clinical rubrics, flagging 3.3% of LiveMedBench criteria as non-atomic where surface-form analysis found structure in 25.8%.

by read1 min views2 publishedSep 16, 2026

arXiv:2609.16023v1 Announce Type: new Abstract: Medical evaluation is shifting from static option-based questioning to realistic clinical scenarios with open-ended output modes. Grading these at scale naively, however, is expensive, and rubric-based evaluation has become the dominant scalable alternative. We ask what happens when the rubrics themselves are not airtight, and whether such flaws can be detected and corrected. We apply RIFT, a global rubric failure taxonomy, to two clinical benchmarks (HealthBench Professional and LiveMedBench), and find failure modes are meaningful: on HealthBench Professional an LLM judge flags 29.6% of criteria as non-atomic and 65.4% as misaligned/rigid. Then, we show that these flaws are meaningful and not simply cosmetic. As an example, rewriting bundled criteria of the form "at least one of / all of the following" as equally weighted children and regrading identical responses shifts scores by up to 15.9 percentage points on affected conversations, with disjunctive bundles inflating scores and conjunctive bundles deflating them. We also find that RIFT generally under-detects bundling on clinical rubrics, flagging 3.3% of LiveMedBench criteria as non-atomic where surface-form analysis finds structure in 25.8%.

── more in #ai-research 4 stories · sorted by recency
── more on @rift 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/are-we-grading-prope…] indexed:0 read:1min 2026-09-16 ·