04:00
2026-09-16
arxiv.org
ai-research
Are We Grading Properly? Understanding Failure Modes in Medical Benchmarks
A study applying the RIFT global rubric failure taxonomy to two clinical benchmarks found that an LLM judge flagged 29.6% of HealthBench Professional criteria as non-atomic and 65.4% as misaligned or …