cd /news/ai-research/where-does-retrieval-based-open-ende… · home › topics › ai-research › article
[ARTICLE · art-140741] src=arxiv.org ↗ pub= topic=ai-research verified=true sentiment=↓ negative

Where Does Retrieval-Based Open-Ended Evaluation Fail? Automatic Taxonomy Induction from Long-Form Medical Answer Factuality Verification

A new arXiv paper (2609.30467v1) introduces two taxonomies for retrieval-based factuality evaluation failures in open-ended medical settings, decomposing errors into retrieval-stage failures across five quality dimensions and verifier-reasoning errors across six consecutive steps, using an LLM-as-Judge pattern induction pipeline on the MedExpert dataset and 3 closed-ended datasets. Stress-testing across 4 retrieval methods and 6 frontier verifier models found that scaling model size, adding reasoning effort, expanding to authoritative web sources, and applying medical fine-tuning do not resolve these failure modes, which the authors call fundamental limitations of the retrieve-then-verify paradigm rather than artifacts of outdated systems. Code and data are released at https://anonymous.4open.science/r/Medical_RAG_eval-4AB5.

by read1 min views2 publishedSep 28, 2026

arXiv:2609.30467v1 Announce Type: new Abstract: Retrieval-based factuality evaluation, where LLM-generated claims are verified against evidence from authoritative medical corpora, has become the dominant paradigm for scalable hallucination detection in high-stakes clinical settings. Despite the urgency of reliable and transparent medical fact verification, most systems measure performance with aggregate metrics like F1, which obscure where and why failures occur. Existing RAG diagnostics require gold answers or annotated gold evidence, neither of which exists in this regime. We introduce two comprehensive taxonomies, grounded in a case study on the open-ended MedExpert dataset and 3 closed-ended datasets, decomposing failures into retrieval-stage errors along five quality dimensions, and verifier-reasoning errors into six consecutive steps. We adapt an automatic pattern induction pipeline using LLM-as-Judge to label evidence quality and classify verifier reasoning errors at scale, and then stress-test our findings across 4 retrieval methods and 6 frontier verifier models. Our analysis reveals that scaling model size, adding reasoning effort, expanding to authoritative web sources, and applying medical fine-tuning do not resolve these failure modes, demonstrating that they represent fundamental limitations of the retrieve-then-verify paradigm in open-ended medical settings rather than artifacts of outdated systems. We release our code and data at https://anonymous.4open.science/r/Medical_RAG_eval-4AB5 for the full reproducibility of our results.

── more in #ai-research 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/where-does-retrieval…] indexed:0 read:1min 2026-09-28 · —