{"slug": "stat-why-benchmarking-clinical-llms-from-openevidence-doximity-is-complicated", "title": "STAT+: Why benchmarking clinical LLMs from OpenEvidence, Doximity is complicated", "summary": "A Nature Medicine study pitting clinical AI systems OpenEvidence and UpToDate Expert AI against general LLMs sparked controversy in the clinical AI world, with STAT health tech correspondent Katie Palmer noting that the way benchmarks are summarized into headlines obscures their limited meaning.", "body_md": "Y*ou’re reading the web edition of STAT’s AI Prognosis newsletter, our subscriber-exclusive guide to artificial intelligence in health care and medicine. Sign up to get it delivered in your inbox every Wednesday.*\n\nI saw “The Odyssey” during its opening weekend. Ever since then, I have been questioning whether I’m illiterate or whether Christopher Nolan is a poor storyteller. This [London Review of Books evaluation of the film](https://www.lrb.co.uk/the-paper/v48/n14/emily-wilson/an-uncomplicated-man), written by the woman whose translation of “The Odyssey” Nolan apparently read, has freed me from my wondering. (h/t to my colleague Matthew Herper)\n\nHot takes on Homer’s epic, or hot tips about Epic Systems: [[email protected]](/cdn-cgi/l/email-protection)\n\n## Benchmark battle bots\n\nYou might recall that in mid-June, there was a [Nature Medicine study](https://www.nature.com/articles/s41591-026-04431-5) that pitted clinical AI systems OpenEvidence and UpToDate Expert AI against general LLMs. It set off a reaction in the clinical AI world like no other paper has. “The results rang out like a gunshot,” as STAT health tech correspondent Katie Palmer describes it.\n\nThe controversy surrounding the study, and everything that [came after](https://www.statnews.com/2026/07/09/openevidence-backs-study-that-finds-it-beats-llms-health-tech/), exemplifies the problems I have with benchmarks.\n\nKatie summed it up well when I talked to her yesterday: “The way that benchmarks have been talked about generally, and specifically in clinical AI, tends to summarize them into the headlines,” she said. “Every study needs a headline and every story needs a headline, but as we both know, and as I think most people in the industry know, an individual benchmark doesn’t mean much.”\n\n## This article is exclusive to STAT+ subscribers\n\n### Unlock this article — plus in-depth analysis, newsletters, premium events, and news alerts.\n\nAlready have an account? [Log in](/login/)\n\n[View All Plans](https://www.statnews.com/stat-plus/)", "url": "https://wpnews.pro/news/stat-why-benchmarking-clinical-llms-from-openevidence-doximity-is-complicated", "canonical_source": "https://www.statnews.com/2026/07/29/benchmarking-clinical-chatbots-openevidence-doximity-ai-prognosis/?utm_campaign=rss", "published_at": "2026-07-29 13:29:39+00:00", "updated_at": "2026-07-29 13:41:07.121905+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-research"], "entities": ["OpenEvidence", "UpToDate Expert AI", "Nature Medicine", "STAT", "Katie Palmer"], "alternates": {"html": "https://wpnews.pro/news/stat-why-benchmarking-clinical-llms-from-openevidence-doximity-is-complicated", "markdown": "https://wpnews.pro/news/stat-why-benchmarking-clinical-llms-from-openevidence-doximity-is-complicated.md", "text": "https://wpnews.pro/news/stat-why-benchmarking-clinical-llms-from-openevidence-doximity-is-complicated.txt", "jsonld": "https://wpnews.pro/news/stat-why-benchmarking-clinical-llms-from-openevidence-doximity-is-complicated.jsonld"}}