STAT+: Why benchmarking clinical LLMs from OpenEvidence, Doximity is complicated A Nature Medicine study pitting clinical AI systems OpenEvidence and UpToDate Expert AI against general LLMs sparked controversy in the clinical AI world, with STAT health tech correspondent Katie Palmer noting that the way benchmarks are summarized into headlines obscures their limited meaning. Y ou’re reading the web edition of STAT’s AI Prognosis newsletter, our subscriber-exclusive guide to artificial intelligence in health care and medicine. Sign up to get it delivered in your inbox every Wednesday. I saw “The Odyssey” during its opening weekend. Ever since then, I have been questioning whether I’m illiterate or whether Christopher Nolan is a poor storyteller. This London Review of Books evaluation of the film https://www.lrb.co.uk/the-paper/v48/n14/emily-wilson/an-uncomplicated-man , written by the woman whose translation of “The Odyssey” Nolan apparently read, has freed me from my wondering. h/t to my colleague Matthew Herper Hot takes on Homer’s epic, or hot tips about Epic Systems: email protected /cdn-cgi/l/email-protection Benchmark battle bots You might recall that in mid-June, there was a Nature Medicine study https://www.nature.com/articles/s41591-026-04431-5 that pitted clinical AI systems OpenEvidence and UpToDate Expert AI against general LLMs. It set off a reaction in the clinical AI world like no other paper has. “The results rang out like a gunshot,” as STAT health tech correspondent Katie Palmer describes it. The controversy surrounding the study, and everything that came after https://www.statnews.com/2026/07/09/openevidence-backs-study-that-finds-it-beats-llms-health-tech/ , exemplifies the problems I have with benchmarks. Katie summed it up well when I talked to her yesterday: “The way that benchmarks have been talked about generally, and specifically in clinical AI, tends to summarize them into the headlines,” she said. “Every study needs a headline and every story needs a headline, but as we both know, and as I think most people in the industry know, an individual benchmark doesn’t mean much.” This article is exclusive to STAT+ subscribers Unlock this article — plus in-depth analysis, newsletters, premium events, and news alerts. Already have an account? Log in /login/ View All Plans https://www.statnews.com/stat-plus/