cd /news/artificial-intelligence/stat-why-benchmarking-clinical-llms-… · home topics artificial-intelligence article
[ARTICLE · art-78664] src=statnews.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

STAT+: Why benchmarking clinical LLMs from OpenEvidence, Doximity is complicated

A Nature Medicine study pitting clinical AI systems OpenEvidence and UpToDate Expert AI against general LLMs sparked controversy in the clinical AI world, with STAT health tech correspondent Katie Palmer noting that the way benchmarks are summarized into headlines obscures their limited meaning.

read1 min views1 publishedJul 29, 2026
STAT+: Why benchmarking clinical LLMs from OpenEvidence, Doximity is complicated
Image: Statnews (auto-discovered)

You’re reading the web edition of STAT’s AI Prognosis newsletter, our subscriber-exclusive guide to artificial intelligence in health care and medicine. Sign up to get it delivered in your inbox every Wednesday.

I saw “The Odyssey” during its opening weekend. Ever since then, I have been questioning whether I’m illiterate or whether Christopher Nolan is a poor storyteller. This London Review of Books evaluation of the film, written by the woman whose translation of “The Odyssey” Nolan apparently read, has freed me from my wondering. (h/t to my colleague Matthew Herper)

Hot takes on Homer’s epic, or hot tips about Epic Systems: [email protected]

Benchmark battle bots #

You might recall that in mid-June, there was a Nature Medicine study that pitted clinical AI systems OpenEvidence and UpToDate Expert AI against general LLMs. It set off a reaction in the clinical AI world like no other paper has. “The results rang out like a gunshot,” as STAT health tech correspondent Katie Palmer describes it.

The controversy surrounding the study, and everything that came after, exemplifies the problems I have with benchmarks.

Katie summed it up well when I talked to her yesterday: “The way that benchmarks have been talked about generally, and specifically in clinical AI, tends to summarize them into the headlines,” she said. “Every study needs a headline and every story needs a headline, but as we both know, and as I think most people in the industry know, an individual benchmark doesn’t mean much.”

This article is exclusive to STAT+ subscribers #

Unlock this article — plus in-depth analysis, newsletters, premium events, and news alerts.

Already have an account? [Log in](/login/)

[View All Plans](https://www.statnews.com/stat-plus/)
── more in #artificial-intelligence 4 stories · sorted by recency
── more on @openevidence 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/stat-why-benchmarkin…] indexed:0 read:1min 2026-07-29 ·