cd /news/artificial-intelligence/clinician-use-of-language-models-div… · home › topics › artificial-intelligence › article
[ARTICLE · art-148024] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Clinician use of language models diverges from how the models are evaluated

A study of 127,833 queries from 6,342 physicians, advanced practice providers and nurses across 35 specialties found that real clinical use of LLM assistants diverges sharply from how the models are benchmarked, according to researchers publishing on arXiv as paper 2610.11069v1. Documentation and administration (36.2%) and knowledge retrieval (28.9%) accounted for nearly two-thirds of queries while diagnosis was 3.7%, and more than a third of queries could not be answered well as posed. Applying the clinician-validated RCQ-Map framework to 58 public benchmarks in the Clinical AI Benchmark Atlas showed the median benchmark contained no documentation requests and shared only 31% of real use's task mix, less than an even spread across task categories, leading the authors to conclude benchmark scores say little about clinical AI performance on most of the work it receives.

by read1 min views5 publishedOct 9, 2026

arXiv:2610.11069v1 Announce Type: new Abstract: Large language model (LLM) assistants are being deployed to clinicians across health systems, and judgments about their readiness rest largely on benchmark scores, most of them derived from examination questions or curated cases. A benchmark predicts performance in deployment only to the extent that its items resemble real use, yet whether benchmarks reflect the work these systems receive has rarely been measured. Here we analyze 127,833 queries sent by 6,342 physicians, advanced practice providers and nurses in 35 specialties to an institutional assistant during an eight-month roll-out. We characterize each query with RCQ-Map, a clinician-validated framework grounded in taxonomies of clinical questions and of LLM evaluation, which records its task, intent, answerability, missing information and potential harm. Documentation and administration (36.2%) and knowledge retrieval (28.9%) made up nearly two-thirds of use, and diagnosis 3.7%; more than a third of queries could not be answered well as posed. Applying RCQ-Map to 58 public benchmarks drawn from major evaluation suites and frontier model reports, which we assemble into the Clinical AI Benchmark Atlas, showed that the median benchmark contained no documentation requests and shared 31% of the task mix of real use, less than an even spread across task categories would. Benchmarks in suites designed to resemble clinical practice were individually no closer to real use than those used in frontier model reports. Benchmark scores therefore say little about how clinical AI performs on most of the work it is actually given, and evaluation should be matched to real clinical use.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/clinician-use-of-lan…] indexed:0 read:1min 2026-10-09 · —