STAT+: Clinical chatbots are taking medicine by storm. Should doctors trust them? A study published in Nature Medicine in June by researchers from NYU Langone Health found that clinical AI models from OpenEvidence and UpToDate Expert AI performed worse than general large language models on clinical question-answering tasks, sparking intense debate in the health AI community. Kaiser Permanente's vice president of AI and emerging technologies called the paper's reactions unprecedented. Hundreds of thousands of U.S. doctors use clinical large language models, pitched by companies like OpenEvidence, Doximity, and UpToDate as an antidote to the dangers of hallucination-prone generalist models from Big Tech. Yet few studies have pitted them against each other — and this summer brought a high-profile head-to-head. Researchers from NYU Langone Health had tested general and clinical models, including OpenEvidence and UpToDate Expert AI, on three sets of clinical questions. The findings, published in Nature Medicine in June: The clinical AI performed worse than the general models. The results rang out like a gunshot. “I’ve never seen a single paper trigger the kind of reactions this one has in the health AI community,” wrote https://www.linkedin.com/posts/danielayang general-purpose-large-language-models-outperform-share-7472786182888771584-17xY/?utm source=share&utm medium=member android&rcm=ACoAABwDShQBLcuRrEDaKuPvyY-y-FMOt2B BHc Kaiser Permanente’s vice president of AI and emerging technologies on LinkedIn. The paper’s findings, like all science, are subject to interpretation and debate — but many online reactions treated them more like a clear victory for general frontier models. This article is exclusive to STAT+ subscribers Unlock this article — and get additional analysis of the technologies disrupting health care — by subscribing to STAT+. Already have an account? Log in /login/ View All Plans https://www.statnews.com/stat-plus/