{"slug": "ai-tool-may-help-doctors-but-sample-size-too-small", "title": "AI tool may help doctors, but sample size too small", "summary": "A randomized trial of nearly 10,000 patient encounters across 16 Kenyan primary clinics found that an AI tool called AI Consult, powered by OpenAI's GPT-4o, helped clinicians produce better diagnoses and treatment plans, but did not find statistically significant evidence of improved patient outcomes. The study, published in Nature Medicine and funded by the Gates Foundation, showed a 23% decrease in treatment failures such as death or unresolved symptoms, but the reduction was not statistically significant due to low baseline rates. Dr. Bilal Mateen, co-author and chief AI officer at PATH, said a trial would need about 139,000 people to detect a meaningful difference.", "body_md": "# This AI tool promises a 'second pair of eyes' to clinicians. Did patients benefit?\n\nBy Joseph Kim\n\nThursday, July 23, 2026 • 7:36 AM EDT\n\nA 4-month-old boy comes into the clinic with a fever and a stuffy nose. Vyonne Njeri thinks it's just a cold. Then a yellow box pops up on her computer telling her to check his heart — because his heart rate is elevated. Njeri is a registered clinical officer in Nairobi, Kenya; she sees patients on her own like a nurse practitioner. When she listens with a stethoscope she hears a whoosh — a sign that the infant could have a congenital heart defect.** **\n\n\"That's something I would have missed on any other day,\" Njeri says of the visit a few months ago. \"That child would have just gone home.\" She credits an AI tool that double checks her work for helping her. Njeri referred him to a specialist who confirmed the diagnosis and started him on medications; the child may have to undergo surgery.\n\nThe tool, called AI Consult, was tested in a randomized trial of nearly 10,000 patient encounters across 16 Kenyan primary clinics operated by Penda Health, and the results were published this summer in [Nature Medicine](https://www.nature.com/articles/s41591-026-04503-6). Half the clinical officers typed their notes into a computer with OpenAI's GPT-4o checking their electronic notes in the background.\n\nGPT-4o is a large language model, or LLM — the kind of AI that powers chatbots like ChatGPT. They've been trained on much of the internet and can generate text in response to instructions.\n\nThe other half used a computer for note-taking without the AI tool.\n\nHere's how the AI program works. It provides three prompts — green, yellow and red, like a traffic light. The AI looks for possible problems in a patient's care or gaps in the notes. Green means everything is OK. Yellow means \"click me\" — the AI found a small problem in the notes and offers some guidance. Red pops up if the AI finds a critical concern based on the notes and alerts the clinician, who may need to act promptly.\n\nFor Njeri, these interventions offered welcome feedback: \"You need to check on that\" or \"you're on the right track.\"\n\nAn independent panel of six Kenyan family physicians graded the notes. They judged that clinicians for whom AI Consult served as what Njeri calls a \"second pair of eyes\" produced better diagnoses and treatment plans.\n\nAnd it cost 4 cents per patient.\n\nThe study captures both the potential of AI to improve healthcare and also the challenge of proving it.\n\nBut although health workers found this AI checkup helpful, the question of patient benefit was left unanswered. The study did not find evidence of improved outcomes.\n\nThere was a 23% decrease in treatment failures such as death or unresolved symptoms, but the decrease was not statistically significant because the number of such cases was low to start with.\n\nTreatment failures \"are just too rare in primary care,\" says [ Dr. Bilal Mateen](https://www.path.org/who-we-are/leadership/bilal-mateen-mph-mbbs/), a co-author of the study and chief AI officer at PATH, the global health nonprofit that sponsored the trial. A trial would need about 139,000 people to detect a meaningful difference, he says.\n\nThe trial was funded by the Gates Foundation, which provides financial support to NPR for its global health team; NPR is solely responsible for all content.\n\nNonetheless, in the continuing exploration of the role that AI can play in medical care, particularly in lower-resource countries — there is praise for the study.** **[ Dr. Jonathan Chen](https://profiles.stanford.edu/jonc101), an associate professor of medicine and biomedical data science at Stanford University, who was not involved with the study, says, \"this is an important study as it goes beyond just running simulated test[s] with AI systems.\" The study is among the first trials to compare the role of this type of AI's potential as an aid to primary care.\n\nMateen says that prompting healthcare providers with what he calls \"information-as-an-intervention\" will have an \"important incremental effect.\"\n\nNjeri shares that sentiment: \"It helps us remember protocols and new guidelines,\" she says. In a country where health resources are not comparable to wealthier nations, she says her team may see five or six patients an hour with a wide range of conditions — often without specialists for backup. \"Having this tool really adds value,\" she says.\n\nAs for the AI recommendations, Njeri says she finds them helpful about half the time — the other half the recommendations aren't as helpful but they are seldom wrong.\n\n\"But I still feel it's nice to have it — to give you that second thumbs-up or thumbs-down, like having a superior who says, 'you could do better there' or 'you're doing okay,'\" she says. Most of the AI's recommendations were rated safe and appropriate by the expert panel.\n\nMateen says he is working toward implementing the tool globally.\n\nThe paper did not look into whether AI could increase access to care, but Chen is most excited about that possibility. As AI gets better, he says it could in fact create a draft for the healthcare worker's notes for editing — which could enable them to save time so they could see more patients. Or the AI tool could even interact with patients themselves.\n\nHe says, \"timely and consistent access to medical care\" is where AI may have the biggest impact.\n\nMateen says he is working toward bringing this tool to healthcare systems around the globe. But there are still concerns about fledgling AI tools. Healthcare AI researcher [ Dr. Nicholas Okumu](https://drnicholasokumu.com/), an orthopedic surgeon with Kenyatta National Hospital who wasn't involved with the study, worries new AI systems could make mistakes or give bad advice, leading to grave errors. He warns \"even AI that's approved can still cause harm, so oversight has to stay active.\"", "url": "https://wpnews.pro/news/ai-tool-may-help-doctors-but-sample-size-too-small", "canonical_source": "https://text.npr.org/g-s1-134929", "published_at": "2026-07-24 03:03:22+00:00", "updated_at": "2026-07-24 03:22:35.672349+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-products", "ai-research", "ai-policy"], "entities": ["AI Consult", "OpenAI", "GPT-4o", "Penda Health", "Nature Medicine", "PATH", "Gates Foundation", "Bilal Mateen"], "alternates": {"html": "https://wpnews.pro/news/ai-tool-may-help-doctors-but-sample-size-too-small", "markdown": "https://wpnews.pro/news/ai-tool-may-help-doctors-but-sample-size-too-small.md", "text": "https://wpnews.pro/news/ai-tool-may-help-doctors-but-sample-size-too-small.txt", "jsonld": "https://wpnews.pro/news/ai-tool-may-help-doctors-but-sample-size-too-small.jsonld"}}