{"slug": "ai-models-invented-medical-diagnoses-in-nearly-1-in-5-no-scan-tests-study-finds", "title": "AI models invented medical diagnoses in nearly 1 in 5 no-scan tests, study finds", "summary": "AI models from Anthropic, OpenAI, and Google invented unsupported medical diagnoses in 18% of tests where the required scan was intentionally omitted, a Carnegie Mellon University study found. The research examined Claude Opus 4.7, GPT-5.4, and Gemini 3.1 Pro across 11,700 prompts involving chest X-rays, brain MRI scans, and skin mole images, with models correctly declining diagnosis in about 82% of responses. Study author Siddharth Vohra said healthcare providers should test AI systems for demographic sensitivity and examine behavior with incomplete patient information before clinical use.", "body_md": "Leading AI systems produced unsupported medical diagnoses in nearly one in five tests where the required scan had been intentionally omitted, according to a Carnegie Mellon University study that raises fresh concerns about patients relying on chatbots for medical advice.\n\nThe research examined Claude Opus 4.7, OpenAI’s GPT-5.4 and Google’s Gemini 3.1 Pro across 11,700 prompts involving chest X-rays, brain MRI scans and images of skin moles.\n\nAlthough each prompt referred to an image, none was actually provided, allowing researchers to test whether the AI models would acknowledge the missing evidence or generate an answer regardless.\n\nAcross the experiment, the models correctly declined to offer a diagnosis in about 82% of responses, but supplied a disease in the remaining 18%, despite having no medical image to examine.\n\n**Demographics and word choice changed the answers**\n\nSiddharth Vohra, a master’s student at Carnegie Mellon’s Robotics Institute and the study’s sole author, tested 12 simulated patient profiles using different combinations of age, sex and race, alongside a control prompt containing no demographic details.\n\nRather than failing consistently, the three systems showed different patterns of behavior. GPT-5.4 generated diagnoses across all 36 patient-and-scan combinations, while Claude refused more frequently but repeatedly produced the same diseases in certain demographic scenarios. Gemini declined most often, although the diagnoses it did provide still changed when patient characteristics were altered.\n\nIn one test, Claude identified melanoma in 94% of responses involving a 65-year-old white man asking about a missing image of a skin mole. GPT-5.4, meanwhile, named sarcoidosis in 77 out of 100 chest X-ray prompts involving a young Black patient, even though no X-ray was attached.\n\nThe findings were also highly sensitive to small changes in wording. When researchers replaced “mole” with “lesion” for the same patient profile, Claude shifted from identifying melanoma in 94% of responses to refusing the task every time, while GPT-5.4’s behavior changed little.\n\n**Refusal language did not always stop a diagnosis**\n\n**Refusal language did not always stop a diagnosis**\n\nThe study also found that some models appeared to refuse a diagnosis in their written explanation while still placing a disease name inside a structured field that could be read by hospital software or another automated system.\n\nIn Claude’s weakest-performing test, 62 of 94 invented diagnoses were accompanied by language resembling a refusal, creating a risk that a human reader might believe the model had safely declined while an unsupported diagnosis continued through the technical workflow.\n\n**Study highlights limits of fluent medical answers**\n\nThe study relied on simulated patient profiles rather than real clinical cases. However, its findings still exposed a key medical AI risk: fluent, authoritative answers can appear clinically sound even when the model lacks the evidence needed to support them.\n\nVohra said healthcare providers should test AI systems for demographic sensitivity and examine how they behave when patient information is incomplete before introducing them into clinical settings.\n\nSuggested safeguards include forcing diagnostic fields to remain blank when an image is missing and comparing model responses generated with and without the underlying scan.", "url": "https://wpnews.pro/news/ai-models-invented-medical-diagnoses-in-nearly-1-in-5-no-scan-tests-study-finds", "canonical_source": "https://thecoinheadlines.com/tech-and-ai/ai-models-invented-medical-diagnoses-in-nearly-1-in-5-no-scan-tests-study-finds/article-27512/", "published_at": "2026-07-28 01:54:19+00:00", "updated_at": "2026-07-28 02:02:36.673019+00:00", "lang": "en", "topics": ["artificial-intelligence", "ai-safety", "ai-ethics", "ai-research", "large-language-models"], "entities": ["Carnegie Mellon University", "Claude Opus 4.7", "OpenAI", "GPT-5.4", "Google", "Gemini 3.1 Pro", "Siddharth Vohra", "Robotics Institute"], "alternates": {"html": "https://wpnews.pro/news/ai-models-invented-medical-diagnoses-in-nearly-1-in-5-no-scan-tests-study-finds", "markdown": "https://wpnews.pro/news/ai-models-invented-medical-diagnoses-in-nearly-1-in-5-no-scan-tests-study-finds.md", "text": "https://wpnews.pro/news/ai-models-invented-medical-diagnoses-in-nearly-1-in-5-no-scan-tests-study-finds.txt", "jsonld": "https://wpnews.pro/news/ai-models-invented-medical-diagnoses-in-nearly-1-in-5-no-scan-tests-study-finds.jsonld"}}