cd /news/artificial-intelligence/ai-models-invented-medical-diagnoses… · home topics artificial-intelligence article
[ARTICLE · art-76233] src=thecoinheadlines.com ↗ pub= topic=artificial-intelligence verified=true sentiment=↓ negative

AI models invented medical diagnoses in nearly 1 in 5 no-scan tests, study finds

AI models from Anthropic, OpenAI, and Google invented unsupported medical diagnoses in 18% of tests where the required scan was intentionally omitted, a Carnegie Mellon University study found. The research examined Claude Opus 4.7, GPT-5.4, and Gemini 3.1 Pro across 11,700 prompts involving chest X-rays, brain MRI scans, and skin mole images, with models correctly declining diagnosis in about 82% of responses. Study author Siddharth Vohra said healthcare providers should test AI systems for demographic sensitivity and examine behavior with incomplete patient information before clinical use.

read3 min views1 publishedJul 28, 2026
AI models invented medical diagnoses in nearly 1 in 5 no-scan tests, study finds
Image: Thecoinheadlines (auto-discovered)

Leading AI systems produced unsupported medical diagnoses in nearly one in five tests where the required scan had been intentionally omitted, according to a Carnegie Mellon University study that raises fresh concerns about patients relying on chatbots for medical advice.

The research examined Claude Opus 4.7, OpenAI’s GPT-5.4 and Google’s Gemini 3.1 Pro across 11,700 prompts involving chest X-rays, brain MRI scans and images of skin moles.

Although each prompt referred to an image, none was actually provided, allowing researchers to test whether the AI models would acknowledge the missing evidence or generate an answer regardless.

Across the experiment, the models correctly declined to offer a diagnosis in about 82% of responses, but supplied a disease in the remaining 18%, despite having no medical image to examine.

Demographics and word choice changed the answers

Siddharth Vohra, a master’s student at Carnegie Mellon’s Robotics Institute and the study’s sole author, tested 12 simulated patient profiles using different combinations of age, sex and race, alongside a control prompt containing no demographic details.

Rather than failing consistently, the three systems showed different patterns of behavior. GPT-5.4 generated diagnoses across all 36 patient-and-scan combinations, while Claude refused more frequently but repeatedly produced the same diseases in certain demographic scenarios. Gemini declined most often, although the diagnoses it did provide still changed when patient characteristics were altered.

In one test, Claude identified melanoma in 94% of responses involving a 65-year-old white man asking about a missing image of a skin mole. GPT-5.4, meanwhile, named sarcoidosis in 77 out of 100 chest X-ray prompts involving a young Black patient, even though no X-ray was attached.

The findings were also highly sensitive to small changes in wording. When researchers replaced “mole” with “lesion” for the same patient profile, Claude shifted from identifying melanoma in 94% of responses to refusing the task every time, while GPT-5.4’s behavior changed little.

Refusal language did not always stop a diagnosis

Refusal language did not always stop a diagnosis

The study also found that some models appeared to refuse a diagnosis in their written explanation while still placing a disease name inside a structured field that could be read by hospital software or another automated system.

In Claude’s weakest-performing test, 62 of 94 invented diagnoses were accompanied by language resembling a refusal, creating a risk that a human reader might believe the model had safely declined while an unsupported diagnosis continued through the technical workflow.

Study highlights limits of fluent medical answers

The study relied on simulated patient profiles rather than real clinical cases. However, its findings still exposed a key medical AI risk: fluent, authoritative answers can appear clinically sound even when the model lacks the evidence needed to support them.

Vohra said healthcare providers should test AI systems for demographic sensitivity and examine how they behave when patient information is incomplete before introducing them into clinical settings.

Suggested safeguards include forcing diagnostic fields to remain blank when an image is missing and comparing model responses generated with and without the underlying scan.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @carnegie mellon university 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/ai-models-invented-m…] indexed:0 read:3min 2026-07-28 ·