cd /news/large-language-models/large-language-models-show-metacogni… · home › topics › large-language-models › article
[ARTICLE · art-100801] src=arxiv.org ↗ pub= topic=large-language-models verified=true sentiment=· neutral

Large Language Models Show Metacognitive Sensitivity in Medical Reasoning

A new arXiv study (2608.14552v1) found that OpenAI's gpt-4.1-nano model achieved 93.5% diagnostic accuracy and 78.4% mean confidence across 135 trials in a controlled medical benchmark, with confidence tracking evidence quality and uncertainty, indicating partial metacognitive sensitivity. However, errors clustered in moderate, conflicting Alzheimer-type neurocognitive disorder cases, where the model retained more confidence than accuracy justified, highlighting localized calibration failure.

read1 min views17 publishedAug 18, 2026

arXiv:2608.14552v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly evaluated and used in medicine, but clinical usefulness depends on answer accuracy and whether confidence tracks evidence quality and uncertainty. We developed a controlled, psychophysics-inspired clinical benchmark to test diagnostic choice and confidence behavior in a medical LLM. The benchmark focused on probable Alzheimer-type neurocognitive disorder (AT-NCD) versus depression-related cognitive impairment (DRCI). We generated 45 synthetic vignettes varying evidence strength, conflicting evidence, and missing information. Each vignette was presented under three prompt variants, yielding 135 trials. In a pilot run with gpt-4.1-nano, all trials produced valid structured outputs. Across forced-choice trials, diagnostic accuracy was 93.5%, mean confidence was 78.4%, and AUROC2 was 0.876. Confidence increased with evidence distance from the diagnostic boundary, decreased when information was missing, and remained higher on correct than incorrect trials after adjustment for evidence strength and prompt format. These findings indicate partial metacognitive sensitivity rather than globally uninformative confidence. However, errors clustered in moderate, conflicting AT-NCD cases, where the model shifted toward DRCI and retained more confidence than empirical accuracy justified. Model comparison suggested that confidence quality should be measured directly rather than inferred from benchmark accuracy or model capability alone. This study establishes a reproducible framework for evaluating evidence sensitivity, metacognitive sensitivity, and localized calibration failure in medical LLMs.

── more in #large-language-models 4 stories · sorted by recency
── more on @openai 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/large-language-model…] indexed:0 read:1min 2026-08-18 · —