cd /news/artificial-intelligence/medqa-mm-shortcuts-behind-medical-vi… · home topics artificial-intelligence article
[ARTICLE · art-121124] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

MedQA-MM: Shortcuts Behind Medical Visual Reasoning

A new arXiv preprint (arXiv:2609.03261v1) finds that medical multimodal AI benchmarks overstate visual reasoning because correct answers can be derived from non-visual cues, a phenomenon the authors call 'reasoning inflation.' Across six medical multimodal MCQ datasets, full-input accuracy was 62.63%, but text-only and options-only settings achieved 53.96% and 29.71%, respectively. The authors introduce MedQA-MM, a 1,000-item shortcut-mitigated subset where text-only and options-only accuracy drop to 5.21% and 12.33%, underscoring the need for route-level evidence in medical image-reasoning claims.

read1 min views1 publishedSep 4, 2026

arXiv:2609.03261v1 Announce Type: new Abstract: A benchmark score credits final answers, but not the route by which an item can be answered. In medical multimodal multiple-choice questions (MCQs), this distinction matters because a correct answer can be supported by the intended image finding or by benchmark-preserved cues in the wording of answers, non-visual clinical text, visible image text, artificial annotations, or device/context artifacts. We call the resulting score-level overinterpretation reasoning inflation. Here, a route is an observable input path that can support answer selection, not a claim about the model's hidden cognition. Across six medical multimodal MCQ datasets, we separate candidate cues from behavioral evidence through prompt- and image-side audits, modality ablations, and matched repairs that preserve the medical target and answer key. In a 13-configuration open-model panel, full-input accuracy is 62.63%, while text-only and options-only settings achieve 53.96% and 29.71%, respectively. Removing length-gap, absolute/conspicuous, and spatial/prepositional cues lowers accuracy by 6.58, 3.50, and 4.77 percentage points. We also construct MedQA-MM, a 1,000-item shortcut-mitigated subset, where text-only and options-only accuracy fall to 5.21% and 12.33%. This does not imply that models never use images; it shows that medical image-reasoning claims require route-level evidence.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/medqa-mm-shortcuts-b…] indexed:0 read:1min 2026-09-04 ·