cd /news/artificial-intelligence/auditing-medical-vision-language-mod… · home topics artificial-intelligence article
[ARTICLE · art-91451] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Auditing Medical Vision-Language Models on Chest Radiographs: Estimating Reference Agreement Across Institutions

A study of three generative vision-language models on chest radiographs found that reference agreement does not transfer reliably across institutions, with the best estimation strategy achieving a mean Brier score of 0.0853 versus 0.1083 for adaptive selection, and a nominal 95% coverage interval achieving only 87.0% at the hardest institution. The authors, from an unspecified research group, evaluated more than 345,000 finding-level predictions across three institutional corpora and six findings, concluding that agreement must be re-evaluated per site and per interface.

read1 min views1 publishedAug 11, 2026

arXiv:2608.07550v1 Announce Type: new Abstract: Vision-language models return structured chest-radiograph findings through interfaces exposing no confidence score, so a receiving institution cannot read off how far to trust an individual judgment. Whether agreement with an institution's reference standard transfers across sites, findings, prediction directions and question formats is largely unmeasured. We evaluated three generative vision-language models on three institutional chest-radiograph corpora and six findings under two elicitation protocols, comprising more than 345,000 finding-level predictions, and estimated finding-by-direction reference agreement at a receiving institution from a small budget of local labels. Estimation strategies were then stress-tested under repeated strict institution-held-out evaluation. Under evaluation excluding the receiving institution from development entirely, adaptive selection among the seven estimators that design admits did not improve on simple fixed alternatives: it achieved a mean Brier score of 0.1083, against 0.0853 for always using a Beta-Binomial empirical-Bayes estimator and 0.0855 for a target-only logistic model. Those two differ by 0.0003, less than this family's own sensitivity to a change of solver version, and each leads in about half the settings, so no default can be recommended. Their advantage over estimators pooling across institutions was concentrated at one site and not confirmatory once clustered by institution, and a plug-in empirical-Bayes posterior-predictive count interval at a nominal 95% level covered 87.0%, less at the hardest institution. Reference agreement therefore has to be re-evaluated per site and per interface; these results concern agreement with institutional labels, not clinical correctness.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/auditing-medical-vis…] indexed:0 read:1min 2026-08-11 ·