Auditing Medical Vision-Language Models on Chest Radiographs: Estimating Reference Agreement Across Institutions
A study of three generative vision-language models on chest radiographs found that reference agreement does not transfer reliably across institutions, with the best estimation strategy achieving a mean Brier score of 0.0…