Performance of large language models in the optical diagnosis of colorectal polyps A retrospective study using the PRIME dataset found that multimodal large language models (MLLMs) achieved F1 scores above 0.9 for distinguishing neoplastic from non-neoplastic colorectal polyps, but none met ESGE standards for clinical use. Google Gemini 2.5 Pro achieved the highest F1 scores for invasive vs. non-invasive polyps (0.560) and low- vs. high-grade adenoma (0.492), while Claude Opus 4 and GPT-5 tied for the highest percent correct score (41.7%) using the Paris classification. The authors call for prospective multicenter trials and human-in-the-loop workflows before deployment. arXiv:2608.07543v1 Announce Type: new Abstract: Background and Study Aims: Accurate optical diagnosis of colorectal polyps guides resection strategy and surveillance, with multimodal large language models MLLMs showing potential for image-based diagnosis. We aimed to evaluate the diagnostic accuracy of MLLMs in classifying colorectal polyps and predicting histology. Methods: We conducted a retrospective diagnostic performance study using the PRIME dataset, a curated set of white light and narrow-band imaging NBI images. We evaluated Claude Opus 4, Google Gemini 2.5 Pro, GPT-o3, GPT-4o, and GPT-5. For Paris, Narrow-band Imaging Colorectal Endoscopic NICE , and predicted histology, we calculated F1 scores, percent correct scores, and accuracy of each MLLM compared to expert responses for 132 cases. Cochran's Q and McNemar's Test were used to determine differences between predicted values of each MLLM. Results: The F1 scores among MLLMs were 0.9 for all models for neoplastic vs. non-neoplastic polyps. Gemini 2.5 Pro demonstrated the highest F1 scores for invasive vs. non-invasive polyps and low- vs. high-grade adenoma, at 0.560 and 0.492 respectively. Claude Opus 4 and GPT-5 had statistically significantly higher percent correct scores than other MLLMs at 41.7%, using Paris classification. Conclusions: Claude Opus 4 and Gemini 2.5 Pro showed the highest accuracy in differentiating polyp subtypes, performing closest to expert consensus. Sensitivity and specificity, however, did not meet ESGE standards, highlighting the need for prospective multicenter trials and the design of human-in-the-loop workflows before clinical deployment.