Do large language models scrutinise what they review? A multimodal audit of scoring calibration, error detection, and author-identity effects
A study of two multimodal large language models, Qwen2.5-VL-72B and Pixtral-Large-124B, found that they assigned peer-review scores of 7.0 to 8.1 to 165 submissions to the 2026 International Conference on Learning Repres…