{"slug": "do-large-language-models-scrutinise-what-they-review-a-multimodal-audit-of-error", "title": "Do large language models scrutinise what they review? A multimodal audit of scoring calibration, error detection, and author-identity effects", "summary": "A study of two multimodal large language models, Qwen2.5-VL-72B and Pixtral-Large-124B, found that they assigned peer-review scores of 7.0 to 8.1 to 165 submissions to the 2026 International Conference on Learning Representations, compared with human mean scores of 3.4 to 6.8. The models detected only 12.1% of 145 inserted errors under natural prompting, rising to 22.2% with a verification instruction, and 78% of errors remained undetected; providing figures reduced error detection and increased scores, while author identity had no effect.", "body_md": "arXiv:2608.28626v1 Announce Type: new\nAbstract: Large language models (LLMs) are increasingly used to generate peer reviews, prompting examination of their capacity for critical evaluation. This study evaluates two multimodal LLMs, Qwen2.5-VL-72B and Pixtral-Large-124B, as reviewers across 165 submissions to the 2026 International Conference on Learning Representations, a venue that postdates both models' training cutoffs. Manuscripts were presented to both models with author identities blinded, replaced with high-prestige affiliations, or replaced with low-prestige affiliations, and in either text-only or text-with-figure format. Additionally, 145 verifiably detectable errors were inserted into 55 manuscripts to assess error identification under natural and verification-oriented prompts. Across all manuscript groups, including rejected submissions, LLM scores ranged from 7.0 to 8.1, whereas human mean scores ranged from 3.4 to 6.8. The models detected 12.1\\% of the verified errors under natural prompting, and a one-sentence verification instruction increased detection to 22.2\\%; however, 78\\% of the errors remained undetected. Providing figures reduced error detection while increasing review scores. No visual error was reliably verified against its corresponding figure, and half of the text-only reviews described figures that were not provided. Author identity did not influence either review scores or error detection. LLM editorial decisions exactly matched those produced by simple score averaging.", "url": "https://wpnews.pro/news/do-large-language-models-scrutinise-what-they-review-a-multimodal-audit-of-error", "canonical_source": "https://arxiv.org/abs/2608.28626", "published_at": "2026-09-01 04:00:00+00:00", "updated_at": "2026-09-01 04:24:49.637704+00:00", "lang": "en", "topics": ["large-language-models", "ai-research", "ai-ethics"], "entities": ["Qwen2.5-VL-72B", "Pixtral-Large-124B", "International Conference on Learning Representations"], "alternates": {"html": "https://wpnews.pro/news/do-large-language-models-scrutinise-what-they-review-a-multimodal-audit-of-error", "markdown": "https://wpnews.pro/news/do-large-language-models-scrutinise-what-they-review-a-multimodal-audit-of-error.md", "text": "https://wpnews.pro/news/do-large-language-models-scrutinise-what-they-review-a-multimodal-audit-of-error.txt", "jsonld": "https://wpnews.pro/news/do-large-language-models-scrutinise-what-they-review-a-multimodal-audit-of-error.jsonld"}}