cd /news/large-language-models/do-large-language-models-scrutinise-… · home topics large-language-models article
[ARTICLE · art-117321] src=arxiv.org ↗ pub= topic=large-language-models verified=true sentiment=· neutral

Do large language models scrutinise what they review? A multimodal audit of scoring calibration, error detection, and author-identity effects

A study of two multimodal large language models, Qwen2.5-VL-72B and Pixtral-Large-124B, found that they assigned peer-review scores of 7.0 to 8.1 to 165 submissions to the 2026 International Conference on Learning Representations, compared with human mean scores of 3.4 to 6.8. The models detected only 12.1% of 145 inserted errors under natural prompting, rising to 22.2% with a verification instruction, and 78% of errors remained undetected; providing figures reduced error detection and increased scores, while author identity had no effect.

read1 min views3 publishedSep 1, 2026

arXiv:2608.28626v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to generate peer reviews, prompting examination of their capacity for critical evaluation. This study evaluates two multimodal LLMs, Qwen2.5-VL-72B and Pixtral-Large-124B, as reviewers across 165 submissions to the 2026 International Conference on Learning Representations, a venue that postdates both models' training cutoffs. Manuscripts were presented to both models with author identities blinded, replaced with high-prestige affiliations, or replaced with low-prestige affiliations, and in either text-only or text-with-figure format. Additionally, 145 verifiably detectable errors were inserted into 55 manuscripts to assess error identification under natural and verification-oriented prompts. Across all manuscript groups, including rejected submissions, LLM scores ranged from 7.0 to 8.1, whereas human mean scores ranged from 3.4 to 6.8. The models detected 12.1% of the verified errors under natural prompting, and a one-sentence verification instruction increased detection to 22.2%; however, 78% of the errors remained undetected. Providing figures reduced error detection while increasing review scores. No visual error was reliably verified against its corresponding figure, and half of the text-only reviews described figures that were not provided. Author identity did not influence either review scores or error detection. LLM editorial decisions exactly matched those produced by simple score averaging.

── more in #large-language-models 4 stories · sorted by recency
── more on @qwen2.5-vl-72b 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/do-large-language-mo…] indexed:0 read:1min 2026-09-01 ·