04:00
2026-09-12
machinebrief.com
large-language-models
NovGauge: A Fine-Grained Benchmark for Diagnosing LLMs' Capability in Paper Novelty Assessment
A new benchmark called NovGauge, built from 619 paper pairs and 50 multi-paper sets drawn from ICLR reviewer overlap claims and survey co-citations, finds that large language models remain unreliable …