# Think-Probe-Respond: Improving Large Language Models as Judges of Research Idea Novelty

> Source: <https://www.machinebrief.com/news/think-probe-respond-improving-large-language-models-as-judge-zpbj>
> Published: 2026-08-27 04:00:00+00:00

arXiv:2608.25660v1 Announce Type: new
Abstract: Automated novelty judgment can accelerate scientific discovery by enabling efficient evaluation, refinement, and comparison of research ideas. While large language models are increasingly adopted for this task, we investigate a previously overlooked limitation in their judgment capabilities: despite generating reasoning rationales that closely mirror those of human experts, their final novelty judgments often diverge substantially. We demonstrate that this miscalibration stems from a systematic bias towards judging ideas as "medium novel". To mitigate this, we propose Think-Probe-Respond (TPR), a lightweight approach that probes latent novelty judgments from hidden states during the reasoning phase and uses the probed judgments to condition the final response. Across strong baselines, TPR improves novelty judgment performance by 22.30% and successfully mitigates the prevalent "medium novelty" bias.
