{"slug": "think-probe-respond-improving-large-language-models-as-judges-of-research-idea", "title": "Think-Probe-Respond: Improving Large Language Models as Judges of Research Idea Novelty", "summary": "Researchers propose Think-Probe-Respond (TPR), a method that improves large language models' ability to judge research idea novelty by probing latent judgments from hidden states, achieving a 22.30% performance improvement over strong baselines. The study, released on arXiv (2608.25660v1), identifies a systematic bias toward judging ideas as 'medium novel' and demonstrates that TPR mitigates this bias.", "body_md": "arXiv:2608.25660v1 Announce Type: new\nAbstract: Automated novelty judgment can accelerate scientific discovery by enabling efficient evaluation, refinement, and comparison of research ideas. While large language models are increasingly adopted for this task, we investigate a previously overlooked limitation in their judgment capabilities: despite generating reasoning rationales that closely mirror those of human experts, their final novelty judgments often diverge substantially. We demonstrate that this miscalibration stems from a systematic bias towards judging ideas as \"medium novel\". To mitigate this, we propose Think-Probe-Respond (TPR), a lightweight approach that probes latent novelty judgments from hidden states during the reasoning phase and uses the probed judgments to condition the final response. Across strong baselines, TPR improves novelty judgment performance by 22.30% and successfully mitigates the prevalent \"medium novelty\" bias.", "url": "https://wpnews.pro/news/think-probe-respond-improving-large-language-models-as-judges-of-research-idea", "canonical_source": "https://www.machinebrief.com/news/think-probe-respond-improving-large-language-models-as-judge-zpbj", "published_at": "2026-08-27 04:00:00+00:00", "updated_at": "2026-08-27 05:19:38.668708+00:00", "lang": "en", "topics": ["large-language-models", "ai-research", "ai-tools"], "entities": ["arXiv", "Think-Probe-Respond"], "alternates": {"html": "https://wpnews.pro/news/think-probe-respond-improving-large-language-models-as-judges-of-research-idea", "markdown": "https://wpnews.pro/news/think-probe-respond-improving-large-language-models-as-judges-of-research-idea.md", "text": "https://wpnews.pro/news/think-probe-respond-improving-large-language-models-as-judges-of-research-idea.txt", "jsonld": "https://wpnews.pro/news/think-probe-respond-improving-large-language-models-as-judges-of-research-idea.jsonld"}}