04:00
2026-09-11
arxiv.org
large-language-models
Rethinking Verbalized Confidence for LLM-as-a-Judge: A Compatibility Shift on Post-2025 Proprietary Models
A new arXiv paper (2609.10996v1) reports that verbalized confidence has become the more robust soft-scoring mechanism for LLM-as-a-Judge on post-2025 proprietary models, reversing the standard advice …