04:00
2026-09-30
arxiv.org
large-language-models
Binarization Flattens the Score Space
A new arXiv paper (2609.35797v1) finds that collapsing large language model judge rewards to pass/fail {0, 1} hides proportional changes in how well a response meets each criterion, and recommends kee…