cd /news/large-language-models/binarization-flattens-the-score-spac… · home › topics › large-language-models › article
[ARTICLE · art-142228] src=arxiv.org ↗ pub= topic=large-language-models verified=true sentiment=· neutral

Binarization Flattens the Score Space

A new arXiv paper (2609.35797v1) finds that collapsing large language model judge rewards to pass/fail {0, 1} hides proportional changes in how well a response meets each criterion, and recommends keeping at least three grades — for example rewarding {0, 0.5, 1} for fully, partially, or not met. On MATH and SciBench outputs from one seven-criterion judge, all 14 constructed criterionwise stretches were invisible after binarization but visible with three grades, and at n=1,024 a test given both population laws had at least 96.5% power at a 1.5x stress. The authors caution that verdicts alone remain insufficient, since arbitrary within-grade changes and fixed-covariance, loading-aligned mean shifts — the signature of a sycophancy-shaped lift the panel reads as competence — are still indistinguishable from genuine improvement, and the shared-factor reference approximation fit MATH and SciBench but not HealthBench.

by read1 min views1 publishedSep 30, 2026
arXiv:2609.35797v1 Announce Type: new 
Abstract: Large language model (LLM) judges are often used as rewards to train policies on objectives that deterministic verifiers cannot capture. However, these rewards are often collapsed to pass/fail ({0, 1}), which reports the verdict but not how well a response met each criterion. We model each pass/fail verdict as a score on an unreported scale, compared with one cutoff. A stretch of that scale moves every score proportionally toward or away from the cutoff, but never across it, so no verdict changes. A policy is therefore free to apply any stretch without changing anything the panel reports. Under a joint-Gaussian model, a third grade adds a second threshold and removes this affine stretch ambiguity. On MATH and SciBench outputs from one seven-criterion judge, all 14 constructed criterionwise stretches were invisible after binarization but visible with three grades. At $n=1{,}024$, a test given both population laws had at least 96.5% power at a $1.5\times$ stress. Retaining grades closes one blind spot created by binarization, but verdicts alone remain insufficient as some changes are still indistinguishable from genuine improvement. These include arbitrary within-grade changes and fixed-covariance, -aligned mean shifts -- the signature of a sycophancy-shaped lift the panel reads as competence. The shared-factor reference approximation fit MATH and SciBench but not HealthBench, delineating its empirical scope. We recommend keeping at least three grades (for example, asking the judge whether each criterion is fully, partially, or not met and rewarding {0, 0.5, 1}), and externally validating gains along the remaining direction, which no finer scale removes.
── more in #large-language-models 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/binarization-flatten…] indexed:0 read:1min 2026-09-30 · —