cd /news/machine-learning/truncate-bad-upweight-good-bon-style… · home topics machine-learning article
[ARTICLE · art-105456] src=machinebrief.com ↗ pub= topic=machine-learning verified=true sentiment=· neutral

Truncate Bad, Upweight Good: BoN-Style Distillation via Rank-Based Classification

Researchers propose TUP (Truncate-bad, Upweight-good Policy), a distillation method that removes low-ranked completions from the support and reweights only the retained upper tail with tunable sharpness, achieving competitive performance with strong offline alignment baselines. The method, described in arXiv:2608.19748v1, uses shifted-truncated win-rates as soft labels and distilled-to-reference log-likelihood ratios as logits, and is trained fully offline via binary cross-entropy.

read1 min views1 publishedAug 21, 2026

arXiv:2608.19748v1 Announce Type: new Abstract: Inference-time selection methods, such as Best-of-N, improve generation by sampling a pool of candidates and selecting the top-ranked completion according to a reward model. Distillation seeks to amortize this procedure into a single policy by replacing raw rewards with in-pool ranks and learning a policy that upweights higher-ranked completions. However, existing rank-based policies typically use smooth full-support reweighting, so low-ranked completions receive less mass but remain in the target support. Although a sharper reweighting reduces lower-tail mass, it also increases reliance on brittle ranking at the top made by a single reward model. We propose TUP: a Truncate-bad, Upweight-good Policy that removes low-ranked completions from the support and reweights only the retained upper tail with a tunable sharpness. TUP admits a closed-form, prompt-independent normalization and can be trained fully offline via binary cross-entropy, using shifted-truncated win-rates as soft labels and distilled-to-reference log-likelihood ratios as logits. Theoretically, under certain assumptions, we show that for any unknown oracle reward, the best monotone rank-reweighting can be matched by a lower-tail truncation rule, providing formal support for removing the lower tail rather than merely downweighting it. Empirically, we show that TUP is competitive with strong offline alignment baselines.

── more in #machine-learning 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/truncate-bad-upweigh…] indexed:0 read:1min 2026-08-21 ·