cd /news/artificial-intelligence/efficient-best-of-n-policy-evaluatio… · home › topics › artificial-intelligence › article
[ARTICLE · art-147378] src=machinebrief.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Efficient Best-of-N policy evaluation for inference-time alignment

A new arXiv paper (2610.09250v1) proposes a sample-only framework for evaluating and selecting Best-of-N (BoN) inference-time alignment policies without access to response likelihoods, which standard off-policy estimators require. The authors introduce BoN-DR, a doubly robust estimator of BoN policy value that reuses a shared auxiliary sample pool across candidate budgets, and prove its efficiency and valid asymptotic inference under reward estimator misspecification. Two selection rules are derived — maximizing estimated policy value and maximizing a lower confidence bound on improvement over the reference policy — and the framework accurately estimates BoN policy values and selects effective sampling budgets across synthetic experiments and GSM8K with multiple reference and reward models.

by read1 min views1 publishedOct 8, 2026

arXiv:2610.09250v1 Announce Type: new Abstract: Best-of-N (BoN) is a common inference-time alignment method that selects the highest-scoring response among N samples from a reference model. Evaluating BoN policies from logged data is challenging under sample-only access because standard off-policy estimators require density ratios that depend on unavailable response likelihoods. In this paper, we propose a sample-only framework for evaluating and selecting BoN policies without access to these likelihoods. We show that the order-statistic structure of BoN allows the required density ratios to be expressed through score-rank probabilities that are estimable from samples alone. We then develop a doubly robust estimator of the BoN policy value (BoN-DR) that efficiently reuses a shared auxiliary sample pool across candidate budgets. We establish valid asymptotic inference even under reward estimator misspecification and prove the efficiency of our BoN-DR estimator. Since larger budgets can amplify errors in the score function and lead to reward overoptimization, we derive two selection rules: (i) maximizing the estimated policy value and (ii) maximizing a lower confidence bound on the improvement over the reference policy, which accounts for estimation uncertainty and provides a no-harm guarantee. Across synthetic experiments and GSM8K with multiple reference and reward models, our framework accurately estimates BoN policy values and selects effective sampling budgets.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/efficient-best-of-n-…] indexed:0 read:1min 2026-10-08 · —