{"slug": "efficient-best-of-n-policy-evaluation-for-inference-time-alignment", "title": "Efficient Best-of-N policy evaluation for inference-time alignment", "summary": "A new arXiv paper (2610.09250v1) proposes a sample-only framework for evaluating and selecting Best-of-N (BoN) inference-time alignment policies without access to response likelihoods, which standard off-policy estimators require. The authors introduce BoN-DR, a doubly robust estimator of BoN policy value that reuses a shared auxiliary sample pool across candidate budgets, and prove its efficiency and valid asymptotic inference under reward estimator misspecification. Two selection rules are derived — maximizing estimated policy value and maximizing a lower confidence bound on improvement over the reference policy — and the framework accurately estimates BoN policy values and selects effective sampling budgets across synthetic experiments and GSM8K with multiple reference and reward models.", "body_md": "arXiv:2610.09250v1 Announce Type: new \nAbstract: Best-of-N (BoN) is a common inference-time alignment method that selects the highest-scoring response among N samples from a reference model. Evaluating BoN policies from logged data is challenging under sample-only access because standard off-policy estimators require density ratios that depend on unavailable response likelihoods. In this paper, we propose a sample-only framework for evaluating and selecting BoN policies without access to these likelihoods. We show that the order-statistic structure of BoN allows the required density ratios to be expressed through score-rank probabilities that are estimable from samples alone. We then develop a doubly robust estimator of the BoN policy value (BoN-DR) that efficiently reuses a shared auxiliary sample pool across candidate budgets. We establish valid asymptotic inference even under reward estimator misspecification and prove the efficiency of our BoN-DR estimator. Since larger budgets can amplify errors in the score function and lead to reward overoptimization, we derive two selection rules: (i) maximizing the estimated policy value and (ii) maximizing a lower confidence bound on the improvement over the reference policy, which accounts for estimation uncertainty and provides a no-harm guarantee. Across synthetic experiments and GSM8K with multiple reference and reward models, our framework accurately estimates BoN policy values and selects effective sampling budgets.", "url": "https://wpnews.pro/news/efficient-best-of-n-policy-evaluation-for-inference-time-alignment", "canonical_source": "https://www.machinebrief.com/news/efficient-best-of-n-policy-evaluation-for-inference-time-ali-dybb", "published_at": "2026-10-08 04:00:00+00:00", "updated_at": "2026-10-08 06:17:07.758204+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "large-language-models", "ai-research", "ai-safety"], "entities": ["arXiv", "Best-of-N", "BoN-DR", "GSM8K"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/efficient-best-of-n-policy-evaluation-for-inference-time-alignment", "markdown": "https://wpnews.pro/news/efficient-best-of-n-policy-evaluation-for-inference-time-alignment.md", "text": "https://wpnews.pro/news/efficient-best-of-n-policy-evaluation-for-inference-time-alignment.txt", "jsonld": "https://wpnews.pro/news/efficient-best-of-n-policy-evaluation-for-inference-time-alignment.jsonld"}}