Cost-Aware Best-LLM Identification using Dueling Feedback A new arXiv paper (2609.30360v1) formulates a variant of the multi-armed bandit problem for identifying the best large language model from a collection with heterogeneous querying costs, using dueling pairwise-comparison feedback. The authors propose a Track-and-Stop style algorithm for best-arm identification with prescribed confidence and prove it almost surely achieves the asymptotically optimal cost as the error tends to zero, assuming a Condorcet winner that they say they empirically validated across multiple real-world datasets. Evaluations on synthetic and real-world instances showed consistent improvements over classical cost-unaware algorithms and their cost-aware extensions. arXiv:2609.30360v1 Announce Type: new Abstract: Inspired by the problem of identifying the best model from a collection of large language models LLMs with heterogeneous querying costs, we formulate and analyse a variant of the multi-armed bandit MAB with i dueling feedback, where pairwise comparisons between model responses provide robust preference signals, and ii heterogeneous sampling costs, reflecting the differing costs of querying different LLMs. Assuming the existence of a Condorcet winner, a condition we empirically validate across multiple real-world datasets, we propose a Track-and-Stop style algorithm for best-arm identification with prescribed confidence. We prove that the algorithm almost surely achieves the asymptotically optimal cost as the error tends to zero. Finally, we extensively evaluate our approach on both synthetic and real-world instances, demonstrating consistent improvements over classical cost-unaware algorithms and their cost-aware extensions.