arXiv:2502.04492v3 Announce Type: replace Abstract: The advancement of LLMs and their accessibility have triggered renewed interest in multi-agent reinforcement learning as robust and adaptive frameworks for dynamically changing environments. This paper introduces \texttt{RL-Focal}, a two-stage RL agent framework that routes and ensembles LLMs. \textit{First}, we develop the Decider RL-agent, which learns to dynamically select an ensemble of small size ($m_i$) among $N$ LLMs ($m_i \ll N$) for incoming queries from a user-defined downstream task $i$, by maximizing both error-diversity and reasoning-performance of the selected ensemble through iterative updates of task-adaptive rewards and policy. \textit{Second}, to enable effective fusion of dynamically selected LLMs, we develop the stage-2 Fusion RL-agent, which learns to resolve reasoning conflicts from different LLMs and dynamically adapt to different ensemble teams composed by the Decider Agent for different downstream tasks. {\em Third}, we introduce the focal diversity metric to better model the error correlations among multiple LLMs further improving the generalization performance of the Decider Agent, which actively prunes the ensemble combinations. By focal diversity, we enhance performance across tasks by effectively promoting reward-aware and policy-adaptive ensemble selection and inference fusion. Extensive evaluations on five benchmarks show that RL-Focal achieves the performance improvement of 8.48% with an ensemble of small size compared to the best individual LLM in a pool and offers stronger robustness. Code is available \href{https://github.com/git-disl/RL-Focal}{here}.
Dynamic Optimizations of LLM Ensembles with Two-Stage Reinforcement Learning Agents
Researchers introduced RL-Focal, a two-stage reinforcement learning agent framework that routes and ensembles large language models, achieving an 8.48% performance improvement over the best individual LLM in a pool across five benchmarks, according to the arXiv paper 2502.04492v3. The framework's first-stage Decider RL-agent selects a small ensemble of size m_i from N LLMs by maximizing error-diversity and reasoning performance, while the second-stage Fusion RL-agent resolves reasoning conflicts and adapts to different ensemble teams. The authors also introduce a focal diversity metric to model error correlations among LLMs and prune ensemble combinations, with code available on GitHub.
Run your AI side-project on zahid.host
EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.