{"slug": "toolsearcher-optimizing-tool-selection-at-scale-via-reinforcement-learning", "title": "ToolSearcher: Optimizing Tool Selection at Scale via Reinforcement Learning", "summary": "A new arXiv paper (2609.30906v1) proposes ToolSearcher, a reinforcement learning framework for large-scale tool selection by LLM agents, addressing settings where real-world tool repositories are too large to fit within context-length constraints. ToolSearcher introduces category-constrained tool discrimination, event-level search modeling, and trajectory-aligned credit allocation, and the authors report it consistently outperforms strong baselines on large-scale tool selection benchmarks involving iterative search and complex tool composition.", "body_md": "arXiv:2609.30906v1 Announce Type: new \nAbstract: Large language models (LLMs) excel at natural language processing but struggle to interact with external environments. Tool learning provides a promising way to extend LLMs into actionable agents, where tool selection is a critical prerequisite for successful tool use. Existing work often assumes a small or predefined set of tools, leaving large-scale tool selection underexplored. Real-world repositories contain a vast and diverse array of tools, making it difficult for LLMs to effectively search, distinguish, and compose tools under context-length constraints. We identify large-scale tool selection as a new challenge for agentic reinforcement learning, highlighting that existing RL methods for knowledge-based question answering are inadequate for selecting tools while considering compatibility. To address this challenge, we propose ToolSearcher, a novel RL framework for effective multi-turn search and fine-grained optimization in large-scale tool selection. Specifically, we introduce category-constrained tool discrimination to improve the model's ability to distinguish functionally similar tools, event-level search modeling to explicitly optimize the discovery of target tools during multi-turn search, and trajectory-aligned credit allocation to provide fine-grained reward signals for different stages of the search-selection process. Extensive experiments on large-scale tool selection benchmarks demonstrate that ToolSearcher consistently outperforms a set of strong baselines in challenging settings involving iterative search and complex tool composition.", "url": "https://wpnews.pro/news/toolsearcher-optimizing-tool-selection-at-scale-via-reinforcement-learning", "canonical_source": "https://www.machinebrief.com/news/toolsearcher-optimizing-tool-selection-at-scale-via-reinforc-mi7o", "published_at": "2026-09-28 04:00:00+00:00", "updated_at": "2026-09-28 04:48:19.240933+00:00", "lang": "en", "topics": ["ai-agents", "artificial-intelligence", "large-language-models", "ai-research", "machine-learning"], "entities": ["ToolSearcher", "arXiv"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/toolsearcher-optimizing-tool-selection-at-scale-via-reinforcement-learning", "markdown": "https://wpnews.pro/news/toolsearcher-optimizing-tool-selection-at-scale-via-reinforcement-learning.md", "text": "https://wpnews.pro/news/toolsearcher-optimizing-tool-selection-at-scale-via-reinforcement-learning.txt", "jsonld": "https://wpnews.pro/news/toolsearcher-optimizing-tool-selection-at-scale-via-reinforcement-learning.jsonld"}}