{"slug": "toollery-scaling-llm-agents-to-thousands-of-skills-and-tools", "title": "Toollery: Scaling LLM Agents to Thousands of Skills and Tools", "summary": "A training-free candidate-compression framework called Toollery improves LLM agent skill and tool selection across libraries scaling to roughly 79,000 capabilities, according to a new arXiv paper (2609.22218v1). Toollery generates user-intent queries from each skill or tool specification and builds a retrieval index that narrows real user requests to a compact top-k candidate set before the LLM makes its final selection. Evaluated on the roughly 79K-capability SkillRouter benchmark, BFCL-V4 with over 440 atomic tools, and 3,396 proprietary smart-cockpit requests over 220 tools, Toollery improved end-to-end selection on the cockpit dataset at a fixed top-10 budget and maintained comparable AST Accuracy on BFCL-V4, though the authors note quality and cost gains depend on workload coverage and provider caching.", "body_md": "arXiv:2609.22218v1 Announce Type: new \nAbstract: As LLM agents are exposed to hundreds to tens of thousands of skills, tools, and API functions, full-library prompting becomes costly, slow, and less reliable: each added candidate increases prompt tokens and latency, while longer candidate lists introduce more distractors for LLM selection. We present \\textbf{Toollery}, a training-free candidate-compression framework for scalable LLM skill/tool selection. Following established document-side query expansion, Toollery generates user-intent queries from each skill/tool specification and builds a retrieval index that maps real user requests to compact candidate sets before final LLM decision-making. By treating high-level skills and atomic tools as selectable capabilities, Toollery can be applied to both skill libraries and tool registries. We evaluate Toollery on the roughly 79K-capability SkillRouter benchmark, BFCL-V4 with over 440 atomic tools, and 3,396 proprietary smart-cockpit requests over 220 tools. Across these settings, Toollery keeps online selection bounded to a compact top-$k$ candidate set and improves recall over ordinary specification retrieval. At a fixed top-10 budget, Toollery improves end-to-end selection on the cockpit dataset, and maintains comparable AST Accuracy on BFCL-V4. These results support Toollery as a practical candidate-compression framework for large and evolving agent capability libraries, while showing that quality and cost gains depend on workload coverage and provider caching.", "url": "https://wpnews.pro/news/toollery-scaling-llm-agents-to-thousands-of-skills-and-tools", "canonical_source": "https://www.machinebrief.com/news/toollery-scaling-llm-agents-to-thousands-of-skills-and-tools-gu8t", "published_at": "2026-09-22 04:00:00+00:00", "updated_at": "2026-09-22 05:24:50.659298+00:00", "lang": "en", "topics": ["ai-agents", "ai-tools", "large-language-models", "ai-research", "natural-language-processing"], "entities": ["Toollery", "SkillRouter", "BFCL-V4", "arXiv"], "alternates": {"html": "https://wpnews.pro/news/toollery-scaling-llm-agents-to-thousands-of-skills-and-tools", "markdown": "https://wpnews.pro/news/toollery-scaling-llm-agents-to-thousands-of-skills-and-tools.md", "text": "https://wpnews.pro/news/toollery-scaling-llm-agents-to-thousands-of-skills-and-tools.txt", "jsonld": "https://wpnews.pro/news/toollery-scaling-llm-agents-to-thousands-of-skills-and-tools.jsonld"}}