{"slug": "selection-is-retrieval-abstention-is-not-on-device-tool-routing-over-70-korean", "title": "Selection Is Retrieval, Abstention Is Not: On-Device Tool Routing over 70 Korean-English Actions", "summary": "A study of on-device tool routing over 600 Korean and English requests and a catalog of 70 local actions found that abstention, not tool selection, is where a neural component is required, according to the arXiv paper 2609.18672v1. Character 3-gram BM25 selected 162 of 164 lexically matched requests but only 85 of 166 paraphrases, and restricting candidates to seven raised the paraphrase figure to a mean of 0.825 over five trials. No classifier over BM25 score features separated in-catalog from out-of-catalog requests above 0.697 area under the curve, while the frozen encoder multilingual-e5-base reached 0.806 and, used for abstention alone, kept 376 requests local and misrouted 9 of the 150 needing delegation.", "body_md": "arXiv:2609.18672v1 Announce Type: new \nAbstract: An AI assistant that calls tools makes two decisions on every request: which tool to invoke, and whether any available tool applies. In the usual design a single language model makes both, by emitting a call or by declining to emit one. On a device that has to answer without a server, the language model is what makes that design expensive, dominating both the latency and the memory of the router. The common alternative is to remove the model completely and rank the catalog of local actions with a retriever instead. That substitution is not symmetric across the two decisions. A retriever returns its highest-scoring candidate for every input and cannot signal that the catalog holds no valid action. Our earlier study found that constraining a decoder to a tool grammar repairs malformed output without improving the choice. What the substitution costs in each decision has not been measured. We evaluate the two decisions separately over 600 Korean and English requests and a catalog of 70 local actions. The router may also ask for a missing slot, reply, or delegate. Half the in-catalog requests reuse catalog vocabulary and half paraphrase it, separating lexical overlap from the action requested. Character 3-gram BM25 selects 162 of 164 lexically matched requests and 85 of 166 paraphrases. Restricting the candidate set to seven raises the paraphrase figure to a mean of 0.825 over five trials. No classifier over its score features separates in-catalog from out-of-catalog above 0.697 area under the curve, where the frozen encoder multilingual-e5-base reaches 0.806. Using that encoder for abstention alone keeps 376 of the requests local and misroutes 9 of the 150 needing delegation. Abstention, not selection, is where a neural component is required. A neural ranker improves every quality metric and is rejected on latency and memory rather than accuracy.", "url": "https://wpnews.pro/news/selection-is-retrieval-abstention-is-not-on-device-tool-routing-over-70-korean", "canonical_source": "https://www.machinebrief.com/news/selection-is-retrieval-abstention-is-not-on-device-tool-rout-4pin", "published_at": "2026-09-17 04:00:00+00:00", "updated_at": "2026-09-17 06:55:04.352729+00:00", "lang": "en", "topics": ["ai-agents", "natural-language-processing", "ai-research", "ai-tools"], "entities": ["arXiv", "BM25", "multilingual-e5-base"], "alternates": {"html": "https://wpnews.pro/news/selection-is-retrieval-abstention-is-not-on-device-tool-routing-over-70-korean", "markdown": "https://wpnews.pro/news/selection-is-retrieval-abstention-is-not-on-device-tool-routing-over-70-korean.md", "text": "https://wpnews.pro/news/selection-is-retrieval-abstention-is-not-on-device-tool-routing-over-70-korean.txt", "jsonld": "https://wpnews.pro/news/selection-is-retrieval-abstention-is-not-on-device-tool-routing-over-70-korean.jsonld"}}