{"slug": "what-stops-a-small-language-model-from-driving-a-database-agent", "title": "What stops a small language model from driving a database agent", "summary": "A study of 39 open-weight language models driving the agent mode of an open-source SQL client over eleven days found that 1,590 of 2,100 model-attributed agent-mode losses, or 75.7%, came from runs that had invoked at least one tool, contradicting the assumption that small open-weight models fail at agentic database work for lack of reasoning capacity. The authors report that transport failures, runs that used tools but never produced a deliverable, were the largest class at 36.2%, while capability failures, runs that invoked no tool at all, were the smallest at 17.3%, and that five server changes touching no model, prompt or sampling setting moved six models by 6 to 21 cells out of 30. The study also flags a confound for published local-model benchmarks: with no context cap, one 7.1 GB model was admitted at its full 262,144-token window and held 51 GB on a 64 GB machine, producing runs indistinguishable in ordinary logs from a model timing out.", "body_md": "# Computer Science > Software Engineering\n\n  [Submitted on 18 Sep 2026]\n\n# Title:What Stops a Small Language Model From Driving a Database Agent\n\n[View PDF](https://arxiv.org/pdf/2609.21341)\n\n[HTML (experimental)](https://arxiv.org/html/2609.21341v1)\n\nAbstract:Small open-weight language models are assumed to fail at agentic database work because they lack the reasoning capacity for it. We test that against a production system. Over eleven days we drove the agent mode of an open-source SQL client with 39 open-weight models served locally and one hosted control, across six task surfaces: 8,199 runs, 110,711 ledger events, 14,008 refused tool calls. Of the 2,100 model-attributed agent-mode losses, 1,590, or 75.7%, came from runs that had invoked at least one tool. That majority is what survives resampling models rather than runs: it holds in 99.7% of clustered resamples and in 15 of the 22 models with at least twenty losses. Within it, transport, a run that used the tools and never got a deliverable through, is the largest class at 36.2% and capability, a run that invoked no tool at all, the smallest at 17.3%; we report that ordering as a property of this corpus rather than a general finding, since clustered by model it holds in only 74.5% of resamples. Transport failures decompose into a few mechanical argument shapes. Production ledgers record refusal codes and never the model's arguments, so these were invisible for ten days; capturing them exposed five server defects, one of which demanded a field on one tool, forbade it on the sibling that composed it, then failed the run for its absence. Five server changes, touching no model, prompt or sampling setting, moved six models by 6 to 21 cells out of 30. We also report a confound we believe affects published local-model benchmarks, ours included: with no context cap, one 7.1 GB model was admitted at its full 262,144-token window and held 51 GB on a 64 GB machine, producing runs indistinguishable in any ordinary log from a model timing out. The corpus, the scorer and a verifier that regenerates every figure are released.\n    \n\n# Bibliographic and Citation Tools\n\nBibliographic Explorer \n\n*(*[What is the Explorer?](https://info.arxiv.org/labs/showcase.html#arxiv-bibliographic-explorer))\nConnected Papers \n\n*(*[What is Connected Papers?](https://www.connectedpapers.com/about))\nLitmaps \n\n*(*[What is Litmaps?](https://www.litmaps.co/))\nscite Smart Citations \n\n*(*[What are Smart Citations?](https://www.scite.ai/))\n# Code, Data and Media Associated with this Article\n\nalphaXiv \n\n*(*[What is alphaXiv?](https://alphaxiv.org/))\nCatalyzeX Code Finder for Papers \n\n*(*[What is CatalyzeX?](https://www.catalyzex.com))\nDagsHub \n\n*(*[What is DagsHub?](https://dagshub.com/))\nGotit.pub \n\n*(*[What is GotitPub?](http://gotit.pub/faq))\nHugging Face \n\n*(*[What is Huggingface?](https://huggingface.co/huggingface))\nScienceCast \n\n*(*[What is ScienceCast?](https://sciencecast.org/welcome))\n# Demos\n\n# Recommenders and Search Tools\n\nInfluence Flower \n\n*(*[What are Influence Flowers?](https://influencemap.cmlab.dev/))\nCORE Recommender \n\n*(*[What is CORE?](https://core.ac.uk/services/recommender))\n# arXivLabs: experimental projects with community collaborators\n\narXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.\n\nBoth individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.\n\nHave an idea for a project that will add value for arXiv's community? [**Learn more about arXivLabs**](https://info.arxiv.org/labs/index.html).", "url": "https://wpnews.pro/news/what-stops-a-small-language-model-from-driving-a-database-agent", "canonical_source": "https://arxiv.org/abs/2609.21341", "published_at": "2026-10-06 22:10:38+00:00", "updated_at": "2026-10-06 22:19:40.943533+00:00", "lang": "en", "topics": ["large-language-models", "ai-agents", "ai-research", "mlops", "ai-tools"], "entities": ["arXiv", "SQL"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/what-stops-a-small-language-model-from-driving-a-database-agent", "markdown": "https://wpnews.pro/news/what-stops-a-small-language-model-from-driving-a-database-agent.md", "text": "https://wpnews.pro/news/what-stops-a-small-language-model-from-driving-a-database-agent.txt", "jsonld": "https://wpnews.pro/news/what-stops-a-small-language-model-from-driving-a-database-agent.jsonld"}}