{"slug": "multi-step-tool-calling-over-korean-open-public-apis-a-benchmark-and-a-data", "title": "Multi-Step Tool-Calling over Korean Open Public APIs: A Benchmark and a Data-Synthesis Recipe", "summary": "Researchers introduced the Korean Open Public API Benchmark (KOPA-Bench) with 145 real-world tasks to measure multi-step tool-calling performance of open-source LLM agents on live Korean government APIs, and presented EDGE, an Execution-grounded Dynamic Graph method that synthesizes executable multi-step trajectories by verifying tool-output links against live APIs. Fine-tuning a 9B model via GRPO on EDGE-generated data nearly matched the untuned 27B model from the same family, with substantial improvements on KOPA-Bench and the BFCL benchmark.", "body_md": "# Computer Science > Artificial Intelligence\n\n  [Submitted on 4 Sep 2026]\n\n# Title:Multi-Step Tool-Calling over Korean Open Public APIs: A Benchmark and a Data-Synthesis Recipe\n\n[View PDF](/pdf/2609.05395v1)\n\n[HTML (experimental)](https://arxiv.org/html/2609.05395v1)\n\nAbstract:Data-sovereignty regulations increasingly require public institutions to deploy open-source, on-premise LLM agents that chain multiple tool-calls across live government APIs. However, open-source models consistently underperform in this multi-step setting, and no existing benchmark measures the gap. We introduce the Korean Open Public API Benchmark (KOPA-Bench), comprising 145 real-world tasks. To close this gap, we present EDGE, an Execution-grounded Dynamic Graph for tool-calling data synthEsis driven by live execution. EDGE builds a graph of how each tool's output can feed another's input, keeps only the links that succeed when actually called against the live APIs, and traverses these verified links to synthesize executable multi-step trajectories. Fine-tuned via GRPO on the resulting dataset, our 9B model nearly matches the untuned 27B model from the same family, improving substantially not only on KOPA-Bench but also on the BFCL benchmark.\n    \n\n### References & Citations\n\nLoading...\n\n# Bibliographic and Citation Tools\n\nBibliographic Explorer \n\n*(*[What is the Explorer?](https://info.arxiv.org/labs/showcase.html#arxiv-bibliographic-explorer))\nConnected Papers \n\n*(*[What is Connected Papers?](https://www.connectedpapers.com/about))\nLitmaps \n\n*(*[What is Litmaps?](https://www.litmaps.co/))\nscite Smart Citations \n\n*(*[What are Smart Citations?](https://www.scite.ai/))\n# Code, Data and Media Associated with this Article\n\nalphaXiv \n\n*(*[What is alphaXiv?](https://alphaxiv.org/))\nCatalyzeX Code Finder for Papers \n\n*(*[What is CatalyzeX?](https://www.catalyzex.com))\nDagsHub \n\n*(*[What is DagsHub?](https://dagshub.com/))\nGotit.pub \n\n*(*[What is GotitPub?](http://gotit.pub/faq))\nHugging Face \n\n*(*[What is Huggingface?](https://huggingface.co/huggingface))\nScienceCast \n\n*(*[What is ScienceCast?](https://sciencecast.org/welcome))\n# Demos\n\n# Recommenders and Search Tools\n\nInfluence Flower \n\n*(*[What are Influence Flowers?](https://influencemap.cmlab.dev/))\nCORE Recommender \n\n*(*[What is CORE?](https://core.ac.uk/services/recommender))\n# arXivLabs: experimental projects with community collaborators\n\narXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.\n\nBoth individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.\n\nHave an idea for a project that will add value for arXiv's community? [**Learn more about arXivLabs**](https://info.arxiv.org/labs/index.html).", "url": "https://wpnews.pro/news/multi-step-tool-calling-over-korean-open-public-apis-a-benchmark-and-a-data", "canonical_source": "http://arxiv.org/abs/2609.05395v1", "published_at": "2026-09-08 14:41:23+00:00", "updated_at": "2026-09-08 14:58:11.108201+00:00", "lang": "en", "topics": ["artificial-intelligence", "ai-research", "ai-agents"], "entities": ["KOPA-Bench", "EDGE", "GRPO", "BFCL"], "alternates": {"html": "https://wpnews.pro/news/multi-step-tool-calling-over-korean-open-public-apis-a-benchmark-and-a-data", "markdown": "https://wpnews.pro/news/multi-step-tool-calling-over-korean-open-public-apis-a-benchmark-and-a-data.md", "text": "https://wpnews.pro/news/multi-step-tool-calling-over-korean-open-public-apis-a-benchmark-and-a-data.txt", "jsonld": "https://wpnews.pro/news/multi-step-tool-calling-over-korean-open-public-apis-a-benchmark-and-a-data.jsonld"}}