cd /news/artificial-intelligence/urbanagent-a-tool-augmented-agent-fo… · home topics artificial-intelligence article
[ARTICLE · art-87146] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=↑ positive

UrbanAgent: A Tool-Augmented Agent for Cross-System Urban Tasks

Researchers propose Urban-Agent, a tool-augmented agent framework that combines a large language model with code execution, API calls, and Model Context Protocol to handle cross-system urban tasks. In evaluations using the new Urban-Eval benchmark, Urban-Agent achieved a 71% task success rate, 10 points above the strongest baseline, across GPT-5-mini, Gemini-2.5-flash, DeepSeek-V4-flash, and Qwen3-235B-A22B.

read1 min views1 publishedAug 5, 2026

arXiv:2608.03018v1 Announce Type: new Abstract: Modern cities rely on an increasing number of digital services to operate, but residents' daily needs are still difficult to meet. Services are fragmented and have little interoperability, placing a heavy operational burden on users. Existing digital platforms, urban foundation models, and intelligent assistants each address only isolated aspects of an urban task. But they struggle to reliably convert complex natural-language requests into executable cross-system workflows. We propose Urban-Agent, a tool-augmented agent framework for cross-system urban tasks. It couples the cognitive and reasoning capabilities of a large language model with a tool-set supporting code execution, API calls, and Model Context Protocol. Through one adaptive closed loop, it clarifies missing information before acting, grounds tool use in live observations, and aligns the final response with observed evidence and task constraints. To address the evaluation gap, we introduce Urban-Eval, a benchmark specifically designed for cross-system urban request. Unlike prior benchmarks that assess either general tool use or urban knowledge and reasoning, Urban-Eval evaluates both task results and execution quality, including required tool coverage, dependency validity, and evidence traceability. Experimental results indicate that Urban-Agent reaches a 71% task success rate, 10 points above the strongest baseline. This lead holds across GPT-5-mini, Gemini-2.5-flash, DeepSeek-V4-flash, and Qwen3-235B-A22B.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @urban-agent 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/urbanagent-a-tool-au…] indexed:0 read:1min 2026-08-05 ·