{"slug": "from-llm-inference-to-agentic-workloads-characterization-and-implications", "title": "From LLM Inference to Agentic Workloads: Characterization and Implications", "summary": "A new benchmark suite, AgentSysBench, reveals that agentic AI workloads differ fundamentally from conventional LLM serving, with non-LLM components dominating latency in 5 of 10 applications and sandbox memory peaking at 28 GB per session. The study, submitted to arXiv on 15 Aug 2026, identifies six distinguishing properties and shows that task-aware serving reduces latency by 29–40%, communication-aware placement by up to 4.5x, state offloading cuts memory by 4.6x, and tool-result caching removes 35.2% of redundant search calls, saving 19.3% of aggregate search latency.", "body_md": "# Computer Science > Operating Systems\n\n[Submitted on 15 Aug 2026]\n\n# Title:From LLM Inference to Agentic Workloads: Characterization and Implications for Serving Systems\n\n[View PDF](/pdf/2608.15127)\n\n[HTML (experimental)](https://arxiv.org/html/2608.15127v1)\n\nAbstract:Agentic applications are shifting AI serving from isolated model inference to long-running workloads in which LLMs coordinate tools, environments, and persistent state. However, the system behavior of these workloads---where latency, cost, and bottlenecks arise---remains poorly characterized, leaving serving systems to rely on assumptions built for conventional inference. We present AgentSysBench, a benchmark suite and measurement toolkit with ten representative agentic applications and unified systems-level instrumentation. Across controlled deployments and production traces, we identify six properties that distinguish agentic workloads from conventional LLM serving: (1) execution is heavyweight and stateful, with non-LLM components dominating latency in 5 of 10 applications and sandbox working-set memory peaking at 28 GB per session; (2) applications compose components with heterogeneous resource affinity---GPU-bound inference, memory-bound retrieval, CPU-bound sandboxes---whose task latencies diverge by up to 32x; (3) bottlenecks shift across requests, models, and deployments; (4) production sessions hold state idle for minutes to hours between active steps; (5) a control-plane tax---auxiliary LLM calls and context overhead from tool schemas and observations---crowds out productive compute and context; and (6) production traces from three applications reveal heavy cross-request redundancy in search queries and web fetches, exposing a large caching opportunity. Four design explorations demonstrate that these findings are actionable: task-aware serving reduces latency by 29--40%, communication-aware placement by up to 4.5x, state offloading reduces memory usage by 4.6x, and tool-result caching removes 35.2% of redundant search calls and saves 19.3% of aggregate search latency.\n\n### Current browse context:\n\ncs.OS\n\n### References & Citations\n\nLoading...\n\n# Bibliographic and Citation Tools\n\nBibliographic Explorer\n\n*(*[What is the Explorer?](https://info.arxiv.org/labs/showcase.html#arxiv-bibliographic-explorer))\nConnected Papers\n\n*(*[What is Connected Papers?](https://www.connectedpapers.com/about))\nLitmaps\n\n*(*[What is Litmaps?](https://www.litmaps.co/))\nscite Smart Citations\n\n*(*[What are Smart Citations?](https://www.scite.ai/))# Code, Data and Media Associated with this Article\n\nalphaXiv\n\n*(*[What is alphaXiv?](https://alphaxiv.org/))\nCatalyzeX Code Finder for Papers\n\n*(*[What is CatalyzeX?](https://www.catalyzex.com))\nDagsHub\n\n*(*[What is DagsHub?](https://dagshub.com/))\nGotit.pub\n\n*(*[What is GotitPub?](http://gotit.pub/faq))\nHugging Face\n\n*(*[What is Huggingface?](https://huggingface.co/huggingface))\nScienceCast\n\n*(*[What is ScienceCast?](https://sciencecast.org/welcome))# Demos\n\n# Recommenders and Search Tools\n\nInfluence Flower\n\n*(*[What are Influence Flowers?](https://influencemap.cmlab.dev/))\nCORE Recommender\n\n*(*[What is CORE?](https://core.ac.uk/services/recommender))# arXivLabs: experimental projects with community collaborators\n\narXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.\n\nBoth individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.\n\nHave an idea for a project that will add value for arXiv's community? [ Learn more about arXivLabs](https://info.arxiv.org/labs/index.html).", "url": "https://wpnews.pro/news/from-llm-inference-to-agentic-workloads-characterization-and-implications", "canonical_source": "https://arxiv.org/abs/2608.15127", "published_at": "2026-08-26 17:01:07+00:00", "updated_at": "2026-08-26 17:17:17.616449+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-infrastructure", "ai-research", "ai-agents"], "entities": ["AgentSysBench", "arXiv"], "alternates": {"html": "https://wpnews.pro/news/from-llm-inference-to-agentic-workloads-characterization-and-implications", "markdown": "https://wpnews.pro/news/from-llm-inference-to-agentic-workloads-characterization-and-implications.md", "text": "https://wpnews.pro/news/from-llm-inference-to-agentic-workloads-characterization-and-implications.txt", "jsonld": "https://wpnews.pro/news/from-llm-inference-to-agentic-workloads-characterization-and-implications.jsonld"}}