cd /news/artificial-intelligence/the-ai-invoice-nobody-planned-for-an… · home topics artificial-intelligence article
[ARTICLE · art-86138] src=blogs.cisco.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

The AI invoice nobody planned for — and how Cisco is solving it from the inside

Cisco Systems Inc. is applying its own AI tokenomics discipline, using Splunk Agent Observability and its Galileo acquisition to track and optimize token consumption, aiming to balance AI cost and value. The company reports that enterprises face unpredictable AI costs due to disconnected vendor dashboards and manual reports, and Cisco's integrated infrastructure, security, and observability stack provides end-to-end visibility to control spend. Cisco's goal is to achieve equilibrium where token cost and business value align, rather than simply reducing AI spending.

read5 min views1 publishedAug 4, 2026
The AI invoice nobody planned for — and how Cisco is solving it from the inside
Image: Blogs (auto-discovered)

*Finance is asking IT to justify AI spend. Most IT leaders don’t have a good answer. Hear from the Cisco leaders driving our AI adoption and infrastructure about how we are building our own — from the inside. *

The economics of enterprise AI at scale

Token consumption has quickly become a core currency of enterprise competitiveness. Your ability to deploy AI economically, efficiently, and at scale will determine not just your AI ROI, but your organization’s success in this operating climate.

Yet as enterprises scale, cost and a lack of visibility and trust become the deciding factors in sustainable AI, not the technology itself. When every autonomous action consumes tokens, costs can become difficult to forecast and even harder to control without the right visibility.

IT leaders and engineering teams are contending with disconnected vendor dashboards and manual reports, lacking a reliable view into which teams are spending what, and why. Finance teams are asking for justification on AI budget spikes, but without unified data, IT leaders are left reacting to costs they cannot explain, attribute, or forecast.

Early on, we saw the same pattern many enterprises face: reacting to token spikes and shifting budget from other engineering priorities just to keep agents running. Rather than manage around it, we set out to solve it, building the visibility needed to better understand and control our AI spend.

This is not a problem unique to Cisco, but it is one we are uniquely positioned to solve. To move from reactive budget-cutting to strategic investment, we needed to make tokenomics a core discipline.

Tokenomics: the defining success factor

In the context of enterprise AI, tokenomics is the measure of efficiency for your AI operations. It is about maximizing “token yield,” the ROI per token. We measure success by the correlation between token consumption and the quality of output, task completion, and business value. Our goal isn’t necessarily to spend less on AI, but to reach an equilibrium where cost and value are in balance.

That equilibrium is impossible to achieve without the right integration. We’ve found that infrastructure, security, and observability all impact tokenomics, and none of them can work in isolation:

Without optimized infrastructure, GPUs can sit idle while you continue to pay for the compute, inflating the infrastructure cost relative to token spend. Without AI governance and security, agents can consume tokens on unauthorized tasks or inefficient processes, leading to budget overruns that are impossible to claw back. And without end-to-end observability, you can’t connect AI activity to financial outcomes, which leads to measuring inefficiency, not eliminating it.

This is where Cisco’s value is different: no other vendor covers every layer of this chain. Because our capabilities span from full stack AI infrastructure to models, security, and agent observability, we have the unique, end-to-end visibility required for true observability of AI. By integrating these into a single operating model, we can help ensure that AI systems are running securely while keeping token usage performant and economically justified.

**How Cisco is mastering AI tokenomics through observability **

At Cisco, we optimized our AI-ready infrastructure and ensured our AI governance and security models operate at enterprise scale. Now, we are achieving economic control with agent observability.

We are building these capabilities on Splunk Agent Observability, supercharged by our acquisition of Galileo, to evaluate and improve agent behavior, observe AI performance, and optimize token costs. Agent Observability evaluates agent outputs using specialized small language models (SLMs) that judge quality at a fraction of the cost of frontier models, detect hallucinations, and enforce guardrails at runtime to block inaccurate and harmful behaviors. In addition, Agent Observability monitors performance across the entire AI stack, including models, GPUs, vector databases, memory and orchestration frameworks, and enables teams to track, forecast and optimize AI token usage and spend.

Centralizing this data into Splunk allows us to correlate infrastructure health, agent behavior, security events, and token costs across the entire enterprise and identify when inefficiencies in usage, infrastructure bottlenecks, or security risks are driving up our cost per token. For example, we’ve seen agents spiraling in a wave of unnecessary tool calls or appending unused skills, easily leading to bloated token consumption without benefiting the task that the agent carries out for the end user.

The power of this solution is in our stronger understanding of the relationship between AI spend and business value. By integrating our observability stack with internal business systems, we attribute AI spend directly to the ROI for specific teams, projects, and budget owners, all aligned to Cisco’s fiscal calendar and budget hierarchy.

This granularity allows engineering and finance leaders to make informed decisions. For example, by mapping token usage to actual code commits, we can measure the true ROI of our AI investments and gain the forecasting controls necessary to scale AI responsibly.

The road ahead: Follow along

At Cisco, we are proving this model at scale. We aren’t testing this in a lab; we are running it in production across our own global environment — consisting of over 2,000 live AI agents and more than 24,000 active users.

In the blogs that follow, we will share more details into our agent observability deployment, including deployment insights, results, and the lessons learned along the way.

Follow this series and use it to inform your own tokenomics strategy.

Richard Delisser is Vice President, Engineering at Cisco, where he leads AI adoption across software engineering — including AI-driven coding, testing, and autonomous technical debt monitoring.

Greg Sylvester is Vice President, Enterprise AI Platforms and Infrastructure at Cisco, where he leads Cisco’s Enterprise AI — including AI platforms, compute, storage, GPUs, cloud and data center operations, and AI observability and service management.

**Resources: **

Explore more Cisco on Cisco storiesCisco AI-Native InfrastructureCisco ObservabilitySplunk Agent ObservabilityCisco 2026 Research Report: AI Impact on Wide Area Networks

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @cisco systems inc. 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/the-ai-invoice-nobod…] indexed:0 read:5min 2026-08-04 ·