How we eliminated $1 million a year of wasted AI agent spend in one hour Databricks engineers used Unity Gateway tracing and Genie One to identify and fix seven MCP-server bugs in one hour, eliminating an estimated $1.2 million a year in wasted AI agent spend, including $499K in wasted tokens and 12,000 engineering hours annually. The fixes were shipped by coding agents after the team analyzed tool call traces and used natural language queries to surface the biggest cost drivers. Unity Gateway tracing plus Genie One turned our agents' tool failures into a ranked, fixable bug list, helping us identify and eliminate an estimated $1.2 million a year in wasted AI spend and lost productivity. • Broken MCP tool calls silently cost real money. Across our agent fleet, seven small MCP-server bugs burned ~$499K/year in tokens and 12,000 eng-hours/year $1.2M lost because agents quietly retry instead of surfacing failures. • Observe, then fix. Unity Gateway traces every MCP tool call, while Genie One lets teams surface the biggest sources of wasted AI spent using natural language. Our coding agents shipped the fixes in one hour from start to finish. • Design tools for how LLMs actually use them. Models make guesses on ambiguous inputs, so tools should handle variations gracefully rather than crash on unexpected inputs. Databricks engineers rely heavily on AI agents to streamline and accelerate their work. In turn, these agents require access not only to different Foundation Models but also to MCP servers with tools that enable access to relevant artifacts e.g., system logs, usage tables, support tickets, wikis . In a previous blog https://www.databricks.com/blog/managing-ai-coding-costs-scale , we shared that managing AI costs at scale requires optimizing not only model selection but also how agents use tools. In this post, we describe how we looked for cost savings in our agents' use of tools, the challenges we hit along the way, and how OTel tracing in Unity Gateway cut the path from analysis to $1.2M/year in savings to a single hour. Enabling our developers to build their agents was a huge unlock on productivity, but as usage ramped up, we also faced increasing costs. We started investigating several optimizations, and one suspicion that we had was the hidden cost of failing tool calls. Specifically, when tools misbehave, the calling agent rarely fails loudly. Instead, it retries, guesses, and eventually works around the problem, quietly burning tokens and developer time the whole way. This type of waste is dangerous: from the outside, the task still completes, and an aggregate cost dashboard may show a 10% bump in token spend that can be easily misinterpreted as usage growth. We investigated this suspicion in our agent fleet using Unity Gateway's tracing https://docs.databricks.com/aws/en/ai-gateway/unified-trace-table and Genie One https://www.databricks.com/product/genie/one . We found seven small bugs in our tool servers that were costing an estimated $499K/year in wasted tokens and about 12,000 engineering hours per year in agent wait time. Overall, this is an estimated $1.2M/year in lost productivity. Finding all seven bugs, quantifying them, and fixing them took about an hour. This post describes the process we followed and what it taught us about building tools for agents. When we first deployed AI agents widely at Databricks for coding and internal workflows, it was impossible to manage or even fully understand costs because we lacked visibility into the agents’ tool calls and overall activity. To solve this, we leveraged Unity Gateway, which automatically emits an OpenTelemetry trace for all MCP tool invocations, including the tool name, arguments, error if any , token counts, latency, and a session ID that ties calls together. Those traces land in a single table that records exactly what our agents did over any time window. No new instrumentation was required, and the gateway already sits on the path of every call, so the data was readily available. This makes AI agent cost management more actionable, where instead of seeing only aggregate token spend, we can attribute wasted spend to specific tools, errors, and agent sessions. Now that the data is available, the next step is exploration: Normally, the expensive part of this kind of analysis is the SQL and the schema spelunking. But with Genie One, we just pointed it at the trace table, asked these exact questions in plain English , and got answers back in minutes. Most of our hour went to reading those answers rather than writing queries. Genie One turned a vague suspicion "agents seem to thrash on Jira calls" into a ranked, quantified bug list in minutes. Here is an example from a single 24-hour window, showing bugs in our Jira and Google Drive/Docs tool servers: Bug | Errors/day | Annual token cost | Annual wait time | Repeat rate | Jira: KeyError: 'fields' get | 137 | $250K | 2,500 h | ~30% | Jira: 'list' object has no attribute 'split' | 535 | $87K | 4,850 h | 30.5% | Jira: KeyError: 'fields' search | 32 | $58K | 580 h | ~30% | GDrive: Invalid field selection | 417 | $46K | 2,740 h | 54.5% | Jira: unexpected analysis prompt kwarg | 121 | $42K | 840 h | 50.0% | GDocs: find text required | 137 | $15K | 440 h | 14.3% | Jira: quote from bytes expected bytes | 30 | $1.2K | 73 h | 66.7% | | | | | n/a | Take the highest-volume bug, 535 failures a day, as an example. The Jira issues.search tool takes a fields parameter, and the server did this: It expected a comma-separated string like "key,summary,status". But an array is the semantically natural JSON type for "a list of fields," and that is what the model inferred from its background knowledge of JSON conventions and from adjacent tool calls in the same session. So it passed the structured value that a reasonable caller would: A list has no .split , so the server raised 'list' object has no attribute 'split', a raw Python traceback that tells the agent nothing about what it did wrong. So the agent guessed again. Sometimes it retried the same list and failed the same way; sometimes it re-read the schema or fell back to trial and error. On average, it took 12 turns to recover, and 30% of sessions hit the error more than once. One .split call was costing an estimated $87K/year in tokens and 4,850 hours of agent wait time. The Google Drive Invalid field selection error was even more striking in volume: 49.6% of all drive file get calls failed , because the model kept passing valid-looking Drive API field names id, name, mimeType that the tool's endpoint did not accept. The obvious takeaway is "write better error messages," and the data backs it up. Recovery cost tracks error-message quality almost perfectly: Error message quality | Example | Repeat rate | Avg turns to recover | Self-documenting | "find text and replace text required" | 14% | 4.6 | Somewhat informative | "Missing required parameters: org, repo" | ~30% | 4 | Cryptic traceback | "'list' object has no attribute 'split'" | 30.5% | 12.1 | Misleading | "unexpected keyword argument 'analysis prompt'" | 50% | 13.1 | But "good error messages help" is old news. The more interesting question is why the model called these tools "wrong" in the first place. In most of these cases, it didn't. MCP tool signatures are often deliberately under-specified. We keep them loose on purpose: partly for generality, and partly to save context tokens, since every parameter description costs tokens the model pays for on every call. The consequence is that when a signature is vague about fields, the model fills the gap with a reasonable guess, and a JSON array is a reasonable guess for a list of fields. The bug was not that the model called the tool incorrectly. It was that the server accepted only one of several reasonable interpretations and crashed on the rest. So the design principle is the reverse of the reflexive one: tools for agents should adapt to the way LLMs naturally call them, e.g., coerce the list into a string, default the omitted parameter, absorb the unexpected argument, and so on. An under-specified signature is a promise of flexibility, and the tool should honor that promise on the receiving end rather than crash on the first input that doesn't match the one shape its author had in mind. The fixes themselves were simple and are not the interesting part of this story. Once Genie One had handed us a ranked list of which errors to fix and what the model was actually sending, applying the fixes across the tool servers was a quick pass with a coding agent. The whole loop find, quantify, fix took about an hour. The scarce, expensive step was never writing the fix. It was knowing what to fix. Tracing plus Genie One turned that step from a research project into a question you can ask out loud. As more real work shifts onto agents, silent tool failures become a first-class cost center, the kind that hides inside "usage growth" and never pages anyone. The loop for catching them is cheap and repeatable: Unity Gateway makes agent behavior observable, and Genie One makes that behavior queryable without SQL. Together, this gives teams a repeatable way to monitor AI agents, diagnose MCP tool failures, and reduce wasted AI spend. If you run agents against your own tools, do the same. Trace the calls and ask Genie One what keeps going wrong. Unity Gateway is Generally Available, and you can now monitor all AI activity using the unified trace table, which is now in Beta. See our docs https://docs.databricks.com/aws/en/ai-gateway/unified-trace-table on how to get started. Subscribe to our blog and get the latest posts delivered to your inbox.