cd /news/ai-agents/introducing-advanced-ai-observabilit… · home › topics › ai-agents › article
[ARTICLE · art-142641] src=konghq.com ↗ pub= topic=ai-agents verified=true sentiment=↑ positive

Introducing Advanced AI Observability: See the Whole Agent Journey

Kong previewed advanced AI observability capabilities in Kong Konnect that make the session, rather than the individual request, the unit of AI observability, according to Alex Drag, Kong's Head of Product Marketing. The preview stitches together model calls, tool and MCP interactions, agent activity, policy and guardrail execution, token consumption, cost, and latency so teams can trace where an agent's behavior diverged across a multi-turn session. Kong says the capabilities let teams analyze model success and failure rates, input and output token consumption, TTFT and TPOT, cache hit rates, and cost per request and session across models, providers, agents, users, and MCP servers.

by read10 min views1 publishedSep 30, 2026
Introducing Advanced AI Observability: See the Whole Agent Journey
Image: Konghq (auto-discovered)

Alex Drag

Head of Product Marketing

Trace complete AI sessions across models, tools, agents, and infrastructure — and understand where performance, cost, and failures accumulate.

Traditional observability was built around requests. A request comes in. A service processes it. A response goes out. You trace what happened in between. That model works well for APIs. It breaks down for AI.

With agents, a single user request can trigger multiple model calls. An agent might invoke a tool, call another agent, query an MCP server, retry a model, hit a guardrail, and call another tool before finally producing a response. And all of this happens non-determinstically.

And then the user asks another question.

The individual requests might all look healthy. Yet by turn seven, the agent has gone completely sideways.

The problem isn’t necessarily in a request. It’s in the journey.

Today, we’re previewing new advanced AI observability capabilities in Kong Konnect that make the session,not the individual request,the unit of AI observability.

Stop debugging AI one request at a time #

When an AI application fails, knowing that an individual model call returned a 200 isn’t enough.

You need to know what led to it.

Kong stitches the interactions associated with an AI session together, giving developers a view across the sequence of model calls, tool invocations, policies, and other operations involved in completing the task.

Instead of reconstructing an AI interaction from disconnected requests and logs, teams can follow the execution path from beginning to end.

For each session, developers can inspect signals including:

  • Model and LLM calls

  • Tool and MCP interactions

  • Agent activity

  • Policy and guardrail execution

  • Inputs and outputs

  • Token consumption

  • Cost

  • Latency and duration

  • Successes and failures

You can move from “the agent failed” to understanding where the behavior diverged and what happened before it did.

Follow the session, not just the request #

Consider an agent tasked with investigating a customer issue.

It might first call a model to interpret the request. Then query customer data through an API. Then invoke an MCP tool. Then call another model. Then execute a PII policy. Then synthesize the final response.

Traditional request-level telemetry gives you fragments of that workflow.

Kong gives you the session.

Developers can start with the complete AI interaction and drill down through the individual operations that comprise it, inspecting inputs, outputs, attributes, models, tokens, costs, and latency along the way.

That makes it possible to investigate AI systems in the way they actually operate.

See where AI performance breaks down #

AI introduces an entirely new set of operational questions.

Which models are failing most frequently? Which agents are driving latency? Where are tokens being consumed? Are MCP tools introducing failures? Is semantic caching actually improving performance? Which requests are outliers?

Kong brings these signals together across models, providers, agents, users, and MCP servers.

Teams can analyze:

  • Model success and failure rates

  • Input and output token consumption

  • TTFT and TPOT

  • Request and session latency

  • Cache hit rates

  • Agent and model usage

  • MCP server activity

  • Cost per request and session

  • Performance and cost outliers

Instead of treating AI telemetry as a separate collection of model-provider dashboards, teams get a common operational view across their AI infrastructure.

See how cost accumulates across the AI journey #

AI cost has the same problem as AI observability.A provider bill tells you what individual model calls cost. It doesn’t necessarily tell you what it cost to accomplish something.

An agentic task might involve five model calls, three tools, a retry, and another agent before producing an outcome.

Kong connects those interactions into the same session, making it possible to see how tokens, latency, and cost accumulate across the complete AI interaction.

That creates a natural connection between AI observability and AI cost management: understand not only what happened, but what the complete journey cost.

From observing AI to debugging it #

Seeing what happened is only part of the problem.

Developers also need a way to investigate why.

Kong is developing a new Konnect Debugger that will complement session-level tracing with deeper, live investigation of AI traffic.

Rather than relying exclusively on post-hoc dashboards after something has gone wrong, developers will be able to move from observing AI behavior to actively interrogating it.

Together, session tracing and the Konnect Debugger create a much deeper debugging experience:

Observe → Trace → Investigate → Debug

See the entire AI journey. Find the point where behavior changed. Inspect what the model, tool, agent, or policy saw. Then investigate the interaction in context.

One observability layer across generative and agentic AI #

The complexity of AI systems will only increase.

Generative AI applications are becoming multi-turn. Agents are becoming multi-tool. Agentic workflows are becoming multi-agent. And all of them increasingly depend on APIs, MCP servers, models, policies, and enterprise infrastructure working together.

Observability needs to evolve with them.

Kong’s advanced AI observability capabilities are designed to provide a common view across that environment, from individual model calls to complete agentic sessions.

Because when an AI system goes wrong, the question isn’t simply:

“Which request failed?”

It’s:

“What happened across the entire journey, and why?”

Kong Advanced AI Observability, including session-level tracing, is coming soon to Kong Konnect.

Multi-agent systems are, at their core, context distribution systems. Every agent in your workflow is a consumer and producer of context. The interesting architectural questions are all about how that context moves. Two operations drive everything:

Hugo Guerrero

API management solved a similar problem for developers. Instead of asking a platform team every time they needed an API, developers could go to a Developer Portal, discover what was available to them, and start building — within the governance bound

Alex Drag

Most AI cost management starts with infrastructure. A provider can tell you that you consumed a certain number of input and output tokens on a particular model. An observability platform can show requests, latency, tokens, and traces. A cloud cost p

Alex Drag

Turning an API into an MCP server sounds straightforward: expose each API operation as an MCP tool and let the agent decide which one to use. For small APIs, that can work well. But enterprise APIs can contain hundreds of operations. And as agents c

Alex Drag

The explosion of AI-native applications is upon us. With each new week, massive innovations are being made in how AI-centric applications are being built. There are a variety of tools developers need to consider, be it supplying live contextual data

If you're running agents in production, you've felt this problem: every agent, every LLM application, every MCP client needs to be configured against a growing sprawl of MCP servers, each with its own endpoint, its own handshake, and its own access Greg Peranich

Kong Mesh 2.14 also introduces improvements to the mesh-scoped zone proxy deployment model. This makes it easier to configure and operate zone proxies for specific meshes, including Helm support for mesh zone proxy configuration. For customers runni

Justin Davies

Multi-agent systems are, at their core, context distribution systems. Every agent in your workflow is a consumer and producer of context. The interesting architectural questions are all about how that context moves. Two operations drive everything:

Hugo Guerrero

API management solved a similar problem for developers. Instead of asking a platform team every time they needed an API, developers could go to a Developer Portal, discover what was available to them, and start building — within the governance bound

Alex Drag

Most AI cost management starts with infrastructure. A provider can tell you that you consumed a certain number of input and output tokens on a particular model. An observability platform can show requests, latency, tokens, and traces. A cloud cost p

Alex Drag

Turning an API into an MCP server sounds straightforward: expose each API operation as an MCP tool and let the agent decide which one to use. For small APIs, that can work well. But enterprise APIs can contain hundreds of operations. And as agents c

Alex Drag

The explosion of AI-native applications is upon us. With each new week, massive innovations are being made in how AI-centric applications are being built. There are a variety of tools developers need to consider, be it supplying live contextual data

If you're running agents in production, you've felt this problem: every agent, every LLM application, every MCP client needs to be configured against a growing sprawl of MCP servers, each with its own endpoint, its own handshake, and its own access Greg Peranich

Kong Mesh 2.14 also introduces improvements to the mesh-scoped zone proxy deployment model. This makes it easier to configure and operate zone proxies for specific meshes, including Helm support for mesh zone proxy configuration. For customers runni

Justin Davies

Multi-agent systems are, at their core, context distribution systems. Every agent in your workflow is a consumer and producer of context. The interesting architectural questions are all about how that context moves. Two operations drive everything:

Hugo Guerrero

API management solved a similar problem for developers. Instead of asking a platform team every time they needed an API, developers could go to a Developer Portal, discover what was available to them, and start building — within the governance bound

Alex Drag

Most AI cost management starts with infrastructure. A provider can tell you that you consumed a certain number of input and output tokens on a particular model. An observability platform can show requests, latency, tokens, and traces. A cloud cost p

Alex Drag

Turning an API into an MCP server sounds straightforward: expose each API operation as an MCP tool and let the agent decide which one to use. For small APIs, that can work well. But enterprise APIs can contain hundreds of operations. And as agents c

Alex Drag

The explosion of AI-native applications is upon us. With each new week, massive innovations are being made in how AI-centric applications are being built. There are a variety of tools developers need to consider, be it supplying live contextual data

If you're running agents in production, you've felt this problem: every agent, every LLM application, every MCP client needs to be configured against a growing sprawl of MCP servers, each with its own endpoint, its own handshake, and its own access Greg Peranich

Kong Mesh 2.14 also introduces improvements to the mesh-scoped zone proxy deployment model. This makes it easier to configure and operate zone proxies for specific meshes, including Helm support for mesh zone proxy configuration. For customers runni

Justin Davies

Ready to see Kong in action? #

Get a personalized walkthrough of Kong's platform tailored to your architecture, use cases, and scale requirements.

── more in #ai-agents 4 stories · sorted by recency
── more on @kong 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/introducing-advanced…] indexed:0 read:10min 2026-09-30 · —