{"slug": "aws-cloudwatch-omni-goes-after-the-hardest-question-in-agentic-ai-why-did-the-do", "title": "AWS CloudWatch Omni goes after the hardest question in agentic AI: Why did the agent do that?", "summary": "Amazon Web Services made Amazon CloudWatch Omni generally available last week, an app-centric, OpenTelemetry-based observability service that embeds AI to answer why an agentic AI system produced a wrong answer, wrong tool call or stale knowledge-base result even when latency and error metrics look healthy. Omni ships with 17 built-in evaluators that score coherence, helpfulness, faithfulness and routing correctness, supports prompt-version comparison in a playground and one-click test-dataset creation from production traces, and offers a Visual Studio Code, Cursor and Kiro extension requiring no AWS account plus a standalone web console with single sign-on via Okta and Microsoft Entra ID. Sony's Masahiro Oba, senior general manager of the AI Acceleration Division, said the company's enterprise-wide agentic AI platform now supports hundreds of proof-of-concept and production workloads, and AWS cites an IDC forecast of more than 1 billion deployed agents by 2029.", "body_md": "### AWS CloudWatch Omni goes after the hardest question in agentic AI: Why did the agent do that?\n\nFor decades, the observability industry has answered one basic question: Is it running? Agentic artificial intelligence breaks that model. An agent can return a clean response, meet its latency target and throw no errors, yet still give a customer the wrong answer, call the wrong tool or pull from a stale knowledge base.\n\nBy every traditional metric, the system is healthy, but the only metric that matters to the business is that it failed. That’s the gap [Amazon Web Services Inc.](https://aws.amazon.com/) is targeting with [Amazon CloudWatch Omni](https://aws.amazon.com/cloudwatch/omni/) (pictured), which [became generally available](https://aws.amazon.com/blogs/aws/introducing-amazon-cloudwatch-omni-ai-powered-observability-for-generative-ai-and-agentic-workloads/) last week. AWS positions Omni as the next generation of CloudWatch.\n\nIt’s app-centric, runs outside the AWS Management Console, is built on OpenTelemetry and embeds AI. This shifts observability from “Is it running?” to “Why did my agent do that?” I’d argue that’s the question standing between most enterprises and production-scale agentic AI. AWS cites an IDC forecast of more than 1 billion deployed agents by 2029. No operations team can manually review that much nondeterministic behavior.\n\nHere are five considerations for information technology and business leaders:\n\n### Evaluation is the new monitoring\n\nThe most important part of Omni isn’t the dashboards; it’s the evaluation engine. Omni captures every trace and ships with 17 built-in evaluators that score coherence, helpfulness, faithfulness and routing correctness, among other metrics. Teams can compare prompt versions in a playground, build test datasets from production traffic, and automatically catch regressions. Evaluators can also run continuously against live traffic, so quality drift is flagged the same way a CPU spike would be. Latency and error rates can’t tell you whether an answer was right, but scoring can.\n\nSony is an early adopter. “At Sony, our enterprise-wide agentic AI platform now supports hundreds of proof-of-concept and production workloads,” said Masahiro Oba, senior general manager of the AI Acceleration Division at Sony. “At this scale, observability and evaluation are essential. With Amazon CloudWatch Omni, I can go from a single trace directly to evaluation, AI analysis, comparison or dataset creation.”\n\nThe key word in that statement is “hundreds.” Most companies I talk to aren’t stuck on building a single agent. They’re stuck on governing dozens or hundreds, each built by a different team with a different idea of what “good” looks like. Oba also noted that assembling evaluation datasets is often a business-side bottleneck, and that one-click dataset creation from live traces removes that bottleneck.\n\n### Getting out of the console is a bigger deal than it sounds\n\nDevelopers get a native extension for Visual Studio Code, Cursor and Kiro, where traces appear as they run an agent locally, with no AWS account required. Operators get a standalone web experience with single sign-on via existing identity providers such as Okta and Microsoft Entra ID. Both share a single data layer, so the trace a developer debugs is the same one an operator investigates.\n\nAWS acknowledges that its console is built for infrastructure administrators, not for the site reliability engineers, AI engineers and application owners who now carry operational responsibility. Meeting developers in the integrated development environment, where AI coding assistants such as Claude Code and Codex can set up instrumentation, shifts quality work to the point where problems are cheapest to fix.\n\n### The unified data layer is the real differentiator\n\nMany startups can trace large language model calls. What’s interesting about Omni is that agent traces, application telemetry and infrastructure signals all live in the same CloudWatch data store. As a result, an investigation can start with an agent receiving a bad tool result, move to an application programming interface error from a capacity-limited service, and end with an exhausted database connection pool.\n\nIn most shops today, that’s three tools, three teams and a lot of meetings to connect the dots. AWS DevOps Agent is enabled by default in investigation sessions, correlating signals and maintaining a full investigation history.\n\nCapital One was a design partner. “Capital One operates one of the largest observability footprints in financial services,” said Parvez Naqvi, managing vice president of cloud platform and resilience engineering at Capital One. “As a design partner for Amazon CloudWatch Omni, we helped shape a single AI-powered observability solution that will give our engineers topology-aware intelligence and natural-language querying across all telemetry from a single surface, with full data ownership through OpenTelemetry.”\n\nFor heavily regulated industries, such as banking, data portability is critical to success, and that requires data ownership. Captured investigation history is also underrated, since in regulated industries it’s audit evidence of how an AI incident was handled.\n\n### Open standards lower the lock-in bar but don’t eliminate it\n\nInstrumentation runs on OpenInference and the AWS Distro for OpenTelemetry, whether agents run on AWS or in other clouds. Omni supports LangChain, LangGraph, CrewAI, the OpenAI Agents SDK, Strands and the Vercel AI SDK, as well as third-party evaluators such as DeepEval. Amazon Bedrock AgentCore agents receive Omni automatically. Azure ingestion is supported today, with deeper multicloud coverage coming.\n\nThat openness is necessary in a crowded field. Datadog, Dynatrace, New Relic, Grafana Labs and Splunk are all expanding into agent observability, as are large language model-focused tools such as LangSmith and Arize AI.\n\nOpenTelemetry makes the data portable, but the intelligence layer isn’t. The topology, evaluators, investigation history and DevOps Agent all run on AWS. Companies already heavily invested in AWS will see Omni as the natural default. Those running mature Datadog or Splunk estates across multiple clouds will likely use it for agent development and evaluation, while keeping their observability system of record where it is.\n\n### Pricing is built for adoption, so watch the telemetry bill\n\nThe IDE extension is free. Customers pay for the telemetry they send and store. Dashboards and alerts are free, and queries up to five times the monthly ingestion volume are included. Eligible accounts receive a 30-day trial and $1,000 in OpenTelemetry ingestion credits.\n\nThe catch is that agents are chatty. Every prompt, model call, tool invocation and sub-agent handoff generates a span. Multiply that across hundreds of workloads and continuous evaluation, and ingestion costs can outpace the AI budget that created them. DevOps Agent is also priced separately.\n\n### What this means for buyers\n\nOmni is a comprehensive agent observability platform that addresses the trust gap that keeps agents in pilot mode. IT leaders should:\n\n- **Define “good” before you buy the evaluator.** Built-in scoring only helps if your business owners have documented what a correct, compliant and helpful answer looks like for each agent. Most haven’t.\n- **Standardize instrumentation now.** Put every agent pilot on OpenTelemetry, regardless of the backend. That keeps your options open and makes future platform decisions much easier.\n- **Model telemetry costs at production scale.** Set policies for sampling, retention and evaluation frequency before agents go live, not after the first bill arrives.\n- **Decide where your system of record resides.** If AWS is your primary cloud, Omni is a strong default choice. If you’re multicloud with an established observability platform, consider Omni for development and evaluation, and integrate it rather than replacing your existing setup.\n- **Treat investigation history as part of governance.** Integrate Omni’s captured investigation trail into AI risk and compliance processes, especially in regulated industries.\n\nThe industry spent a decade learning to observe distributed systems. Agentic AI requires observing decisions, not just systems, and AWS aims to own that layer for its customers. The companies that get the most out of it will be those that treat evaluation as an operational discipline, not a feature to be switched on.\n\n*Zeus Kerravala is a principal analyst at ZK Research, a division of Kerravala Consulting. He wrote this article for SiliconANGLE.*\n\n##### Image: AWS\n\n# A message from John Furrier, co-founder of SiliconANGLE:\n\nSupport our mission to keep content open and free by engaging with theCUBE community. **Join theCUBE’s Alumni Trust Network**, where technology leaders connect, share intelligence and create opportunities.\n\n- **15M+ viewers of theCUBE videos** , powering conversations across AI, cloud, cybersecurity and more\n- **11.4k+ theCUBE alumni** — Connect with more than 11,400 tech and business leaders shaping the future through a unique trusted-based network\n\n### Are you an AWS customer?  Support SiliconANGLE financially by buying your AWS services from our Marketplace portal page and links: [https://siliconangle.com/aws-marketplace/](https://siliconangle.com/aws-marketplace/)\n\n##### **About SiliconANGLE Media**\n\n[SiliconANGLE](https://cts.businesswire.com/ct/CT?id=smartlink&url=https%3A%2F%2Fsiliconangle.com%2F&esheet=54119777&newsitemid=20240910506833&lan=en-US&anchor=SiliconANGLE&index=9&md5=646b1b564e2259100a2b8638aab0a552),\n\n[theCUBE Network](https://cts.businesswire.com/ct/CT?id=smartlink&url=https%3A%2F%2Fwww.thecube.net%2F&esheet=54119777&newsitemid=20240910506833&lan=en-US&anchor=theCUBE+Network&index=10&md5=7de2a85f95ab4a4a495cede20b8cb1da),\n\n[theCUBE Research](https://cts.businesswire.com/ct/CT?id=smartlink&url=https%3A%2F%2Fthecuberesearch.com%2F&esheet=54119777&newsitemid=20240910506833&lan=en-US&anchor=theCUBE+Research&index=11&md5=7bb33676722925eb57d588ec343e4f6f),\n\n[CUBE365](https://cts.businesswire.com/ct/CT?id=smartlink&url=https%3A%2F%2Fwww.cube365.net%2F&esheet=54119777&newsitemid=20240910506833&lan=en-US&anchor=CUBE365&index=12&md5=d310fb35919714e66ad8d42c9c0c1bc6),\n\n[theCUBE AI](https://cts.businesswire.com/ct/CT?id=smartlink&url=https%3A%2F%2Fwww.thecubeai.com%2F&esheet=54119777&newsitemid=20240910506833&lan=en-US&anchor=theCUBE+AI&index=13&md5=b8b98472f8071b23ebb10ab9a8dd0683)and theCUBE SuperStudios — with flagship locations in Silicon Valley and the New York Stock Exchange — SiliconANGLE Media operates at the intersection of media, technology and AI.\n\nFounded by tech visionaries John Furrier and Dave Vellante, SiliconANGLE Media has built a dynamic ecosystem of industry-leading digital media brands that reach 15+ million elite tech professionals. Our new proprietary theCUBE AI Video Cloud is breaking ground in audience interaction, leveraging theCUBEai.com neural network to help technology companies make data-driven decisions and stay at the forefront of industry conversations.", "url": "https://wpnews.pro/news/aws-cloudwatch-omni-goes-after-the-hardest-question-in-agentic-ai-why-did-the-do", "canonical_source": "https://siliconangle.com/2026/09/27/aws-cloudwatch-omni-goes-after-the-hardest-question-in-agentic-ai-why-did-the-agent-do-that/", "published_at": "2026-09-28 02:06:54+00:00", "updated_at": "2026-09-28 02:46:39.340132+00:00", "lang": "en", "topics": ["ai-agents", "ai-infrastructure", "mlops", "developer-tools", "artificial-intelligence"], "entities": ["Amazon Web Services", "Amazon CloudWatch Omni", "Sony", "Masahiro Oba", "IDC", "Visual Studio Code", "Cursor", "Okta"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/aws-cloudwatch-omni-goes-after-the-hardest-question-in-agentic-ai-why-did-the-do", "markdown": "https://wpnews.pro/news/aws-cloudwatch-omni-goes-after-the-hardest-question-in-agentic-ai-why-did-the-do.md", "text": "https://wpnews.pro/news/aws-cloudwatch-omni-goes-after-the-hardest-question-in-agentic-ai-why-did-the-do.txt", "jsonld": "https://wpnews.pro/news/aws-cloudwatch-omni-goes-after-the-hardest-question-in-agentic-ai-why-did-the-do.jsonld"}}