For decades, the observability industry has answered a fundamental question: Is it running? Agentic artificial intelligence breaks this model. An agent can return a clean response, meet its latency target, and not throw errors, but still give a customer the wrong answer, invoke the wrong tool, or rely on an outdated knowledge base.
By any traditional metric the system is healthy, but the only metric that matters to the company is that it has failed. That’s the gap Amazon Web Services Inc. aims with Amazon CloudWatch Omni (pictured), which became generally available last week. AWS positions Omni as the next generation of CloudWatch.
It is application-centric, runs outside the AWS management console, is based on OpenTelemetry, and integrates AI. This shifts the observability of “Is it running?” to “Why did my agent do that?” I would argue that this is the question that stands between most companies and production-scale agent AI. AWS cites an IDC prediction of more than 1 billion deployed agents by 2029. No operations team can manually inspect that much nondeterministic behavior.
Here are five considerations for information technology and business leaders:
Evaluation is the new monitoring
The most important part of Omni isn’t the dashboards; It is the evaluation engine. Omni captures every trace and comes with 17 built-in evaluators that evaluate coherence, helpfulness, fidelity, and routing correctness, among other things. Teams can compare prompt versions in a playground, create test datasets from production traffic, and automatically detect regressions. Evaluators can also work continuously with live traffic, so quality deviations are flagged in the same way a CPU spike would be. Latency and error rates can’t tell you whether an answer was correct, but scoring can.
Sony is an early adopter. “At Sony, our enterprise-wide agent AI platform now supports hundreds of proof-of-concept and production workloads,” said Masahiro Oba, senior general manager of the AI Acceleration Division at Sony. “At this scale, observability and evaluation are essential. With Amazon CloudWatch Omni, I can go from a single track directly to evaluation, AI analysis, comparison or data set creation.”
The key word in this statement is “hundreds.” Most companies I talk to don’t limit themselves to building a single agent. They rely on governing dozens or hundreds, each built by a different team with a different idea of what “good” looks like. Oba also noted that compiling evaluation datasets is often a business bottleneck and one-click dataset creation from live traces eliminates this bottleneck.
Leaving the console is a bigger deal than it sounds
Developers get a native extension for Visual Studio Code, Cursor, and Kiro that displays traces when they run an agent locally without requiring an AWS account. Operators get a standalone web experience with single sign-on via existing identity providers such as Okta and Microsoft Entra ID. Both share a single data layer, so the trace a developer debugs is the same one an operator examines.
AWS recognizes that its console was designed for infrastructure administrators, not site reliability engineers, AI engineers, and application owners who now have operational responsibility. Engaging developers in the integrated development environment, where AI coding wizards like Claude Code and Codex can set up instrumentation, shifts the quality of work to the point where problems are most cost-effective to fix.
The unified data layer is the real differentiator
Many startups can track large language model calls. What’s interesting about Omni is that agent tracking, application telemetry, and infrastructure signals are all stored in the same CloudWatch data store. As a result, an investigation can start with an agent receiving a bad tool result, move from a service with limited capacity to an application programming interface error, and end with an exhausted database connection pool.
In most businesses today that means three tools, three teams and lots of meetings to clarify the connections. AWS DevOps Agent is enabled by default in investigation sessions, correlates signals, and maintains a complete investigation history.
Capital One was the design partner. “Capital One has one of the largest observability footprints in financial services,” said Parvez Naqvi, managing vice president of cloud platform and resilience engineering at Capital One. “As a design partner for Amazon CloudWatch Omni, we helped create a single AI-powered observability solution that provides our engineers with topology-aware information and natural language queries across all telemetry from a single interface, with full data ownership through OpenTelemetry.”
For highly regulated industries like banking, data portability is critical to success, and that requires data ownership. The investigation history recorded is also underestimated, as in regulated industries it is audit evidence of how an AI incident was handled.
Open standards reduce the level of lock-in, but do not eliminate it
Instrumentation runs on OpenInference and the AWS Distro for OpenTelemetry, regardless of whether agents are running on AWS or in other clouds. Omni supports LangChain, LangGraph, CrewAI, the OpenAI Agents SDK, Strands, and the Vercel AI SDK, as well as third-party evaluators such as DeepEval. Amazon Bedrock AgentCore agents receive Omni automatically. Azure ingestion is supported today, with broader multicloud coverage to follow.
This openness is necessary in a crowded environment. Datadog, Dynatrace, New Relic, Grafana Labs, and Splunk are all expanding into agent observability, as are major language model-focused tools like LangSmith and Arize AI.
OpenTelemetry makes the data portable, but the intelligence layer does not. Topology, evaluators, investigation history, and DevOps agent all run on AWS. Companies that have already invested heavily in AWS will see Omni as a natural default. Those running mature Datadog or Splunk environments across multiple clouds will likely use it for agent development and evaluation while keeping their observability system of record where it is.
Prices are based on acceptance, so pay attention to the telemetry bill
The IDE extension is free. Customers pay for the telemetry they send and store. Dashboards and alerts are free and queries up to five times the monthly ingestion volume are included. Eligible accounts will receive a 30-day trial and $1,000 OpenTelemetry enrollment credit.
The catch is that agents are talkative. Every prompt, model call, tool call, and subagent handoff generates a margin. Multiply that by hundreds of workloads and continuous evaluation, and the cost of ingestion can exceed the AI budget that incurred it. The DevOps agent is also charged separately.
What this means for buyers
Omni is a comprehensive agent observability platform that closes the trust gap that keeps agents in pilot mode. IT managers should:
- Define “good” before purchasing the evaluator. Built-in scoring only helps if your business owners have documented what an accurate, compliant, and helpful response looks like for each agent. Most didn’t.
- Standardize instrumentation now. Put every agent pilot on OpenTelemetry, regardless of the backend. This keeps your options open and makes future platform decisions much easier.
- Model telemetry costs at production scale. Establish guidelines for sampling, retention, and evaluation frequency before agents go live, rather than after the first invoice is received.
- Decide where your recording system is located. If AWS is your primary cloud, Omni is a good default choice. If you run a multicloud with an established observability platform, consider and integrate Omni for development and evaluation rather than replacing your existing setup.
- Treat investigation history as part of governance. Integrate Omni’s investigation path into AI risk and compliance processes, especially in regulated industries.
The industry spent a decade learning to observe distributed systems. Agentic AI requires observation of decisions, not just systems, and AWS wants to own that layer for its customers. The companies that get the most value from this will be those that treat assessment as an operational discipline rather than a feature that needs to be enabled.
Zeus Kerravala is a Principal Analyst at ZK Research, a division of Kerravala Consulting. He wrote this article for SiliconANGLE.
Image: AWS
Support our mission to keep content open and free by interacting with theCUBE community. Join theCUBE Alumni Trust Networkwhere technology leaders connect, share information and create opportunities.
- Over 15 million viewers of theCUBE videosto spark conversations about AI, cloud, cybersecurity and more
- Over 11.4k theCUBE alumni — Connect with more than 11,400 technology and business leaders shaping the future through a unique, trusted network
Founded by tech visionaries John Furrier and Dave Vellante, SiliconANGLE Media has built a dynamic ecosystem of industry-leading digital media brands, reaching over 15 million elite tech professionals. Our new proprietary theCUBE AI Video Cloud is a game-changer in audience engagement, leveraging the theCUBEai.com neural network to help technology companies make data-driven decisions and stay at the forefront of the industry conversation.