September 28, 2026·8 min read·AIgentic.media

AWS CloudWatch Omni rewrites the question for AI agents: 'Why did the agent do that?'

awsai-observabilityagentic-aicloudwatchtooling
AWS CloudWatch Omni rewrites the question for AI agents: 'Why did the agent do that?'

For decades, observability asked one question. Agentic AI just made it the wrong one.

Every few years, a technology arrives that quietly rewrites a question everyone had stopped asking.

For thirty years, observability tools answered: "Is it running?" The answer was a dashboard green — latency under threshold, error count zero, uptime 99.9%. That was enough. Software was deterministic: the same input produced the same output, and a healthy system was a correct system.

Agentic AI breaks that model entirely.

An agent can return a clean response, meet every latency target, and throw no errors — yet still give a customer the wrong answer, call the wrong tool, or pull from a stale knowledge base. By every traditional metric, the system is healthy. By the only metric that matters to the business, it failed.

AWS is launching Amazon CloudWatch Omni to change the question. Instead of "Is it running?", the new question is: "Why did my agent do that?" And that shift might be the single biggest bottleneck standing between most enterprises and production-scale agentic AI.

The old question was built for deterministic software

Observability as an industry was born in the age of monoliths and matured through distributed microservices, containers, and serverless. In every paradigm, the core question stayed the same: is the system healthy? Developers instrumented code, dashboards lit up green or red, and pager duty triggered on anomaly thresholds.

This worked because software was predictable. A web server returning HTTP 200 was serving the right content. A database query completing in 50ms was returning the right rows. There was no gap between "system is healthy" and "output is correct."

Agents destroy that assumption.

An agent orchestrates LLM calls, tool invocations, sub-agent handoffs, and reasoning chains to fulfill an open-ended goal. A prompt change that seems innocent can degrade response quality without any visible impact on latency or error rates. The agent can pick the wrong tool, pull from the wrong knowledge source, or call the wrong API — all while reporting healthy metrics.

An IDC forecast cited by AWS projects more than 1 billion deployed agents by 2029. No operations team can manually review that much non-deterministic behavior. But trust requires visibility, and visibility for agents requires a fundamentally different observability model.

AWS CloudWatch Omni: what it actually does

CloudWatch Omni — which became generally available last week — is AWS's answer to this trust gap. It is an app-centric, AI-powered observability platform built on OpenTelemetry that operates outside the AWS Management Console.

The most important part is not the dashboards. It is the evaluation engine.

CloudWatch Omni captures every agent trace — every LLM call, tool invocation, reasoning step, and sub-agent handoff — as structured spans in a hierarchical timeline. Then it runs 17 built-in evaluators that score each response on dimensions like coherence, helpfulness, faithfulness, routing correctness, and retrieval quality.

Teams can compare prompt versions side by side in a playground, build test datasets directly from production traces, and run experiments across different model and prompt configurations to catch regressions before they reach customers. Evaluators can also run continuously against live traffic, so quality drift is flagged the same way a CPU spike would be.

Instrumentation runs on OpenInference and the AWS Distro for OpenTelemetry, and works whether agents run on AWS or in other clouds. The platform supports LangChain, LangGraph, CrewAI, the OpenAI Agents SDK, Strands, and the Vercel AI SDK. Amazon Bedrock AgentCore agents receive Omni instrumentation automatically.

Two surfaces, one data layer

CloudWatch Omni delivers observability through two complementary interfaces.

Developers get a native extension for Visual Studio Code and Kiro, where traces appear as they run an agent locally — with no AWS account required for basic use. The extension integrates with AI coding assistants like Claude Code and Codex, which can set up instrumentation automatically. A "Cloud Login" feature optionally connects the local environment to an AWS account for persistent storage and team sharing.

Operators get a standalone web experience with single sign-on via existing identity providers such as Okta and Microsoft Entra ID. Both surfaces share a single data layer, so the trace a developer debugs locally is the same one an operator investigates in production.

AWS acknowledges that its Management Console was built for infrastructure administrators — not for the site reliability engineers, AI engineers, and application owners who now carry operational responsibility for agents. Meeting developers in the IDE, where quality problems are cheapest to fix, is a deliberate design choice.

Early adopters: Sony, Capital One, and the reality of "hundreds" of agents

Sony is an early adopter. Masahiro Oba, senior general manager of the AI Acceleration Division at Sony, described the company's agentic AI platform as running "hundreds of proof-of-concept and production workloads."

The key word is "hundreds." Most companies are not stuck on building a single agent. They are stuck on governing dozens or hundreds, each built by a different team with a different idea of what "good" looks like. Oba noted that assembling evaluation datasets is often a business-side bottleneck, and that one-click dataset creation from live traces removes that friction.

Capital One participated as a design partner. Parvez Naqvi, managing vice president of cloud platform and resilience engineering at Capital One, said the bank helped shape a single AI-powered observability solution that provides "topology-aware intelligence and natural-language querying across all telemetry from a single surface." For regulated industries, the captured investigation history doubles as audit evidence — a critical feature when compliance teams need to reconstruct how an AI incident was handled.

The competitive landscape

CloudWatch Omni enters a crowded field. Datadog, Dynatrace, New Relic, Grafana Labs, and Splunk are all expanding into agent observability. LLM-focused tools like LangSmith and Arize AI have specialized tracing capabilities.

AWS's differentiation is integration depth. Because agent traces, application telemetry, and infrastructure signals all live in the same CloudWatch data store, an investigation can start with an agent receiving a bad tool result, move to an API error from a capacity-limited service, and end with an exhausted database connection pool. In most shops today, that is three tools, three teams, and a lot of meetings.

The AWS DevOps Agent is enabled by default in investigation sessions, correlating signals across the entire stack and maintaining a full investigation history.

But the intelligence layer is not portable. The topology, evaluators, investigation history, and DevOps Agent all run on AWS. Companies already heavily invested in AWS will see Omni as the natural default. Those running mature Datadog or Splunk estates across multiple clouds will likely use it for agent development and evaluation while keeping their observability system of record where it is.

The catch: agents are chatty

There is a catch: agents generate an extraordinary volume of telemetry. Every prompt, model call, tool invocation, and sub-agent handoff creates a span. Multiply that across hundreds of workloads with continuous evaluation running, and ingestion costs can outpace the AI budget that created them.

The IDE extension is free. Customers pay for the telemetry they send and store. Dashboards and alerts are free, and queries up to five times the monthly ingestion volume are included. Eligible accounts receive a 30-day trial and $1,000 in OpenTelemetry ingestion credits. The DevOps Agent is priced separately.

AWS recommends that organizations model telemetry costs at production scale and set policies for sampling, retention, and evaluation frequency before agents go live — not after the first bill arrives.

What this means for enterprise AI adoption

CloudWatch Omni is a recognition from the largest cloud provider that agentic AI requires a new operational discipline. The industry spent a decade learning to observe distributed systems. Observing AI agents requires observing decisions, not just system health.

The companies that get the most out of Omni will be those that treat evaluation as an operational discipline, not a feature to be switched on. They will define what "good" looks like before they buy the evaluator. They will standardize instrumentation on OpenTelemetry regardless of the backend. And they will integrate captured investigation history into AI risk and compliance processes.

The trust gap between "the agent ran successfully" and "the agent did the right thing" is the single largest barrier to production AI agent deployment right now. AWS just positioned itself to own the layer that closes that gap.

Sources

Want to learn more?

Let's discuss how AI can transform your business.

Explore AI Agents