cd /news/ai-agents/i-gave-three-broken-aws-environments… · home › topics › ai-agents › article
[ARTICLE · art-149287] src=dev.to ↗ pub= topic=ai-agents verified=true sentiment=↑ positive

I Gave Three Broken AWS Environments to an AI On-Call Engineer. Here's What AWS DevOps Agent Found

A developer tested the AWS DevOps Agent against three deliberately broken AWS environments — an Auto Scaling CPU spike, a port scan and HTTP flood, and an RDS MySQL connection exhaustion — and the agent traced each incident to its underlying configuration gap in roughly 10 to 25 minutes. In the EC2 lab it found the CloudWatch alarm had no actions attached and the ASG had no scaling policy; in the network lab it flagged a security group open to 0.0.0.0/0 and identified three of four prior alarm firings as false positives from S3 return traffic; in the RDS lab it attributed the overload to a runaway bastion-host workload and spotted broken CloudWatch log export. The agent tied each fix back to the CloudFormation stack and produced a five-phase mitigation plan: prepare, pre-validate, apply, post-validate, rollback.

by read4 min views1 publishedOct 11, 2026

It’s 2:00 AM. A CloudWatch alert triggers. The on-call engineer wakes up, opens five console tabs, and begins reconstructing the timeline of events from scratch.

The AWS DevOps Agent is designed to handle that initial hour of painstaking analysis. AWS refers to it as a "frontier agent"- one that not only resolves incidents but prevents them before they occur. It analyzes CloudWatch metrics, CloudTrail data, VPC Flow Logs, RDS Performance Insights, and IaC stacks, subsequently providing a root-cause analysis and an action plan for remediation.

I was curious to see how this works in practice, so I attended an AWS DevOps Agent workshop focused on incident investigation. The workshop consisted of four labs: one on system setup and three involving specially prepared environments with pre-injected faults. I attended this engaging event - organized by the AWS User Group 3City community in collaboration with AWS-some time ago.

The AWS DevOps Agent console: "Resolve. Prevent. Improve."

Scenario What I broke What the agent concluded Time to root cause
EC2 / Auto Scaling Ran a CPU-burn script on the ASG through SSM The CPU alarm had no actions attached and the ASG hadno scaling policy , so the group couldn't scale out ~10 min
Network security Port scan + HTTP flood from my laptop A security group open to 0.0.0.0/0 on 22/80/443/8080 since stack creation. It also found that3 of 4 earlier alarm firings were false positives caused by S3 return traffic ~20–25 min
RDS MySQL 140 connections + a terrible CROSS JOIN … ORDER BY RAND() A runaway workload from the bastion host using the admin user. It also spottedbroken CloudWatch log export ~16 min

In all three labs the agent did more than restate the alarm. It found the configuration gap underneath it, tied the fix back to the CloudFormation stack so the fix wouldn't become drift, and wrote a mitigation plan in five phases: prepare → pre-validate → apply → post-validate → rollback.

You create an Agent Space. This is the boundary for what the agent can see and do. In the setup wizard you:

Once the space exists, you can add more capabilities: secondary AWS accounts, Azure, GitHub/GitLab/Azure DevOps pipelines, Datadog/Dynatrace/New Relic/Splunk/Grafana telemetry, Slack/ServiceNow/PagerDuty, any MCP server, and even remote agents over the A2A protocol.

👉 Full walkthrough: Part 1: Setting up AWS DevOps Agent A Lambda used SSM Run Command to push one t3.micro instance to 100% CPU. The CloudWatch alarm went red, and the group stayed at one instance.

About 1m48s in, the agent had its first finding: ScalingPolicies: [] and AlarmActions: []. The alarm worked. It just wasn't wired to anything.

It also:

CPUCreditBalance stayed at the t3.micro maximum of 288

👉 Part 2: EC2 CPU spike investigation I ran a script from my laptop that port-scanned the web server, enumerated admin pages, tried SSH logins and flooded it with HTTP requests.

The agent rated it Severity: high. Using VPC Flow Logs, it identified the source IP, listed the scanned ports, and confirmed that 2,764 accepted packets on port 80 were the real attack.

The part that stood out was historical context I never asked for. Looking at earlier firings of the same alarm, it worked out that 3 of the 4 were benign S3 return traffic, which means the alarm itself needs tuning.

👉 Part 3: Network security investigation A script on the bastion host opened about 140 connections against a max_connections=150 limit and pinned CPU with Cartesian joins.

The investigation was excellent. It split the load using Performance Insights into 108 sessions sitting in SLEEP() and 38 AAS of the CROSS JOIN query, and traced all of it to a single host and DB user.

Then I asked follow-up questions in chat and found two contradictions between the chat and the mitigation plan. The two disagreed on whether 150 is above or below the default max_connections, and on whether wait_timeout would help. Part 4 covers which one was right.

👉 Part 4: RDS performance investigation Screenshots are from my own run of the AWS DevOps Agent workshop on 3 October 2026. All resources were in a temporary workshop account that has since been deleted.

── more in #ai-agents 4 stories · sorted by recency
── more on @aws 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/i-gave-three-broken-…] indexed:0 read:4min 2026-10-11 · —