# I Gave Three Broken AWS Environments to an AI On-Call Engineer. Here's What AWS DevOps Agent Found

> Source: <https://dev.to/dmitriy_trunov_9a09a497b1/i-gave-three-broken-aws-environments-to-an-ai-on-call-engineer-heres-what-aws-devops-agent-found-3lk4>
> Published: 2026-10-11 19:50:16+00:00

It’s 2:00 AM. A CloudWatch alert triggers. The on-call engineer wakes up, opens five console tabs, and begins reconstructing the timeline of events from scratch.

The **AWS DevOps Agent** is designed to handle that initial hour of painstaking analysis. AWS refers to it as a "frontier agent"- one that not only resolves incidents but prevents them before they occur. It analyzes CloudWatch metrics, CloudTrail data, VPC Flow Logs, RDS Performance Insights, and IaC stacks, subsequently providing a root-cause analysis and an action plan for remediation.

I was curious to see how this works in practice, so I attended an AWS DevOps Agent workshop focused on incident investigation. The workshop consisted of four labs: one on system setup and three involving specially prepared environments with pre-injected faults. I attended this engaging event - organized by the **AWS User Group 3City** community in collaboration with **AWS**-some time ago.

*The AWS DevOps Agent console: "Resolve. Prevent. Improve."*

| Scenario | What I broke | What the agent concluded | Time to root cause | 
|---|---|---|---|
| **EC2 / Auto Scaling** | Ran a CPU-burn script on the ASG through SSM | The CPU alarm had **no actions attached** and the ASG had**no scaling policy** , so the group couldn't scale out | ~10 min | 
| **Network security** | Port scan + HTTP flood from my laptop | A security group open to `0.0.0.0/0` on 22/80/443/8080 since stack creation. It also found that**3 of 4 earlier alarm firings were false positives** caused by S3 return traffic | ~20–25 min | 
| **RDS MySQL** | 140 connections + a terrible `CROSS JOIN … ORDER BY RAND()` | A runaway workload from the bastion host using the `admin` user. It also spotted**broken CloudWatch log export** | ~16 min | 

In all three labs the agent did more than restate the alarm. It found the **configuration gap underneath it**, tied the fix back to the **CloudFormation stack** so the fix wouldn't become drift, and wrote a mitigation plan in five phases: *prepare → pre-validate → apply → post-validate → rollback*.

You create an **Agent Space**. This is the boundary for what the agent can see and do. In the setup wizard you:

Once the space exists, you can add more **capabilities**: secondary AWS accounts, Azure, GitHub/GitLab/Azure DevOps pipelines, Datadog/Dynatrace/New Relic/Splunk/Grafana telemetry, Slack/ServiceNow/PagerDuty, any **MCP server**, and even **remote agents over the A2A protocol**.

👉 Full walkthrough: [Part 1: Setting up AWS DevOps Agent](https://dev.to/dmitriy_trunov_9a09a497b1/aws-devops-agent-part-1-setting-up-your-first-agent-space-2ib3)

A Lambda used SSM Run Command to push one t3.micro instance to 100% CPU. The CloudWatch alarm went red, and the group stayed at one instance.

About 1m48s in, the agent had its first finding: **`ScalingPolicies: []`** and **` AlarmActions: []`**. The alarm worked. It just wasn't wired to anything.

It also:

`CPUCreditBalance` stayed at the t3.micro maximum of 288
👉 [Part 2: EC2 CPU spike investigation](https://dev.to/dmitriy_trunov_9a09a497b1/aws-devops-agent-part-2-why-didnt-my-auto-scaling-group-scale-ma0)

I ran a script from my laptop that port-scanned the web server, enumerated admin pages, tried SSH logins and flooded it with HTTP requests.

The agent rated it **Severity: high**. Using VPC Flow Logs, it identified the source IP, listed the scanned ports, and confirmed that **2,764 accepted packets on port 80** were the real attack.

The part that stood out was **historical context** I never asked for. Looking at earlier firings of the same alarm, it worked out that 3 of the 4 were **benign S3 return traffic**, which means the alarm itself needs tuning.

👉 [Part 3: Network security investigation](https://dev.to/dmitriy_trunov_9a09a497b1/aws-devops-agent-part-3-i-attacked-my-own-web-server-heres-what-the-ai-found-in-the-vpc-flow-2hdc)

A script on the bastion host opened about 140 connections against a `max_connections=150` limit and pinned CPU with Cartesian joins.

The investigation was excellent. It split the load using **Performance Insights** into 108 sessions sitting in `SLEEP()` and 38 AAS of the `CROSS JOIN` query, and traced all of it to a single host and DB user.

Then I asked follow-up questions in chat and found **two contradictions** between the chat and the mitigation plan. The two disagreed on whether 150 is above or below the default `max_connections`, and on whether `wait_timeout` would help. Part 4 covers which one was right.

👉 [Part 4: RDS performance investigation](https://dev.to/dmitriy_trunov_9a09a497b1/aws-devops-agent-part-4-rds-connection-exhaustion-and-why-you-should-still-check-the-ais-answers-2ai6)

*Screenshots are from my own run of the AWS DevOps Agent workshop on 3 October 2026. All resources were in a temporary workshop account that has since been deleted.*
