Can AI Replace Cloud Engineers? I Put It to the Test A developer tested how much cloud engineering work could be delegated to AI across AWS, Azure, and Google Cloud, running live Level 2 deployments of a Layer 7 load balancer fronting two private VMs and a Level 3 design-document trial that scored 77/100 and failed its acceptance criteria. The AI handled Terraform, GitHub Actions, federated authentication, remote state, monitoring, failure tests, recovery, and cleanup of 139 Terraform-managed items, while the human retained requirements, identity boundaries, and approvals. Three failures surfaced: unreachable startup dependencies broke Nginx on AWS, a CI apply silently stripped tags from 12 Azure resources, and an untracked backend declaration made 18 existing Google Cloud resources appear new. I tested how much of my cloud engineering work I could delegate to AI in two experiments. The live Level 2 experiments completed deployment, selected failure tests, recovery, and cleanup across AWS, Azure, and Google Cloud. The latest Level 3 trial, D5, scored 77/100 in an independent AI review: FAIL against the acceptance criteria, HOLD for adoption . The distinction matters throughout this article. Level 2 has live execution records. Level 3 evaluates design documents. Neither experiment measures the percentage of a profession that can be replaced. “Level” is my own classification, not an industry standard or an official mapping to cloud certifications. I did not run the introductory Level 1 stage. The common target was a Layer 7 load balancer serving HTTP through two private VMs. The VMs had no public IP addresses and no direct SSH access from the internet. | Component | AWS | Azure | Google Cloud | |---|---|---|---| | Public entry point | Regional ALB | Application Gateway | Global external Application Load Balancer | | Backend | Two EC2 instances in separate AZs | Two Linux VMs in separate zones | Two VMs in a regional managed instance group | | Outbound path | S3 Gateway Endpoint | NAT Gateway | Cloud Router and Cloud NAT | | Remote state | S3 | Blob Storage | Cloud Storage | These were implementations of common requirements, not identical architectures. Google Cloud used a managed instance group, while AWS and Azure used individual VMs. A global load balancer did not make the backend multi-region. The human retained responsibility for requirements, allowed changes, identity and exposure boundaries, and approvals. AI handled Terraform, GitHub Actions, short-lived federated authentication, remote state, monitoring, selected failure tests, recovery, and cleanup. I looked beyond a successful terraform apply : The execution records report completion across all three clouds. Cleanup removed 139 Terraform-managed items : AWS 48, Azure 53, and Google Cloud 38. These totals include the application infrastructure and bootstrap resources for state and authentication. That is not a count of VMs or individually billable resources. “Zero remaining” refers to the managed active resources checked in this experiment. It does not mean an empty account, physical erasure of every retained copy, or a guarantee of no later charges. See the Level 2 final report https://github.com/moruku36/cloud-validation-level2-multicloud/blob/c5d25b637c2a71ffb84b6a7dab6c5f0f994062a0/docs/final-report.md for the evaluation method and counts. AWS: the instance was private, but its startup dependencies were unreachable. Insufficient outbound connectivity prevented Nginx installation, and the ALB returned 502. AI diagnosed the problem and added narrowly scoped HTTPS access to the S3 prefix list. The lesson was to review startup dependencies alongside inbound isolation. This S3-based path depended on the selected OS package source; it was not general internet access. Azure: a successful apply removed tags from 12 resources. Local and CI inputs differed. The CI apply completed successfully while removing additional tags. The inputs were aligned, and a tag-only repair followed human approval. The exit status alone did not establish that the change matched the intended result. Google Cloud: missing backend configuration made 18 existing resources look new. The backend declaration was not tracked in Git. CI could not use the existing remote state and planned 18 additions. The process stopped before apply and repaired the declaration. Unexpected creation, deletion, or replacement needs to be a stop condition, not just another line in the output. These incidents and their recovery are documented in the failure analysis https://github.com/moruku36/cloud-validation-level2-multicloud/blob/c5d25b637c2a71ffb84b6a7dab6c5f0f994062a0/docs/final-report.md . Monitoring checks covered selected signals and alert states. External notification receivers were not configured, so delivery to an operator was not demonstrated. The failure scenarios also differed between providers. The experiments ran sequentially: AWS, then Azure, then Google Cloud, carrying lessons forward. Human-intervention counts used different units. That rules out a clean provider ranking or a defensible intervention-reduction percentage. There was also no standardized human-only baseline for time, cost, or quality. The final report identifies code and CI controls that need checking before reuse; successful historical execution is not proof that the current checkout reproduces unchanged. The supported conclusion is narrower: an experienced engineer could delegate a substantial implementation and operations lifecycle in small, disposable environments while retaining design, review, and approval. Production readiness, long-term SLOs, and unattended operation remain outside that result. The second experiment started with a member-facing application: registration, login, profile changes, list/detail views, favorites, and administrative updates. Its main constraints were: The RTO/RPO targets covered process, runtime, database, and single-AZ failures. Region-wide failure was excluded. These were design requirements , not observed service metrics. The customer answers https://github.com/moruku36/cloud-validation-level3-architect/blob/e1ad7ab1baa2dec111259dd37a8eef405d418746/docs/sources/business-answers-2026-09-08.md and requirements document https://github.com/moruku36/cloud-validation-level3-architect/blob/e1ad7ab1baa2dec111259dd37a8eef405d418746/docs/requirements.md preserve the input. Budget compliance could not depend on free tiers, temporary credits, or long-term commitment discounts. The tension between a 30-minute recovery target and weekday daytime staffing is a useful example: selecting managed services does not by itself establish who can act, which actions are authorized, or whether unattended recovery works. The September work was self-assessed: A moved from 61 to 65, and B scored 66. October introduced a reviewer separate from the designer. D4 and D5 used anonymously reviewed, frozen design-and-evidence packages. A passing design had to meet every condition: The mandatory categories were requirements, architecture, security/IAM, availability, cost, and backup/disaster recovery. No human scoring or adoption approval was obtained. These were independent AI document reviews , not production acceptance tests. | Trial | Requested designer configuration | Independent score | |---|---|---| | D1 | GPT-6 Luna / Low | 53 | | D2 | GPT-6 Luna / Medium | 72 | | D3 | GPT-6 Luna / High | 69 | | D4 | GPT-6.1 Sol / Medium | 78 | | D5 | GPT-6.1 Sol / High | 77 | These are requested settings . Actual runtime models and applied reasoning settings are UNKNOWN, including for the reviewers. The series also carries earlier findings forward. D5 revised D4, added controls and evidence, and used a fresh reviewer. The scores cannot isolate a model or reasoning-effort effect. September self-scores are a different evaluation method and are not part of this series. The latest authoritative trial is D5, dated October 2, 2026. It failed the acceptance criteria and remained on hold for adoption. The methods and results https://github.com/moruku36/cloud-validation-level3-architect/blob/e1ad7ab1baa2dec111259dd37a8eef405d418746/docs/research/METHODS-RESULTS.en.md and D5 report https://github.com/moruku36/cloud-validation-level3-architect/blob/e1ad7ab1baa2dec111259dd37a8eef405d418746/docs/reports/2026-10-02-d5/FINAL-REPORT.en.md explain the chronology and comparison limits. Requirements changed separately on October 2 remain unscored and were not D5 input. The score of 77 does not apply to those changed requirements. Cost was 3/5, below its mandatory minimum of 4. The design did not establish an all-inclusive budget fit with headroom, adequately supported effort for custom controls and rehearsals, or sufficiently concrete cheaper alternatives. Keeping unknown prices explicit was credited; unknowns did not become zero or evidence of affordability. Deletion gate G3 remained HOLD. Deletion-trigger events, retention of linking information, independent custody of evidence, and responsibility for external services or recipient copies were not approved. Operations gate G4 also remained HOLD. Named primary and backup owners, approved effort, absence coverage, and authority for unattended recovery were unresolved. Added controls brought an evidence burden. D5 added a SQL-backed AdmissionCoordinator to principal request paths. Evidence for operation counts, contention, latency, and recovery ordering was insufficient. Complexity fell from 4 to 3, accounting for the one-point total decrease from D4. Complexity 3 meets the general category floor; it is distinct from Cost failing a mandatory minimum. Unresolved P0 items concerned owner decisions, source access, and prerequisites for live validation, rather than identified P0 desk-design defects. The D5 report https://github.com/moruku36/cloud-validation-level3-architect/blob/e1ad7ab1baa2dec111259dd37a8eef405d418746/docs/reports/2026-10-02-d5/FINAL-REPORT.en.md and gate record https://github.com/moruku36/cloud-validation-level3-architect/blob/e1ad7ab1baa2dec111259dd37a8eef405d418746/docs/reports/2026-10-02-d5/evidence/independent-review-D5/gates.csv keep these distinctions explicit. Level 2 demonstrated more than initial code generation. AI investigated provider constraints, authentication mismatches, partial applies, and state inconsistencies, then made scoped repairs and completed cleanup. Level 3 showed increasingly specific design documents, but no accepted design. Missing prices, owner decisions, and empirical evidence remained missing. A score near 80 does not establish that acceptance is close when mandatory conditions remain unresolved. My practical takeaway is to specify the verification and decision process alongside the delegated task: A useful next experiment would hold inputs, evidence, and evaluation conditions constant while verifying actual model settings. Live performance, recovery, and cost validation would be a separate step. Those are research proposals, not work already completed. For now, the evidence supports broad delegation of predefined implementation under human control. It does not establish replacement of cloud engineering as a whole, or that a stronger model alone would satisfy the remaining adoption conditions. Source review date: October 6, 2026. References are pinned to Level 2 commit c5d25b6 and Level 3 commit e1ad7ab. The existing Japanese and English slides were also cross-checked. Preparing this article did not rerun cloud operations or produce a new assessment. AI assisted the structure, drafting, and English wording of this article. The experiment conditions, results, and limitations are based on the linked records.