cd /news/ai-infrastructure/google-cloud-postmortem-for-us-west1… · home topics ai-infrastructure article
[ARTICLE · art-113610] src=status.cloud.google.com ↗ pub= topic=ai-infrastructure verified=true sentiment=· neutral

Google Cloud postmortem for us-west1 August incident

Google Cloud reported that a scheduled fiber optic maintenance on August 20, 2026, from 08:00 to 10:22 US/Pacific caused a 2-hour-22-minute outage in its us-west1 region, impacting multiple services including Persistent Disk, Google Kubernetes Engine, Google Compute Engine, Cloud Run, Google Cloud Bigtable, Cloud Storage, and Identity and Access Management. The root cause was compromised network capacity between data centers, which automated rerouting failed to mitigate, leading to packet loss and cascading failures. Google Cloud is implementing safety checks and optimizing traffic routing to prevent recurrence.

read10 min views1 publishedAug 27, 2026

contact Support. Learn more about what's posted on the dashboard in this FAQ. For additional information on these services, please visit

[https://cloud.google.com/](https://cloud.google.com/).

For incidents related to Google Security Products, visit [https://status.cloud.google.com/security](https://status.cloud.google.com/security). For incidents related to Looker (original), visit [https://status.cloud.google.com/looker](https://status.cloud.google.com/looker).

Incident affecting AlloyDB for PostgreSQL, Apigee Edge Public Cloud, Apigee X, Artifact Registry, BigQuery Data Transfer Service, Cloud Build, Cloud Data Fusion, Cloud Filestore, Cloud Key Management Service, Cloud Monitoring, Cloud Run, Contact Center AI Platform, Dataproc Metastore, Google App Engine, Google BigQuery, Google Cloud Bigtable, Google Cloud Composer, Google Cloud Dataflow, Google Cloud Dataproc, Google Cloud Pub/Sub, Google Cloud SQL, Google Cloud Storage, Google Compute Engine, Google Kubernetes Engine, Identity and Access Management, Managed Service for Apache Kafka, Persistent Disk

We are investigating an issue where customers may experience timeouts, service degradations, errors, and elevated latencies across multiple products in the us-west1 region. #

Incident began at **2026-08-20 08:40** and ended at **2026-08-20 12:20** (all times are **US/Pacific**).

### Previously affected location(s)

GlobalOregon (us-west1)
Date Time Description
27 Aug 2026 14:45 PDT ## Incident Report## SummaryOn Thursday, 20 August 2026, from 08:00 to 10:22 US/Pacific (15:00 to 17:22 UTC), multiple Google Cloud services in the us-west1 region experienced elevated latency, provisioning failures, increased error rates, and widespread service degradations for a total duration of 2 hours and 22 minutes. The incident impacted a wide range of core services — including Persistent Disk, Google Kubernetes Engine, Google Compute Engine, Cloud Run, Google Cloud Bigtable, Cloud Storage, and Identity and Access Management — across both control plane operations and data plane requests. We sincerely apologize for the disruption this incident caused to your business operations and critical workloads. We recognize the vital role Google Cloud plays in supporting your organization, and we deeply regret the impact on your business operations and critical workloads. Engineering and infrastructure teams are actively implementing measures to address the root causes and strengthen network resiliency to prevent recurrences in the future. Specifically, teams are monitoring recovery progress, refining safety checks for planned optical maintenance, and optimizing regional traffic routing mechanisms to safeguard against unexpected capacity constraints. ## Root CauseThe disruption originated during scheduled fiber optic maintenance, which unexpectedly compromised network capacity between data centers within the us-west1 region. Automated rerouting mechanisms failed to properly redistribute traffic to alternate capacity, resulting in network congestion as volumes exceeded available bandwidth in the affected area. This underlying network degradation subsequently impacted higher-level service components through severe packet loss, request throttling, and increased latency across inter-campus dependencies. Core infrastructure services, including Spanner Paxos consensus and the Unified Metadata Server (UMS), experienced significant latency spikes, which cascaded into timeouts and elevated error rates for downstream dependent products such as Cloud Storage, Cloud IAM, Persistent Disk, and Google Kubernetes Engine. Consequently, both control plane operations and data plane requests failed to execute successfully across multiple Google Cloud services in us-west1 throughout the duration of the incident. ## Remediation and PreventionInternal monitoring systems initially detected widespread service anomalies and alerted Google engineers, who promptly confirmed that multiple core services operating across the us-west1 region were severely impacted. In response to the immediate operational risks, engineering teams promptly executed emergency traffic draining protocols to reroute active service workloads away from the compromised inter-campus network infrastructure and minimize further customer disruption. Once teams restored inter-campus fiber network capacity, engineers systematically reintroduced production traffic back to us-west1 through controlled validation stages, ultimately confirming full service normalization across all impacted platforms. Longer term engineering remediations and architectural enhancements are actively being finalized and assigned to owning teams. Key focus areas include reforming scheduled maintenance protocols, establishing stricter safety checks and circuit redundancy requirements prior to routine maintenance, and refining automated regional traffic failover mechanisms to automatically handle sudden capacity degradation without incurring severe network congestion. ## Detailed Description of ImpactOn Thursday, August 20, between 08:00 and 10:22 US/Pacific, Google Cloud customers in the us-west1 region encountered elevated latency, provisioning failures, and increased error rates.
The following services experienced elevated latencies and/or increased error rates, across their respective data planes and control planes:
The regional incident in us-west1 followed a structured three-phase recovery dictated by platform dependency layers: While core platform connectivity and live request serving were restored by 10:22 US/Pacific, certain services took additional time to fully recover due to asynchronous backlog processing and localized control-plane state reconciliation for a very small set of customers. For event-driven and pipeline services, inbound error rates dropped to 0 immediately, but some operations experienced elevated latency while workers processed through backlogs accumulated during the outage, clearing later for Cloud Pub/Sub, Cloud Build & Deploy, Cloud Dataflow, and Cloud Storage lifecycle deletions. Concurrently, while primary read/write traffic was healthy across the region, specific long-running lifecycle workflows required extra time to clear locks and reconcile distributed state machines. Compute Engine VM provisioning in zone us-west1-c and Cloud Filestore control plane instance allocation locks and resource validation checks took longer to normalize across regional storage backends.
24 Aug 2026 23:20 PDT ## Preliminary Incident ReportWe sincerely apologize for the disruption this incident caused to your business operations. Recognizing your reliance on Google Cloud, we express our sincere regrets for any operational impact experienced. Our engineering teams are actively addressing the underlying root cause to prevent future recurrences. Please note that the information provided herein reflects our current understanding as of the time of publication and remains subject to revision as the investigation progresses. A comprehensive Incident Report detailing preventive measures will be issued upon conclusion of our inquiry. If you have experienced impact outside of what is listed below, please reach out to Google Cloud Support using ## Date/Time of the Issue (All time US/Pacific)

SummaryOn Thursday, 20 August 2026, multiple Google Cloud services encountered elevated latency, provisioning failures, increased error rates, and service degradations lasting for a duration of 2 hours and 22 minutes. ## Preliminary Root CauseThe disruption originated during scheduled fiber optic maintenance, which unexpectedly compromised network capacity between data centers within the us-west1 region. Automated rerouting mechanisms failed to properly redistribute traffic to alternate capacity, resulting in network congestion as volumes exceeded available bandwidth in the affected area. This underlying network degradation subsequently impacted higher-level service components through request throttling, increased latency, and cascading retries. Consequently, both control plane operations and data plane requests failed to execute successfully across several dependent Google Cloud services, producing the observed latency and elevated error rates. ## RemediationInternal monitoring alerted Google engineers to the incident, confirming that multiple services across us-west1 were affected. To mitigate the immediate operational impact, engineers initiated traffic draining protocols, re-routing service workloads away from the degraded network infrastructure. Following the restoration of inter-campus network capacity and the resolution of the core issue, engineers systematically restored traffic to the us-west1 region and confirmed service normalization. Long-term engineering solutions are currently being finalized and assigned to address the root cause and strengthen system resilience regarding fiber maintenance procedures. ## Description of ImpactOn Thursday, August 20, between 08:00 and 10:22 US/Pacific, Google Cloud customers in the us-west1 region encountered elevated latency, provisioning failures, and increased error rates. #

The following services experienced elevated latencies and/or increased error rates, across their respective data planes and control planes: | | | 20 Aug 2026 | 12:37 PDT | The issue causing timeouts, degradations, errors, and latencies across multiple us-west1 products has been mitigated. We have mitigated the issue impacting multiple products in our us-west1 region as of Thursday, 2026-08-20 10:22 PDT. Our engineering teams have restored capacity on a planned optical maintenance that caused unexpected congestion in Dalles, Oregon metro /us-west1 region. Our systems stabilized and services have recovered once capacity was restored. Our teams are continuing to monitor for any residual impact. We will publish an analysis of this incident once we have completed our internal investigation. We thank you for your patience while we worked on resolving the issue. Customers in us-west1 may have experienced timeouts, service degradations, errors, and elevated latencies across multiple products. This issue is now mitigated. | | | 20 Aug 2026 | 12:21 PDT | The issue causing timeouts, degradations, errors, and latencies across multiple us-west1 products has been mitigated and we are working to recover all products. Mitigation actions have been completed by our engineering teams and we are seeing recovery from multiple products. We are continuing to work to recover the remaining products. We will provide an update by Thursday, 2026-08-20 13:30 PDT with details. Customers in us-west1 may experience timeouts, service degradations, errors, and elevated latencies across multiple products. No workarounds needed at this time. | | | 20 Aug 2026 | 11:53 PDT | The issue causing timeouts, degradations, errors, and latencies across multiple us-west1 products has been mitigated and we are working to recover all products. Mitigation actions have been completed by our engineering teams and we are seeing recovery from multiple products. We are continuing to work to recover the remaining products. We will provide an update by Thursday, 2026-08-20 12:30 PDT with details. Customers in us-west1 may experience timeouts, service degradations, errors, and elevated latencies across multiple products. No workarounds needed at this time. | | | 20 Aug 2026 | 11:03 PDT | The issue causing timeouts, degradations, errors, and latencies across multiple us-west1 products has been mitigated and we are working to recover all products. Mitigation actions have been completed by our engineering teams and we are seeing recovery from multiple products. We are continuing to work to recover the remaining products. We will provide an update by Thursday, 2026-08-20 11:45 PDT with details. Customers in us-west1 may experience timeouts, service degradations, errors, and elevated latencies across multiple products. No workarounds needed at this time. | | | 20 Aug 2026 | 10:32 PDT | We are investigating an issue where customers may experience timeouts, service degradations, errors, and elevated latencies across multiple products in the us-west1 region. Mitigation actions are implemented by our engineering teams, recovery trends have been observed across infrastructure layers. Active efforts remain underway to bring impacted cloud services back to full operation. We will provide an update by Thursday, 2026-08-20 11:00 PDT with details. Customers in us-west1 may experience timeouts, service degradations, errors, and elevated latencies across multiple products. We recommend customers to failover to other regions where feasible. | | | 20 Aug 2026 | 10:13 PDT | We are investigating an issue where customers may experience timeouts, service degradations, errors, and elevated latencies across multiple products in the us-west1 region. Mitigation actions are implemented by our engineering teams, recovery trends have been observed across infrastructure layers. Active efforts remain underway to bring impacted cloud services back to full operation. We will provide an update by Thursday, 2026-08-20 10:45 PDT with details. Customers in us-west1 may experience timeouts, service degradations, errors, and elevated latencies across multiple products. We recommend customers to failover to other regions where feasible. | | | 20 Aug 2026 | 09:44 PDT | We are investigating an issue where customers may experience timeouts, service degradations, errors, and elevated latencies across multiple products in the us-west1 region. We are experiencing an issue with multiple products, beginning on Thursday, 2026-08-20 08:40 PDT. Our engineering team continues to investigate the issue. We will provide an update by Thursday, 2026-08-20 10:30 PDT with details. We apologize to all who are affected by the disruption. Customers in us-west1 may experience timeouts, service degradations, errors, and elevated latencies across multiple products. None at this time. |

  • All times are US/Pacific
── more in #ai-infrastructure 4 stories · sorted by recency
── more on @google cloud 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/google-cloud-postmor…] indexed:0 read:10min 2026-08-27 ·