On October 8, 2026, customers of Google Cloud Storage in the us‑central1 region experienced prolonged write latency increases and higher error rates that persisted for more than half a day.
On October 8, 2026, Google Cloud Storage customers in the us‑central1 region faced a twelve‑hour surge in write latency accompanied by elevated error rates, disrupting data‑ingest workflows across numerous enterprises.
Incident timeline
The performance degradation began shortly after 02:00 UTC and persisted until roughly 14:30 UTC, affecting both standard and near‑line storage classes. Users reported write operations taking several seconds longer than typical, with some attempts failing outright.
Root cause analysis
Google’s post‑mortem identified a misconfiguration in the region’s internal load‑balancing layer that throttled write requests. The issue was isolated to a subset of storage nodes handling high‑throughput workloads, leading to queue buildup and timeout errors.
The misconfiguration was introduced during a routine firmware rollout and was not caught by automated validation tests, highlighting a gap in the deployment pipeline for regional services.
Mitigation steps
Engineers rolled back the faulty configuration and re‑balanced traffic across healthy nodes, restoring normal write latency within an hour. Additional monitoring alerts were added to detect similar load‑balancing anomalies more quickly.
- Reverted the problematic load‑balancer settings
- Redistributed traffic to unaffected storage nodes
- Enhanced validation checks for future rollouts
- Implemented real‑time latency alerting for the us‑central1 region
Impact on customers
The extended latency window primarily affected batch processing jobs and real‑time data pipelines that rely on timely writes. While no data loss was reported, some customers experienced delayed analytics and increased operational costs due to retry logic.
We saw our nightly ETL jobs run over an hour longer than expected, which pushed downstream reporting into the next business day.
Google has offered affected customers a service credit for the incident period and is working with them to review retry strategies and error handling best practices.
Incident report on Fru.dev detailing Google Cloud Storage latency spike
Comments
No comments yet.