Skip to Main Content

February 2026

7

Watch our 'Azure Incident Retrospective' video about this incident:

What happened?

Between 07:58 UTC on 07 February 2026 and 04:24 UTC on 08 February 2026, impacted customers experienced intermittent service unavailability, timeouts and/or higher than normal latency for services in the West US region. The event began following a power interruption affecting one of the datacenters within the region, after which impact manifested as infrastructure availability loss and service disruptions across multiple dependent workloads in the region.

As power stabilization progressed, recovery proceeded in phases – however, a subset of storage and compute infrastructure components did not return to a healthy state immediately after power restoration, which slowed recovery for dependent components, and contributed to ongoing symptoms such as delayed telemetry and resource recovery.

Impacted services included:

  • Azure AI Search: Customers using Azure AI Search in West US may have experienced query failures, indexing delays, or temporarily unavailable search endpoints.
  • Azure App Service: Between 07:58 UTC and approximately 09:45 UTC on 07 February 2026, customers using Azure App Service in West US may have experienced HTTP 500 errors, request timeouts, or failed deployments.
  • Azure Backup: During the impact window, backup jobs in West US may have experienced delays or temporary failures due to storage unavailability.
  • Azure Cache for Redis: Customers may have observed connectivity failures, increased latency, or transient cache unavailability in West US.
  • Azure Container Registry: Customers may have experienced delays or failures when pulling or pushing container images in West US.
  • Azure Cosmos DB: Customers may have experienced a degradation in service availability and/or request latency. Some requests may have resulted in server errors or timeouts.
  • Azure Data Factory: Customers may have observed pipeline execution failures or delays in West US. Affected integration runtimes recovered following restoration of underlying networking and storage services.
  • Azure Database for MySQL – Flexible Server: Between 09:40 UTC and 20:50 UTC on 07 February 2026, customers may have experienced connection failures or degraded database availability in West US.
  • Azure Database for PostgreSQL – Flexible Server: Customers may have experienced database connectivity interruptions or degraded performance in West US.
  • Azure Databricks: Customers may have experienced workspace unavailability, cluster startup failures, or job execution interruptions in West US.
  • Azure DevOps: Customers may have experienced delays in log ingestion and analytics reporting in West US.
  • Azure Event Hubs: Between 07:58 UTC and 19:05 UTC on 7 February 2026, customers may have experienced failures or high latency when sending and processing events. Control plane operations like creating/updating/deleting Event Hubs or Consumer Groups were also impacted.
  • Azure IoT Hub: Customers may have observed delayed message processing or temporary endpoint unavailability in West US.
  • Azure Kubernetes Service (AKS): Between 07:58 UTC and 21:46 UTC on 07 February 2026, customers using AKS in West US may have experienced control plane instability, node health reporting delays, or pod scheduling failures.
  • Azure Key Vault: Customers may have experienced intermittent authentication or secret retrieval failures in West US.
  • Azure Monitor: Customers may have observed delayed metrics, log ingestion latency, or temporary alerting gaps in West US.
  • Azure Monitor (Application Insights): Between 07:58 UTC on 07 February 2026 and 04:24 UTC on 08 February 2026, customers may have experienced delayed telemetry ingestion and incomplete monitoring data in West US.
  • Azure Service Bus: Between 07:58 UTC and 19:05 UTC on 7 February 2026, customers may have experienced transient messaging failures, delayed message processing, or control plane disruptions impacting queues and topics in West US.
  • Azure SQL Database: Customers may have experienced SQL error 40613 (database unavailable) or connectivity interruptions in West US. Service mitigation included database failover and host recovery.
  • Azure Storage: Customers may have experienced read/write failures, increased latency, or temporary storage account unavailability in West US. Storage services were restored in phases as power and network infrastructure stabilized.
  • Azure Stream Analytics: Customers may be experiencing service availability disruptions, including delays or failures in data ingestion, processing, or output delivery.
  • Azure Virtual Machines: Customers may have experienced VM start failures, allocation errors, or temporary unavailability of running instances in West US.
  • Azure Virtual Machines (Confidential VM types): Customers may have experienced VM deployment or startup failures in West US, due to dependent storage disruption.
  • Microsoft Defender for Cloud Apps: Customers may have observed delayed alerts or telemetry ingestion in West US during the recovery period.

What went wrong and why?

Under normal operating conditions, Azure datacenter infrastructure is designed to tolerate utility power disturbances through redundant electrical systems and automated failover to on-site generators. During this incident, although utility power was still active, an electrical failure in an onsite transformer resulted in the loss of utility power to the datacenter. Although generators started as designed, a cascading failure within a control system prevented the automated transfer of load from utility power to generator power. As a result, Uninterruptible Power Supply (UPS) batteries carried the load for several minutes – until they were fully depleted, leading to customer impact from this power loss as early as 07:58 UTC.

Once power loss occurred, recovery was not uniform across all dependent infrastructure. Our datacenter operations team restored power to IT racks by leveraging our onsite generators, with 90% of IT racks powered by 09:31 UTC on 07 February. While infrastructure generally came back online as expected after power restoration, subsets of networking, storage and compute infrastructure required further electrical control system troubleshooting, before power could be restored. We were fully restored and running generator-backed power by 11:29 UTC. Power restoration was sequenced to avoid destabilizing power that had already been brought back online, and due to the inconsistent state of the control system.

Following power restoration at the datacenter, recovery of customer workloads progressed in stages based on dependencies. Network and Storage recovery were critical requirements, as compute hosts must be able to access persistent storage to complete boot, validate system state, and reattach customer data before workloads can resume. Network devices and services recovered as expected once power was restored.

The power loss affected six storage scale units within the datacenter, of which four recovered as expected once power was restored. Two storage scale units experienced prolonged recovery, which became a primary factor contributing to extended impact as these contained a large subset of storage nodes that failed to complete the boot process. During the boot process, critical artifacts need to be downloaded, but these were delayed or timed out – resulting in these nodes entering a state for which there was not an automated recovery action. Manual operations were attempted, and eventually succeeded in recovering the necessary nodes to bring these storage scale units back online. Due to the dependencies that many compute and platform services have on these storage scale units, there was a delay in overall service restoration. Storage recovery completed by 19:05 UTC.

Compute recovery began once power and storage dependencies were partially restored. Automated recovery systems were available and functioning, but the scale of the event resulted in a surge of required recovery activities across the datacenter. This placed sustained pressure on regional shared infrastructure components responsible for coordinating health checks, power state validation, and repair actions. During the initial recovery window, this elevated load reduced the speed at which individual compute nodes could be recovered. While many nodes returned to service as expected, some nodes required repeated recovery attempts, due to insufficient prioritization across concurrent fault categories during the surge of recovery activities. We introduced optimizations to accelerate recovery, including deprioritizing non-critical recovery activities. As pressure on these shared systems decreased, compute recovery progressed more predictably and remaining nodes were gradually restored. Compute recovery completed by 23:30 UTC.

How did we respond?

  • 07:54 UTC on 07 February 2026 – Initial electrical failure in an onsite transformer, although various UPS batteries carried the load for several minutes.
  • 07:58 UTC on 07 February 2026 – Customers began experiencing unavailability and delayed monitoring/log data.
  • 08:07 UTC on 07 February 2026 – Datacenter and engineering teams engaged and initiated coordinated investigation across power, network, storage, and compute recovery workstreams.
  • 09:31 UTC on 07 February 2026 – 90% of facility power restored on generators.
  • 11:26 UTC on 07 February 2026 – 93% of facility power restored on generators.
  • 11:29 UTC on 07 February 2026 – 100% of facility power restored on generators with continued recovery of hosted services.
  • 12:15 UTC on 07 February 2026 – Storage recovery started and dependent services started seeing recovery.
  • 15:00 UTC on 07 February 2026 – Targeted remediation actions continued for remaining unhealthy services.
  • 15:25 UTC on 07 February 2026 – Applied configuration changes to optimize processing of recovery events, to assist in expediting the recovery of compute services.
  • 19:05 UTC on 07 February 2026 – Storage recovery was complete.
  • 19:24 UTC on 07 February 2026 – Compute recovery system stabilized and fully operational, continued gradually restoring remaining nodes.
  • 23:30 UTC on 07 February 2026 – Long tail of Compute nodes impact mitigated.
  • 04:24 UTC on 08 February 2026 – All affected services confirmed mitigated.
  • 03:42 UTC on 09 February 2026 – Power was transitioned back to utility power, restoring the datacenter to normal operations.

How are we making incidents like this less likely or less impactful?

  • The control system and dependent components have been inspected, reconfigured, and validated to support transfer-to-generator events more reliably. (Completed)
  • We have reevaluated how best to prioritize critical infrastructure earlier in the power restoration sequence, to ensure a smoother recovery across dependent services. (Completed)
  • We are working to enhance the recovery system through prioritization and throttling, to improve stability during surge scenarios like this one. (Estimated completion: July 2026)
  • Finally, we are continually refining how our infrastructure and services come back online after datacenter power loss events, to make these more robust to different failure scenarios.

How can customers make incidents like this less impactful?

How can we make our incident communications more useful?

You can rate this PIR and provide any feedback using our quick 3-question survey: