Producto:
Región:
Fecha:
julio de 2026
23
Join either of our upcoming 'Azure Incident Retrospective' livestreams discussing this incident (to hear from our engineering leaders, and to get any questions answered by our experts) or watch a recording of the livestream (available the following week, on YouTube):
- Option #1 - 17:30 to 18:00 UTC on 27 August 2026 - https://aka.ms/air/ZJV6-SGG/1
- Option #2 - 05:30 to 06:00 UTC on 28 August 2026 - https://aka.ms/air/ZJV6-SGG/2
What happened?
Between 14:44 and 19:41 UTC on 23 July 2026, a subset of customers experienced connectivity failures, increased latency, and/or difficulty accessing Azure services hosted within the West US region. Impact was limited to network traffic entering or exiting the region, traffic that remained entirely within the region was not affected.
Impacted services included: App Service, Application Insights, Azure AD B2C, Azure AI Bot Service, Azure AI Search, Azure API Management, Azure Application Gateway, Azure Bastion, Azure Cosmos DB, Azure Data Explorer, Azure Database for PostgreSQL, Azure Databricks, Azure ExpressRoute, Azure Firewall, Azure Kubernetes Service (AKS), Azure Monitor, Azure Speech in Foundry Tools, Azure Virtual Desktop, Azure Virtual WAN, Azure VMware Solution, Azure VPN Gateway, as well as Microsoft Sentinel and Power BI Embedded. Downstream impact to Microsoft 365 services has been communicated via MO1437424.
What went wrong and why?
Azure datacenters connect through redundant network paths, ensuring that traffic can continue flowing even when one path is unavailable for maintenance or repair. When unplanned maintenance – including ‘break-fix’ hardware repair – is needed on an optical device along one of these paths, the repair process starts by creating system-readable instructions, verifying that the redundant path remains healthy, and performing safety checks to confirm that the work will not affect customer traffic. Because break-fix activities on a single device are generally expected to be impactless, these requests are processed automatically without human engagement or triage.
On 23 July 2026, a break-fix repair was initiated on an optical device to address a network reliability risk. A defect in our blast radius analysis system incorrectly expanded the scope of the repair event to include all optical devices egressing a specific datacenter. The safety validation step, which is designed to confirm that at least one of the two redundant datacenter paths remains available, ran but incorrectly concluded the operation was safe. The checks validated each device individually rather than evaluating the aggregate effect of isolating all devices at once, a scenario that was not accounted for because the system was never designed to process a full datacenter's worth of devices in a single request. As a result, routes were withdrawn from multiple devices simultaneously, disrupting connectivity between the datacenter and the WAN – therefore impacting traffic entering or leaving the West US region.
Once the route withdrawals took effect at 14:44 UTC, physical links and routing adjacencies continued to appear healthy, which initially masked the correlation between the break-fix activity and the connectivity disruption. The impact presented as a WAN routing anomaly, as third-party networks could not reach Azure in the region, rather than as a datacenter connectivity failure. Although all physical work in the region was stopped, our engineers could not correlate to this recent change because the preparation activities in advance of the break-fix did not succeed, so the physical layer and traffic appeared healthy.
Meanwhile, strong signals indicating WAN impact drew much of the investigative focus. Two parallel investigation streams developed: one focused on the WAN routing anomaly, and another examining network device health from within the datacenter. It took additional time to work past these misleading signals and eventually identify the break-fix activity as the source. It was not until 17:19 UTC that our engineers were able to correlate the routing anomaly with the break-fix activity, at which point we had identified the trigger event and began planning a rollback.
Our automated recovery and rollback system detected the device failures, and attempted multiple retries to restore the affected devices. However, because that system depended on the same datacenter connectivity that had been disrupted, its automated rollback attempts were unsuccessful. After our engineers correlated the routing anomaly with the break-fix activity, we identified that the incomplete preparation activities established the trigger conditions. We initiated a manual rollback at 17:45 UTC, then datacenter-to-WAN connectivity was restored when the rollback completed at 18:26 UTC, and all impacted services fully recovered by 19:41 UTC.
How did we respond?
- 14:44 UTC on 23 July 2026 – Device break-fix began. Earliest potential customer impact time.
- 14:45 UTC on 23 July 2026 – Service alerts fired and multiple Azure service teams started investigating the traffic anomalies.
- 14:54 UTC on 23 July 2026 – Automated recovery system started performing multiple retries to recover the devices that were affected by the over scoping that happened during the break-fix issue.
- 15:26 UTC on 23 July 2026 – Routing anomaly alert incorrectly triaged to WAN issue and was sent to engineers to investigate further.
- 16:58 UTC on 23 July 2026 – After multiple retries, automated recovery system raised an alert to engineers to investigate. Engineers tried to recover the devices manually.
- 17:19 UTC on 23 July 2026 – Engineers correlated the routing anomaly with the initial break-fix trigger event and began attempting auto-rollback.
- 17:45 UTC on 23 July 2026 – Engineers initiated manual rollback of the break-fix.
- 18:26 UTC on 23 July 2026 – Rollback was successful on all affected devices, engineers continued monitoring recovery of affected services.
- 19:41 UTC on 23 July 2026 – Customer impact confirmed as mitigated, after all services returned to expected operating levels.
How are we making incidents like this less likely or less impactful?
- We have fixed the defect in the blast radius analysis system that incorrectly expanded the scope of this ‘break-fix’ unplanned maintenance to include additional devices. Input validation now ensures that both diverse paths for a datacenter cannot undergo break-fix requests at the same time. (Completed)
- We have scanned all existing break-fix requests in the system to identify similar conditions that could have allowed the same defect to produce similar impact elsewhere. (Completed)
- We have enhanced safety checks to block any break-fix or maintenance activity that would affect all diverse paths for a datacenter, to de-risk this failure mode. (Completed)
- We are improving change visibility, break-fix requests will be recorded in the standard Azure change ledger, to reduce the time required to correlate maintenance activity with network issues during future investigations. (Estimated completion: August 2026)
- We are enhancing our datacenter traffic monitoring to detect abnormalities in inbound and outbound traffic more quickly, and surface these directly to engineers for evaluation – enabling faster customer notification, and mitigation efforts. (Estimated completion: September 2026)
- We are improving change tooling guardrails so that multiple devices cannot be isolated simultaneously during maintenance. (Estimated completion: September 2026)
- In the longer term, we are improving our automated recovery system to fail more quickly and escalate to engineers more promptly when retries are unsuccessful, rather than continuing to retry against disrupted connectivity. (Estimated completion: October 2026)
- Finally, we are improving our automated rollback so it can recover more robustly, even when the underlying connectivity on which it depends is degraded. (Estimated completion: October 2026)
How can customers make incidents like this less impactful?
- For mission-critical workloads, customers should consider a multi-region deployment strategy to maintain availability during events that affect a single region: https://learn.microsoft.com/azure/architecture/patterns/geodes and https://learn.microsoft.com/azure/well-architected/design-guides/regions-availability-zones
- More generally, consider evaluating the reliability of your applications using guidance from the Azure Well-Architected Framework and its interactive Well-Architected Review: https://aka.ms/AzPIR/WAF
- The impact times represent the full incident duration, so are not specific to any individual customer. Actual impact to service availability varied between customers and resources. For guidance on implementing monitoring to understand granular impact: https://aka.ms/AzPIR/Monitoring
- Finally, consider ensuring that the right people in your organization will be notified about any future service issues by configuring Azure Service Health alerts. These can trigger emails, SMS, push notifications, webhooks, and more: https://aka.ms/ash-alerts
How can we make our incident communications more useful?
You can rate this PIR and provide any feedback using our quick 3-question survey: https://aka.ms/AzPIR/ZJV6-SGG